蚂蚁百灵发布新一代原生混合推理模型Ling-3.0-Flash
7月24日,蚂蚁百灵正式发布新一代原生混合推理模型 Ling-3.0-Flash。该模型总参数量 124B,单次计算激活 5.1B。在大幅精简体量与降低算力开销的同时,Ling-3.0-Flash 在基础推理、指令遵循及长文本处理等核心指标上,对标甚至超越了参数规模达其 2 至 3 倍的行业领先模型,展现出更高的“智效比”与落地性价比。据了解,该模型已上线OpenRouter,限时免费一周,之后将正式开源。

作为该版本的重点,Ling-3.0-Flash 针对Agent的实际应用场景进行了深度打磨。模型扩展了超 10000 个真实交互训练环境,并优化了复杂任务的自我纠错与长程规划机制。无论是在代码编写、复杂任务拆解,还是多源信息深度调研等场景中,Ling-3.0-Flash 均展现出更强的自主闭环交付能力,解决了传统模型在大任务中容易“跑偏”或断档的难题。
据介绍,实现“小体量、高智能”的关键,源于模型底层计算架构的重新设计。Ling-3.0-Flash 放弃了单纯堆砌参数的思路,从预训练阶段即采用混合注意力架构,以 5:1 的比例交替堆叠 KDA 线性注意力层与 MLA 层,让模型在长上下文效率与模型能力之间实现更优平衡。
此外,Ling-3.0-Flash将上一代的 Lightning Attention 升级为 KDA(Kimi Delta Attention),在 Delta Rule 的状态更新中引入细粒度对角门控,让模型在处理长文档和代码库时能更精准地记住关键信息。同时,Ling-3.0-Flash进一步压缩了单次计算的专家激活比例,每个 Token 的专家激活比例由前代的 1/32 进一步压缩至 1/64,并由此带来了更高的“效率杠杆”。

为了让 AI Agent 运行更快、更稳,百灵还为其匹配了配套的工程与协作架构。在响应速度上,通过引入集群级分级缓存,避免了长对话和多轮交互中的重复计算,将长输入下的首字响应延迟降低了 60% 至 80% 以上;在任务稳定性上,升级后的多智能体协同架构允许不同 Agent 分工协作、相互校验,有效减少了单一模型的误判风险,为高频线上服务与真实业务落地提供了支撑。
本文由蚂蚁百灵提供,量子位获授权转载,观点归原作者所有。
版权所有,未经授权不得以任何形式转载及使用,违者必究。
Related Articles
AI leaders sign statement asking the government to do something about automated AI
Employees of OpenAI and Anthropic, as well as Google, Meta, Thinking Machines, Microsoft, Mistral, and other leading AI labs, have written a statement to the US government supporting a potential slowdown of...
AI’s finally expensive enough to make Wall Street nervous
It’s earnings season, and investors got an unpleasant surprise from Google: an increase on its spending estimate, to as much as $205 billion — from the last quarter’s projection of up to $190 billion. Even...
Perplexity’s Personal Computer turns Windows PCs into AI agents
Jess Weatherbed is a news writer focused on creative industries, computing, and internet culture. Jess started her career at TechRadar, covering news and hardware reviews.Perplexity has expanded its agentic...
Smart rings are looking like my kind of AI gadget
Over the last few months, I’ve spent a lot of time talking to my computer. One underrated feature of the LLM revolution has been a remarkable leap in all kinds of dictation technology — even the fastest,...