🚀生成速度(吞吐)最快的 AI 模型排行
按生成吞吐(每秒输出 token 数)从高到低排序,反映模型把长回答「吐完」的速度,与首字延迟是两个独立维度。
适合长文生成、批量处理、需要尽快拿到完整输出的场景。
数据日期 2026-09-23 · 共 20 个模型上榜 · 当前第一名 Celeris-1(生成速度 1892.8 tok/s)
| # | 模型 | 厂商 | 生成速度 | 官方输入价 | 中转站最低可信价 |
|---|---|---|---|---|---|
| 1 | Celeris-1 | Celeris | 1892.8 tok/s | $0.2/M | — |
| 2 | Mercury 2 | Inception | 746.5 tok/s | $0.25/M | — |
| 3 | Ling 3.0 Flash | InclusionAI | 371.6 tok/s | $0.075/M | — |
| 4 | Gemini 3.5 Flash-Lite | 350.4 tok/s | $0.3/M | ¥¥1.47/M (1折) | |
| 5 | Trinity Large Thinking | Arcee AI | 335.2 tok/s | $0.25/M | — |
| 6 | Gemini 3.8 FlashNEW | 315.7 tok/s | $0.75/M | ¥¥3.675/M (1折) | |
| 7 | Nemotron 3.5 Lightning | NVIDIA | 276.9 tok/s | $0.06/M | — |
| 8 | Nemotron 3 Nano Omni 30B A3B Reasoning | NVIDIA | 271.2 tok/s | $0.195/M | — |
| 9 | gpt-oss-20b | OpenAI | 268.8 tok/s | $0.07/M | — |
| 10 | Nova Micro | Amazon | 241.2 tok/s | $0.035/M | — |
| 11 | Command A+ | Cohere | 239.0 tok/s | 暂未公布 | — |
| 12 | DeepSeek V4 Flash Vision | DeepSeek | 237.2 tok/s | $0.44/M | — |
| 13 | Qwen3.5 Omni Flash | Alibaba | 234.9 tok/s | $0.1/M | — |
| 14 | Ministral 3 3B | Mistral | 234.1 tok/s | $0.1/M | — |
| 15 | Granite 4.2 3BNEW | IBM | 232.4 tok/s | $0.03/M | — |
| 16 | DeepSeek V4.1 FlashNEW | DeepSeek | 230.9 tok/s | $0.3/M | — |
| 17 | Inkling Small | Thinking Machines | 224.8 tok/s | $0.3/M | — |
| 18 | Nova 2.0 Lite | Amazon | 224.5 tok/s | $0.3/M | — |
| 19 | Muse Spark 1.3NEW | Meta | 221.0 tok/s | $1.25/M | — |
| 20 | gpt-oss-120b | OpenAI | 213.6 tok/s | $0.15/M | — |
#1Celeris-1
Celeris官方 $0.2/M暂未收录中转站
#2Mercury 2
Inception官方 $0.25/M暂未收录中转站
#3Ling 3.0 Flash
InclusionAI官方 $0.075/M暂未收录中转站
#4Gemini 3.5 Flash-Lite
Google官方 $0.3/M中转最低 ¥¥1.47/M (1折)
#5Trinity Large Thinking
Arcee AI官方 $0.25/M暂未收录中转站
#6Gemini 3.8 FlashNEW
Google官方 $0.75/M中转最低 ¥¥3.675/M (1折)
#7Nemotron 3.5 Lightning
NVIDIA官方 $0.06/M暂未收录中转站
#8Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA官方 $0.195/M暂未收录中转站
#9gpt-oss-20b
OpenAI官方 $0.07/M暂未收录中转站
#10Nova Micro
Amazon官方 $0.035/M暂未收录中转站
#11Command A+
Cohere官方 暂未公布暂未收录中转站
#12DeepSeek V4 Flash Vision
DeepSeek官方 $0.44/M暂未收录中转站
#13Qwen3.5 Omni Flash
Alibaba官方 $0.1/M暂未收录中转站
#14Ministral 3 3B
Mistral官方 $0.1/M暂未收录中转站
#15Granite 4.2 3BNEW
IBM官方 $0.03/M暂未收录中转站
#16DeepSeek V4.1 FlashNEW
DeepSeek官方 $0.3/M暂未收录中转站
#17Inkling Small
Thinking Machines官方 $0.3/M暂未收录中转站
#18Nova 2.0 Lite
Amazon官方 $0.3/M暂未收录中转站
#19Muse Spark 1.3NEW
Meta官方 $1.25/M暂未收录中转站
#20gpt-oss-120b
OpenAI官方 $0.15/M暂未收录中转站
常见问题
这个榜单是怎么排序的?
按生成吞吐(每秒输出 token 数)从高到低排序,反映模型把长回答「吐完」的速度,与首字延迟是两个独立维度。适合长文生成、批量处理、需要尽快拿到完整输出的场景。
数据多久更新一次?
模型库(Artificial Analysis)每日同步,本榜单随每日巡检自动重新计算,当前数据日期 2026-09-23。
为什么有些知名模型没有出现在榜单里?
数据源没有提供该指标的真实跑分时,我们不会用 0 或估算值顶替——缺数据的模型直接不参与该场景的排序,避免用假分数误导。
同一模型有 High / Max Effort 等多个推理档位,算的是哪一档?
同一模型的不同推理档位(low/medium/high/xhigh 等)会合并为一行,按其中表现最高的档位计入本榜单,不会让同一模型的多个档位重复占位刷榜。
数据来自 Artificial Analysis 真实基准测试,排序规则见评分方法论。价格与跑分随官方更新可能变化,请以模型服务商公告为准。