⚡响应最快(首字延迟最低)的 AI 模型排行
按首字延迟 TTFT(Time To First Token,从发起请求到收到第一个 token 的秒数)从低到高排序,越低代表用户感知的「卡顿感」越小。
适合对话类产品、实时客服、语音交互等对首字延迟敏感的场景。
数据日期 2026-09-23 · 共 20 个模型上榜 · 当前第一名 Command A+(首字延迟 (TTFT) 0.18s)
| # | 模型 | 厂商 | 首字延迟 (TTFT) | 官方输入价 | 中转站最低可信价 |
|---|---|---|---|---|---|
| 1 | Command A+ | Cohere | 0.18s | 暂未公布 | — |
| 2 | North Mini Code | Cohere | 0.20s | 暂未公布 | — |
| 3 | Tiny Aya Global | Cohere | 0.23s | 暂未公布 | — |
| 4 | Granite 4.2 3BNEW | IBM | 0.24s | $0.03/M | — |
| 5 | Nemotron 3 Nano Omni 30B A3B Reasoning | NVIDIA | 0.28s | $0.195/M | — |
| 6 | Granite 4.2 30BNEW | IBM | 0.30s | $0.16/M | — |
| 7 | Command A | Cohere | 0.34s | $2.5/M | — |
| 8 | Phi-4 Mini Instruct | Microsoft | 0.35s | 暂未公布 | — |
| 9 | Granite 4.2 8BNEW | IBM | 0.36s | $0.06/M | — |
| 10 | Phi-4 Multimodal Instruct | Microsoft | 0.36s | 暂未公布 | — |
| 11 | Muse Glimmer | Meta | 0.36s | $0.35/M | — |
| 12 | Gemma 4 E4B | 0.37s | $0.02/M | — | |
| 13 | NVIDIA Nemotron 3 Nano 30B A3B | NVIDIA | 0.41s | $0.05/M | — |
| 14 | Nemotron 3.5 Lightning | NVIDIA | 0.45s | $0.06/M | — |
| 15 | Ministral 3 8B | Mistral | 0.47s | $0.15/M | — |
| 16 | Celeris-1 | Celeris | 0.49s | $0.2/M | — |
| 17 | Mistral Small 4 | Mistral | 0.51s | $0.15/M | — |
| 18 | Claude 4.5 Haiku | Anthropic | 0.51s | $1/M | — |
| 19 | gpt-oss-20b | OpenAI | 0.52s | $0.07/M | — |
| 20 | gpt-oss-120b | OpenAI | 0.52s | $0.15/M | — |
#1Command A+
Cohere官方 暂未公布暂未收录中转站
#2North Mini Code
Cohere官方 暂未公布暂未收录中转站
#3Tiny Aya Global
Cohere官方 暂未公布暂未收录中转站
#4Granite 4.2 3BNEW
IBM官方 $0.03/M暂未收录中转站
#5Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA官方 $0.195/M暂未收录中转站
#6Granite 4.2 30BNEW
IBM官方 $0.16/M暂未收录中转站
#7Command A
Cohere官方 $2.5/M暂未收录中转站
#8Phi-4 Mini Instruct
Microsoft官方 暂未公布暂未收录中转站
#9Granite 4.2 8BNEW
IBM官方 $0.06/M暂未收录中转站
#10Phi-4 Multimodal Instruct
Microsoft官方 暂未公布暂未收录中转站
#11Muse Glimmer
Meta官方 $0.35/M暂未收录中转站
#12Gemma 4 E4B
Google官方 $0.02/M暂未收录中转站
#13NVIDIA Nemotron 3 Nano 30B A3B
NVIDIA官方 $0.05/M暂未收录中转站
#14Nemotron 3.5 Lightning
NVIDIA官方 $0.06/M暂未收录中转站
#15Ministral 3 8B
Mistral官方 $0.15/M暂未收录中转站
#16Celeris-1
Celeris官方 $0.2/M暂未收录中转站
#17Mistral Small 4
Mistral官方 $0.15/M暂未收录中转站
#18Claude 4.5 Haiku
Anthropic官方 $1/M暂未收录中转站
#19gpt-oss-20b
OpenAI官方 $0.07/M暂未收录中转站
#20gpt-oss-120b
OpenAI官方 $0.15/M暂未收录中转站
常见问题
这个榜单是怎么排序的?
按首字延迟 TTFT(Time To First Token,从发起请求到收到第一个 token 的秒数)从低到高排序,越低代表用户感知的「卡顿感」越小。适合对话类产品、实时客服、语音交互等对首字延迟敏感的场景。
数据多久更新一次?
模型库(Artificial Analysis)每日同步,本榜单随每日巡检自动重新计算,当前数据日期 2026-09-23。
为什么有些知名模型没有出现在榜单里?
数据源没有提供该指标的真实跑分时,我们不会用 0 或估算值顶替——缺数据的模型直接不参与该场景的排序,避免用假分数误导。
同一模型有 High / Max Effort 等多个推理档位,算的是哪一档?
同一模型的不同推理档位(low/medium/high/xhigh 等)会合并为一行,按其中表现最高的档位计入本榜单,不会让同一模型的多个档位重复占位刷榜。
数据来自 Artificial Analysis 真实基准测试,排序规则见评分方法论。价格与跑分随官方更新可能变化,请以模型服务商公告为准。