256 tok/s
kimi-k2.7-code-highspeed
fastest model on the board — 6× the z.ai field, and it's already the primary
| Model |
↓ tok/s |
|
TTFT |
Tools |
| KIMIk2.7-code-highspeed |
256
|
|
1.1s |
YES |
| DSv4-flash |
112
|
|
0.6s |
YES |
| KIMIk2.7-code |
94
|
|
0.9s |
YES |
| KIMIk2.6 (legacy) |
76
|
|
0.8s |
YES |
| DSv4-pro |
50
|
|
0.4s |
YES |
| Z.AIglm-4.7 |
40
|
|
1.3s |
YES |
| KIMIk3 (2.8T flagship) |
39
|
|
2.1s |
YES |
| Z.AIglm-5.3-flash |
37
|
|
1.2s |
YES |
| Z.AIglm-5-turbo |
31
|
|
1.6s |
YES |
Reasoning effort — glm-4.7
Same problem at each effort level. Tool for the trade: what thinking costs you.
| Effort | Reply | Think tok | Verdict |
| none | 3.1s | 39 | 4.7 thinks ~40 tok regardless — can't fully disable |
| low ★ | 3.0s | 38 | free over none — the right standing default |
| medium | 7.8s | 38 | pays double latency, buys nothing |
| high | 7.2s | 169 | 4× tokens for the same answer — save for hard problems |
Recommendations
Fastest overall
kimi-k2.7-highspeed · 256 tok/s
Already the primary routing. Light-reasoning tier — right for chat and routine code.
Best on z.ai
glm-4.7 · ~40 tok/s
Most consistent GLM, cheapest on Coding Plan credits, tools verified.
Deep work
kimi-k3 · ds-v4-pro
39–50 tok/s but deeper reasoning. Swap in for hard design & analysis, then out.
Tool calling
9 / 9 pass
Every model returned a clean function call. k3 was sharpest — inferred city + unit from context.
Drift since early Sep
Same benchmark re-run ~3 weeks apart — everything got faster.
| glm-5.3-flash | 15 → 37 tok/s | 2.5× |
| kimi-k2.7-code | 43 → 94 tok/s | 2.2× |
| deepseek-v4-flash | 97 → 112 tok/s | 1.2× |
| kimi-k2.7-highspeed | ~250 → 256 tok/s | ≈ |
| kimi-k3 | 39 → 39 tok/s | ≈ |