Benchmark Board

LLM Speed & Tool Calling

9 models · 3 providers · measured live from the OVH VPS · Sep 6, 2026
256 tok/s
kimi-k2.7-code-highspeed fastest model on the board — 6× the z.ai field, and it's already the primary
Model ↓ tok/s TTFT Tools
KIMIk2.7-code-highspeed 256
1.1s YES
DSv4-flash 112
0.6s YES
KIMIk2.7-code 94
0.9s YES
KIMIk2.6 (legacy) 76
0.8s YES
DSv4-pro 50
0.4s YES
Z.AIglm-4.7 40
1.3s YES
KIMIk3 (2.8T flagship) 39
2.1s YES
Z.AIglm-5.3-flash 37
1.2s YES
Z.AIglm-5-turbo 31
1.6s YES

Reasoning effort — glm-4.7

Same problem at each effort level. Tool for the trade: what thinking costs you.
EffortReplyThink tokVerdict
none3.1s394.7 thinks ~40 tok regardless — can't fully disable
low ★3.0s38free over none — the right standing default
medium7.8s38pays double latency, buys nothing
high7.2s1694× tokens for the same answer — save for hard problems

Recommendations

Fastest overall
kimi-k2.7-highspeed · 256 tok/s

Already the primary routing. Light-reasoning tier — right for chat and routine code.

Best on z.ai
glm-4.7 · ~40 tok/s

Most consistent GLM, cheapest on Coding Plan credits, tools verified.

Deep work
kimi-k3 · ds-v4-pro

39–50 tok/s but deeper reasoning. Swap in for hard design & analysis, then out.

Tool calling
9 / 9 pass

Every model returned a clean function call. k3 was sharpest — inferred city + unit from context.

Drift since early Sep

Same benchmark re-run ~3 weeks apart — everything got faster.
glm-5.3-flash15 → 37 tok/s2.5×
kimi-k2.7-code43 → 94 tok/s2.2×
deepseek-v4-flash97 → 112 tok/s1.2×
kimi-k2.7-highspeed~250 → 256 tok/s
kimi-k339 → 39 tok/s