deepseek-flash (V4.1-Flash) · 60 cells · 180 API calls
· 10 scenarios × 6 levels × 3 passes · temperature 0 · deterministic graders| Level | Mean score | Worst cell | Perfect cells | Avg latency | Avg reasoning tok | Starved |
|---|---|---|---|---|---|---|
| none | 0.992 | 0.917 | 9/10 | 1.73s | 0 | 0 |
| minimal pick | 1.000 | 1.000 | 10/10 | 3.45s | 433 | 0 |
| low | 1.000 | 1.000 | 10/10 | 3.15s | 406 | 0 |
| medium | 1.000 | 1.000 | 10/10 | 3.76s | 531 | 0 |
| high | 1.000 | 1.000 | 10/10 | 4.21s | 576 | 0 |
| xhigh | 1.000 | 1.000 | 10/10 | 3.90s | 530 | 0 |
none (reasoning disabled), which dropped a pass on professional judgment.
That means for tasks at this difficulty, turning reasoning on at all is what matters —
and the specific tier above minimal is not distinguishable.
| Scenario | none | minimal | low | medium | high | xhigh | Verdict |
|---|---|---|---|---|---|---|---|
|
Multi-step math word problemReasoning |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Code generation to specCoding |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Find the bug in codeCoding |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Structured JSON extractionExtraction |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Support ticket classificationClassification |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Constrained customer replyWriting |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Faithful summarizationSummarization |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Multi-step logic deductionReasoning |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Strict format instruction-followingInstruction-following |
1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | all levels tie |
|
Professional judgment / boundaryJudgment |
0.92 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | minimal |
none averages
0.992 with visible variance (std across passes up to 0.12);
every thinking level averages 1.000 with zero variance.none and up to
4.2s at higher tiers — same quality, least thinking time.xhigh spent fewer tokens than high — effort sets a ceiling,
not a quota, so the level names are not a reliable cost knob.max_tokens=2048, a long thinking prompt returned
2048 reasoning tokens, finish_reason=length, and zero visible characters —
a silent empty response, HTTP 200, no error. The identical prompt at
max_tokens=8192 returned a full answer. Any integration that caps output tokens
tightly will intermittently return blanks as effort rises. Keep generous output headroom
(≥8K) whenever reasoning is enabled.
minimal effort with perfect reliability, and that
disabling reasoning is the only choice that measurably hurt. It does not prove the
levels are equivalent in general: with 3 passes and saturated scores, this design has no
headroom to detect a difference. Separating minimal from xhigh
would need substantially harder problems (long-context, multi-file code, novel math) where
the ceiling is not already reached.