| Model | Needs in time | Too late | No need | Repeat | Silent | Latency | Cost / run |
|---|---|---|---|---|---|---|---|
| gpt-5.6-solopenai · default | 3 / 5Recording cut, Follow-ups, Unit economics | 0 | 1 | 1 | 15 / 24 | p50 9.4 · p90 18.3 · max 26.3 s | $0.478 |
| gemini-3.8-flashgoogle · default | 2 / 5Recording cut, Unit economics | 0 | 0 | 0 | 22 / 24 | p50 9.1 · p90 15.9 · max 21.4 s | $0.232 |
| kimi-k3moonshotai · default | 2 / 5Recording cut, Unit economics | 2 | 1 | 2 | 15 / 24 | p50 41.8 · p90 80.6 · max 97.3 s | $1.523 |
| glm-5.3-flashz-ai · reasoning effort low | 2 / 5Follow-ups, Unit economics | 2 | 2 | 2 | 12 / 24 | p50 15.8 · p90 49.2 · max 408.7 s | $0.026 |
| gpt-5.6-lunaopenai · default | 3 / 5Recording cut, Menu, Unit economics | 7 | 2 | 0 | 9 / 24 | p50 6.9 · p90 8.4 · max 9.1 s | $0.056 |
| deepseek-v4.1-flashdeepseek · reasoning off | 2 / 5Menu, Unit economics | 10 | 3 | 3 | 5 / 24 | p50 6.8 · p90 65.6 · max 103.8 s | $0.045 |
Ranked by needs in time, minus half a point per wrong card (too late or no labelled need), minus one point if p90 latency exceeds 20 s. Latency is the measured model round trip per decision. Cost is for all 24 decisions of this recording.