A silent assistant listens along and may show one card on the phone. Every card is scored against a hand-labelled need window with the real wall clock: transcript lag, gate, model latency, push. A card that arrives after the need is gone counts against the model.
1 · Speech to text
+15 s
transcript lag (streaming ASR)
→
2 · Gate
0.3 s
jev-latest (TypeSafe.ai), every 30 s on the last 90 s
→
3 · Card model
10 to 409 s
only when the gate fires
→
4 · Push
+5 s
phone out of the pocket
=
Speech to visible card
30 to 429 s
anything a person finishes faster is out of reach
Recordings
| Recording | Status | Gate ticks | Needs | Best model | Decisions | |
|---|---|---|---|---|---|---|
◈ Evening (Cafeteria), Sep 20, 2026Late evening in a cafeteria, loud surroundings, German. |
Evaluated | 24 of 149 | 5 long · 3 short | gpt-5.6-sol3 of 5 in time | 64 cards · 80 silent | |
◈ Evening (Cafeteria) — Stan's pipelineThe same evening recording, processed with Stan's pipeline. Speakers are only distinguished within each section. | Not evaluated | – | – | – | – | |
◈ Tenten Founder Meeting, Sep 20, 2026Office conversation, mostly English. Four known voices and one new person. | Not evaluated | – | – | – | – | |
◈ Dinner (Indian), Sep 17, 2026Five people at the table. German with English mixed in, lots of talking over each other. | Not evaluated | – | – | – | – |
Models across evaluated recordings
| Model | Needs in time | Too late | No need | Repeat | Silent | Latency p50 | Latency max | Cost |
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-solopenai | 3 / 5 | 0 | 1 | 1 | 15 / 24 | 9.4 s | 26.3 s | $0.478 |
| gpt-5.6-lunaopenai | 3 / 5 | 7 | 2 | 0 | 9 / 24 | 6.9 s | 9.1 s | $0.056 |
| gemini-3.8-flashgoogle | 2 / 5 | 0 | 0 | 0 | 22 / 24 | 9.1 s | 21.4 s | $0.232 |
| kimi-k3moonshotai | 2 / 5 | 2 | 1 | 2 | 15 / 24 | 41.8 s | 97.3 s | $1.523 |
| glm-5.3-flashz-ai | 2 / 5 | 2 | 2 | 2 | 12 / 24 | 15.8 s | 408.7 s | $0.026 |
| deepseek-v4.1-flashdeepseek | 2 / 5 | 10 | 3 | 3 | 5 / 24 | 6.8 s | 103.8 s | $0.045 |