Proactivity

A silent assistant listens along and may show one card on the phone. Every card is scored against a hand-labelled need window with the real wall clock: transcript lag, gate, model latency, push. A card that arrives after the need is gone counts against the model.

1 · Speech to text
+15 s
transcript lag (streaming ASR)
2 · Gate
0.3 s
jev-latest (TypeSafe.ai), every 30 s on the last 90 s
3 · Card model
10 to 409 s
only when the gate fires
4 · Push
+5 s
phone out of the pocket
=
Speech to visible card
30 to 429 s
anything a person finishes faster is out of reach

Recordings

RecordingStatusGate ticksNeedsBest modelDecisions
Evening (Cafeteria), Sep 20, 2026Late evening in a cafeteria, loud surroundings, German.
Evaluated24 of 1495 long · 3 short gpt-5.6-sol3 of 5 in time64 cards · 80 silent
Evening (Cafeteria) — Stan's pipelineThe same evening recording, processed with Stan's pipeline. Speakers are only distinguished within each section.
Not evaluated
Tenten Founder Meeting, Sep 20, 2026Office conversation, mostly English. Four known voices and one new person.
Not evaluated
Dinner (Indian), Sep 17, 2026Five people at the table. German with English mixed in, lots of talking over each other.
Not evaluated

Models across evaluated recordings

ModelNeeds in timeToo lateNo needRepeatSilentLatency p50Latency maxCost
gpt-5.6-solopenai3 / 501115 / 249.4 s26.3 s$0.478
gpt-5.6-lunaopenai3 / 57209 / 246.9 s9.1 s$0.056
gemini-3.8-flashgoogle2 / 500022 / 249.1 s21.4 s$0.232
kimi-k3moonshotai2 / 521215 / 2441.8 s97.3 s$1.523
glm-5.3-flashz-ai2 / 522212 / 2415.8 s408.7 s$0.026
deepseek-v4.1-flashdeepseek2 / 510335 / 246.8 s103.8 s$0.045