Passes out of 30 by outcome verification (the device check), and the four cells once the transcript judge is added. Same tasks, same grader, same simulated user; the Apple control was re-run the same day as the arm it is compared with.
| configuration | pass / 30 | verified success | visible failure | hallucinated success | silent misconduct |
|---|---|---|---|---|---|
| Text · Apple model, typed | 13 | 10 | 13 | 4 | 3 |
| Cascaded · Apple speech → Apple model | 11 | 9 | 12 | 7 | 2 |
| Cascaded · MiniCPM speech → Apple model | 11 | 9 | 15 | 4 | 2 |
| Speech-to-speech · MiniCPM-o 4.5, hears + thinks | 16 | 11 | 12 | 2 | 5 |
| Speech-to-speech · MiniCPM-o 4.5, also speaks | 15 | 10 | 11 | 4 | 5 |
The transcript judge reads only the words, never the device; it is told that on this phone a message is sent by placing a draft. Two judging passes agree on 92% of verdicts; read those cells as ±2–3. The speech-to-speech model is also a larger model than Apple's; five of its seven gains are tasks the smaller model fails even when typed.