Field notes · 17 Aug 2026
The audible clock, or: how to judge a dance contest over video call
The central idea of our white paper, explained without maths: score the beat the player heard, not the beat the software scheduled.
Every rhythm game lives with the same physics. Between the moment software decides "the beat happens now" and the moment sound leaves your speaker, there is a pipeline — audio buffers, the operating system's mixer, maybe a Bluetooth hop to your earbuds. On wired audio that pipeline costs around 30 milliseconds. On Bluetooth it can cost 300. You never notice it, because you have nothing to compare it to. But you play to the far end of it: players tap what they hear.
Here's the trap. If the game records your tap against the clock the software used for scheduling, every honest tap arrives "late" by exactly the length of your device's pipeline. Play perfectly on Bluetooth earbuds and you can miss every single beat — not because you were off, but because the referee was listening to a different clock than you were.
The dance contest
Imagine judging a dance contest over video call, and scoring every dancer against the music playing in your room. The dancers with fast connections do fine. The dancers with slow connections all look off-beat — every step landing a beat behind your music, with perfect consistency. Are they bad dancers? Obviously not. The music reached them late, they danced to what reached them, and your scoring never asked the only fair question: were they on the beat of the music as it reached them?
Swap the video call for an audio pipeline and that is exactly what most timing scoring does — and why it quietly punishes whole categories of hardware.
Scoring the heard beat
OneDrum's answer is to stop asking the scheduling clock what time it is. The audio system itself can report which sample is leaving the speaker at which moment — a running statement of "what the player is hearing right now." From that report we derive the audible clock: a timeline pinned to the sound as heard, not the sound as queued.
Both sides of the comparison move onto that clock. The cue's time is when the cue was audible. The tap's time is the input event's own timestamp, converted onto the same timeline. Once cue and tap share a reference frame, the pipeline cancels out of the subtraction entirely: a tap on the heard beat scores as on the beat, whether the sound took 30 milliseconds to arrive or 300.
What it doesn't solve — on purpose
The audible clock removes the device's systematic delay. It deliberately says nothing about the player's — your personal tendency to answer a shade behind the beat, or your touchscreen's input lag. That is handled by calibration with no calibration step: the first bar of play doubles as the measurement, silently (white paper, §4). And when a device's reported latency wanders mid-session — which our production data caught Android doing — drift re-centring keeps the correction honest. The one thing none of this touches is the anti-spoof gate, which is computed on values no calibration can shift, so compensating for slow hardware never becomes a way to sneak a bot past the referee.
That's the whole idea. The referee and the player finally listen to the same clock — everything else in the engine is making sure that clock stays true.
Hear it for yourself — the drum calls, you answer: onedrum.io
The full mechanism, with the maths and the field data, is in the white paper. この記事の日本語版はこちら。 · More field notes