Skip to content

Accuracy, published

How well SQUELCH understands you

A wrong runway or callsign was caught in 382 of 384 deliberately-wrong readbacks, and clean recognition ran from 83 to 96 percent across eight accents.

Measured on 2026-07-28 over 768 synthetic takes, scored entirely on device.

How we measure

You speak a radio call. Your iPhone transcribes it on device with Apple's speech framework, with no cloud and no account.

The transcript is normalized (niner becomes nine, common mishearings are recovered), aligned against the expected call slot by slot, and each slot is graded on its own.

A wrong runway or callsign is never smoothed over. Catching the dangerous mistake is the point, so detection is reported as its own number.

The on-device neural voice you hear in the app is ATC playback and plays no part in these figures: the numbers below measure the grader scoring recorded speech, a separate path.

Conditions

  • Clean. Quiet-room speech, no added noise.
  • Cabin noise. A noise bed at roughly 10 dB signal-to-noise, calibrated to a realistic loud cabin or room as heard through a phone with typical noise suppression.

Synthetic robustness by accent

Voice (accent)CleanCabin noiseWrong runway/callsign caught
US General American96%92%100% (48/48)
US Southern85%74%98% (47/48)
British RP94%92%100% (48/48)
Irish89%78%100% (48/48)
Indian English92%72%100% (48/48)
East Asian83%79%98% (47/48)
Latino91%80%100% (48/48)
Australian94%84%100% (48/48)
Overall90%82%99% (382/384)

These figures come from synthetic text-to-speech voices across eight accents. Synthetic voices are cleaner and steadier than real speech and tend to overstate what a real speaker would score, so read this as a robustness signal across accents, not a promise of your personal accuracy.

Real-speaker figures, recorded on device across a range of voices and noise, are in testing and will be published here alongside these.

Reproduce it

We resolve each scenario through the same engine the app uses, synthesize the readback in eight accents, mix a cabin-noise bed, and score every take on device with the shipping grader.

The corpus metadata (which call, which accent, which condition, and the expected slots) is kept so the table can be regenerated and checked.

Why we publish this

We publish how well SQUELCH understands you and how we measure it. Most radio trainers publish nothing.