Offline transcription that never uploads your audio
Offline transcription means the speech recognition runs on the device that holds the audio. Not “encrypted in transit”. Not “deleted after 30 days”. The recording is never sent, because the work happens where it already is. Kaseta transcribes 99 languages this way on an iPhone – models downloaded once, then recognition on the phone after you tap Stop – and the difference from a cloud service is not a promise. It is a network request that does not exist.
Three things offline has to mean
It works with the radios off. There is no account, because an account implies a server that knows you. And nothing is queued for a quiet upload later, once you are back on Wi-Fi. Any one of those missing and “offline” is a mode, not an architecture. Kaseta’s recognition has no server behind it, which is a stronger statement than a promise not to use one: the models live on the phone after a one-time download, and the audio has nowhere to go.
What is actually running on the phone
Speech recognition models, downloaded once and executed by the phone’s own hardware. Fast is the everyday one; Quality is a larger model for difficult audio, with a monthly allowance on the free plan and none on Pro. Both live on the device once downloaded, and whether Quality is available depends on the iPhone. Neither of them phones home, which is also why the app has nothing useful to log about you.
What it costs you to do it this way
Honest ledger. Long files take real time and real battery, and a five-year-old iPhone will be slower at it than a current one. There is no web dashboard, no team workspace, and no bot that joins your meetings – a call with remote participants reaches the phone only as whatever the room’s speakers play back. And the free plan transcribes the first 5 minutes of each recording, though the recording itself is never cut short – audio keeps going and is saved in full.
When a cloud service is the better tool
- Meetings with a team. Otter and Notta send a bot into the call, take its audio feed at source instead of through a microphone in a room, and give everyone a shared searchable archive. That is a genuinely different product, and if it is what you need, on-device is the wrong shape entirely.
- A large archive on a computer. Hundreds of hours of interviews will process faster on a desktop, and a local Whisper build there is free and equally offline. A phone is the right tool for the recording you just made, not for a back catalogue.
- Collaboration on the text. If three people need to edit the transcript at once, you want something with a server in it. We do not have one, on purpose.
Why there is no subscription
The costs of a cloud transcription service scale with your minutes, so it charges you by the minute. Ours do not exist, so we do not meter them. Pro is $19.99 a year or $39.99 once, and the once-priced version is the point rather than an upsell.
Next: how to transcribe a voice memo is the practical version of this page, and the homepage has the short version of the same argument.