Kaseta
Menu

Offline transcription that never uploads your audio

Offline transcription means the speech recognition runs on the device that holds the audio. Not “encrypted in transit”. Not “deleted after 30 days”. The recording is never sent, because the work happens where it already is. Kaseta transcribes 99 languages this way on an iPhone – models downloaded once, then recognition on the phone after you tap Stop – and the difference from a cloud service is not a promise. It is a network request that does not exist.

Three things offline has to mean

It works with the radios off. There is no account, because an account implies a server that knows you. And nothing is queued for a quiet upload later, once you are back on Wi-Fi. Any one of those missing and “offline” is a mode, not an architecture. Kaseta’s recognition has no server behind it, which is a stronger statement than a promise not to use one: the models live on the phone after a one-time download, and the audio has nowhere to go.

What is actually running on the phone

Speech recognition models, downloaded once and executed by the phone’s own hardware. Fast is the everyday one; Quality is a larger model for difficult audio, with a monthly allowance on the free plan and none on Pro. Both live on the device once downloaded, and whether Quality is available depends on the iPhone. Neither of them phones home, which is also why the app has nothing useful to log about you.

What it costs you to do it this way

Honest ledger. Long files take real time and real battery, and a five-year-old iPhone will be slower at it than a current one. There is no web dashboard, no team workspace, and no bot that joins your meetings – a call with remote participants reaches the phone only as whatever the room’s speakers play back. And the free plan transcribes the first 5 minutes of each recording, though the recording itself is never cut short – audio keeps going and is saved in full.

When a cloud service is the better tool

Why there is no subscription

The costs of a cloud transcription service scale with your minutes, so it charges you by the minute. Ours do not exist, so we do not meter them. Pro is $19.99 a year or $39.99 once, and the once-priced version is the point rather than an upsell.

Next: how to transcribe a voice memo is the practical version of this page, and the homepage has the short version of the same argument.

Steps

  1. 01

    Let the app download its models once

    The first launch fetches the recognition model over the internet. That is the only step that needs a network, and it happens once.

  2. 02

    Put the phone in airplane mode

    Open Control Centre and switch the radios off, Wi-Fi included. The point is to leave the app with nowhere to send anything, not to trust that it would not.

  3. 03

    Record something and tap Stop

    Press record and talk, then stop. There is no session to open with a server first, and no text while you speak – the phone starts on the transcript when you stop.

  4. 04

    Wait for the text – it arrives anyway

    The transcript is computed by the phone after you stop. If the work were happening in a data centre, this is the step that would fail.

  5. 05

    Turn the network back on

    Nothing uploads late, because nothing was queued and there is no account for it to sync to. The test you just ran is the whole architecture.

Frequently asked questions

Can I transcribe audio offline?

Yes. Offline transcription means the speech recognition model runs on the device holding the audio instead of on a remote server, so it works in airplane mode, on a plane, or in a building with no signal. Kaseta does this on iPhone; on a desktop, a local Whisper build does the same thing.

What is on-device transcription?

It is speech recognition performed by the phone's own processor, using models downloaded once and kept on the device, with no network request in the loop. The practical difference from a cloud service is not the wording of a privacy policy – it is that the upload never happens, which is something you can verify yourself by cutting the network.

Is offline transcription as accurate as the cloud?

Usually close, sometimes behind. Cloud services can run very large models on server hardware, which shows on hard audio: heavy accents, crosstalk, a noisy room. Kaseta has a Fast model and a Quality model, both on the phone – Quality has a monthly allowance on the free plan and none on Pro – and on clear speech recorded near the phone the gap is small enough not to matter.

Do I need an account to use it?

No. There are no accounts, because there is no server to hold one. Nothing to sign up for, no email address to hand over, and no cookies on this website either.

Related reading

Get Kaseta on the App Store