Majka Solutions ← Writing
Engineering

Shipping a voice translator that never touches a server

English 10 September 2026 ~9 min

ULAK is an iPhone app that lets two people who don't share a language hold a conversation through one phone. You pick two languages, hit start, and talk. It transcribes, works out which of the two languages was just spoken, translates, and speaks the result aloud.

It does all of that with the network switched off. Not "offline mode as a fallback" — there is no URLSession anywhere in the app. Zero. The whole thing runs on frameworks Apple already ships on the device.

That decision paid for itself in ways we expected (privacy, no inference bill, works on a boat) and cost us in ways we didn't. This is what we measured, what we built around, and what we'd tell you before you try the same thing.

Why on-device, honestly

The privacy argument is the one you put on the website. It's true — nothing said into the phone leaves it — but it isn't what made the decision.

The real reasons were duller. A cloud translator has a per-request cost, and a conversation app makes a request every few seconds; the unit economics of a €5 app with a chatty user are terrible. And the moment you have a server, you have a server: uptime, keys, a proxy so the API secret isn't in the bundle, a region choice, a data-processing agreement. For a two-person studio, "no backend" is not an aesthetic preference. It's the difference between shipping and not.

The privacy line turned out to matter commercially anyway, just not for the reason we assumed. Tourists don't ask about it. Businesses do.

What Apple's stack actually gives you

Before writing a line of product code we probed the frameworks and wrote the numbers down. If you're evaluating this, don't trust the marketing pages — the gap between "supported" and "supported on-device, in your language" is where the project lives or dies.

Speech recognition

SFSpeechRecognizer supported locales:  63
Turkish (tr-TR):                       supported
Turkish on-device recognition:         true
English on-device recognition:         true

Good news, and the foundation of the whole app. SFSpeechRecognizer is the older API, and it does on-device recognition for a wide set of locales at no cost and with no network.

The new stack, and why we couldn't use it

iOS 26 introduced SpeechAnalyzer / SpeechTranscriber — better API, better results. We measured its language list:

SpeechTranscriber supported locales:  30
Turkish:                              []

Turkish isn't there. Which killed our plan to build on the new API — and, in the same stroke, told us something more useful: if Apple's own new transcription stack doesn't do Turkish, then iOS 26's Live Translation doesn't either. Apple's published AirPods language list agrees.

The absence in the platform's own feature list was the market gap. The app exists because we checked an array's length instead of reading a press release.

So the architecture became: build on the older SFSpeechRecognizer, keep the newer engine behind the same protocol, and let both run where both are available.

Translation

Translation framework locales:  38
Turkish (tr-Latn-TR):           supported
tr → en language pack:          installed
en → tr language pack:          installed

The critical finding wasn't in the language list, it was in the SDK header:

convenience public init(installedSource: Locale.Language,
                        target: Locale.Language?)

TranslationSession is documented in a SwiftUI context, and every sample wires it to a view via .translationTask. But if the language pack is already installed, that initialiser lets you construct a session straight from code and hold it for the lifetime of the conversation. That's the difference between a coordinator that owns its session and a view hierarchy that owns your architecture.

One caveat that will bite you: if the pack is not installed, you still have to go through the SwiftUI download flow. So the first-run experience has to handle it, and — see below — that flow is not fast.

Speed

40 sequential translations, one session, one language pair:
total 14.11 s   ·   mean 353 ms   ·   errors 0   ·   no rate limiting

353 ms, on-device, no network, no quota. Faster than most round trips to a cloud endpoint, and it doesn't degrade on hotel Wi-Fi.

Where it stops

Here is the part that doesn't make it into most write-ups. We ran a fixed set of sentences through the translator and graded them by hand:

Turkish inputOutputVerdict
Yarın saat kaçta buluşalım?What time should we meet tomorrow?Correct
Hesabı alabilir miyim lütfen?Can I take the account please?Wrong
Dişimde üç gündür geçmeyen bir ağrı var…I have a pain in my jaw…Wrong
Saç ekimi için randevumu bir hafta erteleyebilir miyiz?Can we postpone my hair transplantation appointment by a week?Correct

"Hesap" is both an account and a restaurant bill; the model picked the wrong one in the one context where a tourist would say it. "Diş" (tooth) became "jaw" — fine in a novel, not fine when you're describing pain to a dentist.

The conclusion we wrote down at the time still holds: good enough for tourist conversation, not good enough for medical or commercial terminology. That single sentence set the product roadmap. It's why there's a glossary layer — a domain dictionary that overrides the general model for terms that matter in a specific setting — rather than a promise that the translation is always right.

Text-to-speech has the same shape of limitation. Turkish ships with exactly one voice:

AVSpeechSynthesizer Turkish voices:  1  (Yelda, tr-TR, quality: compact)
English (US) voices:                28

Compact is the low-quality tier. The enhanced voice exists but the user has to download it from iOS Settings — you cannot do it for them. If a large part of your product's felt quality is how the output sounds, and your language gets one compact voice, that is a first-run problem, not a nice-to-have.

The problem nobody warns you about: the phone hears itself

A translator speaks out loud, through the same phone whose microphone is listening for the next sentence. Unless you stop it, the app transcribes its own voice, decides someone just spoke the other language, translates that, speaks that, and you have built an infinite loop with a speaker.

We solved it in three layers, because no single layer holds.

Layer one is the type system. The conversation is a single state enum, not a set of booleans:

idle → preparing → listening → finalizing
     → arbitrating → translating → speaking
     → resuming → listening …

The invariant that matters is that speaking and listening can never both be active. With separate isSpeaking / isListening flags that invariant can break at runtime; with one enum it cannot be expressed at all. Note resuming: there is always a deliberate gap between the speaker stopping and the microphone reopening.

Layer two is the hardware — the audio session's acoustic echo cancellation, which handles most of what leaks through.

Layer three is a text filter, for the tail that still gets through when the speaker's decay overlaps the moment we reopen the mic. After speaking, we hold on to what we just said. The first transcript after returning to listening is compared against it — normalised Jaccard similarity, threshold 0.7, five-second window — and dropped if it's too similar.

The subtle part is that this filter is deliberately single-shot. It never examines the second transcript. If it stayed armed, it would swallow a real user genuinely repeating themselves, which is exactly what people do when a translation didn't land the first time.

Every layer here is cheap on its own. The reason there are three is that each one fails differently, and the failure of an echo guard is not a glitch — it's a loop the user has to force-quit.

Two engines, and the scoring trap

Where both recognition engines are available we run both and pick a winner. The obvious implementation — compare confidence scores, highest wins — is wrong, and wrong in a way that looks like it works.

SFSpeechRecognizer's segment-average confidence and SpeechTranscriber's confidence attribute are not measuring the same thing and don't share a distribution. Comparing them raw doesn't pick the better transcript; it picks whichever engine happens to score generously.

So each engine gets a calibration — a floor and a ceiling measured on real hardware — that maps its raw score onto a shared 0…1 scale. On top of that sit three thresholds:

The default calibration in the code is explicitly marked as un-measured and is kept only to document the pre-measurement state. Production uses the values taken off a device. If you build one of these, budget for that measurement session — it is not a constant you can reason your way to.

The thing that actually hurt: downloads

On-device translation needs the language pack on the device, and the speech model too. iOS downloads both, and iOS decides when.

On Wi-Fi that's minutes. On cellular we watched a spinner turn for over half an hour, because iOS holds large system assets back on expensive connections. There is no "download over cellular" switch you can offer. The app cannot fix this.

What it can do is stop lying about it. We added a network-path observer whose entire job is to write an honest sentence on the download screen — you're on cellular, iOS may hold this back, this is why it looks stuck. It also detects Low Data Mode, where iOS stops background downloads outright.

That observer is the only piece of networking code in the app, and it doesn't make a single request. It reads the path and picks better words.

What we'd tell you before you start

  1. Probe the frameworks before you design. Print the language arrays, check the on-device flags, time forty calls. Half a day of measurement reshaped this entire project — and the gap it found in Apple's own coverage was the product opportunity.
  2. Grade the output by hand, in your language, with sentences from the real setting. Automated quality scores would not have caught "bill" becoming "account". That one mistranslation defined the roadmap.
  3. Assume the phone will hear itself, and defend in layers. Put the invariant in the type, not in a flag.
  4. Never compare confidence scores across engines without calibrating them to a shared scale first.
  5. Model download is part of your product, not part of the OS. You can't speed it up. You can stop it feeling broken.
  6. On-device removes a server, not the work. It removes the bill, the proxy, the region question and the DPA. It adds echo control, arbitration, calibration and a first-run download flow. It was still the right trade for us — but it is a trade, not a shortcut.

Postscript

ULAK is on the App Store — free, 24 languages, and it works with the phone in airplane mode once the packs are down. We build software like this for other people too: if you're weighing an on-device approach against a cloud one, we're happy to tell you where the line is for your case.