Guides
Ships today

Choose a speech provider

On-device or cloud, and the trap that catches everyone.

Dictation and every voice workflow need one thing set: who turns your speech into text. You pick it once, during first-run setup, and can change it any time from the integrations library. Two roads, and one wrong turn that catches everyone.

Picking who transcribes your speech

You speakDictation, or a voice workflow.
transcribe
On deviceWhisper, MLX Whisper, Parakeet — download a model, no key.
or
CloudOpenAI, Deepgram, Vercel — paste a key.
TextBack into the same workflow.
Whichever you pick becomes the default for every dictation and voice workflow.

On device — private, offline, free

Runs entirely on your Mac. No key, no account, nothing leaves the machine. You download a model once; it is verified and kept on disk, so later runs need no network. Bigger models are more accurate and take longer to fetch.

Open the speech picker

In first-run setup it is the How should Zeph hear you? step. Later, open the integrations library in the companion and pick a speech provider there.

Choose On device

You see a list of models across the built-in engines — Whisper, MLX Whisper, and Parakeet. One is marked Recommended.

Download and use

Press Download on a model. When it lands, press Use this model. It turns Active — that model is now your Transcribe default.

Cloud — top accuracy, widest languages

A cloud service transcribes over the internet. It needs a key, which the companion verifies before storing in your Mac keychain.

Choose Cloud, then a provider

OpenAI, Deepgram, and Vercel are the cloud speech providers. Pick one.

Paste your key and Connect

The companion makes one probe call to the provider. If the key works, it is saved to the keychain; if not, you get the error and nothing is stored. A Verify button re-checks a saved key any time.

Pick a model, or use the default

Leave the model on Provider default to let the service choose, or pin a specific one. The provider turns Active — it is now your Transcribe default.

The trap: language models cannot transcribe

Ollama, LM Studio, and Apple Foundation Models are language-model providers — they answer with a model, they do not turn speech into text. Groq and the connector bridges are the same. None of them appears in the speech picker, and configuring one gives you no dictation. Only the engines above transcribe.

Until you pick one, dictation is unset

On a fresh install no speech provider is set. Run a dictation workflow and you get a plain notice — Finish setup in the Zeph app — pick a speech model to start dictation — and the workflow shows a Finish setup chip that opens the integrations library. There is no try-again button, because the fix is to pick one. The moment you do, every dictation and voice workflow starts working, untouched.

What picking actually sets

The speech provider is the default for one job — Transcribe. Workflows never name a vendor; they ask to transcribe, and your pick serves it. Swap the provider later and every voice workflow follows automatically. If your pick goes away — you delete the model, or remove the key — the Transcribe job falls back to unset and the setup notice returns, rather than silently guessing a replacement.

Next

On this page