How Zeph works
Ships today

Provider

A source of AI capability the system routes work to — cloud or local.

A provider is a source of AI capability the system routes work to: the engine that transcribes your speech, the model that answers, the one that turns text into searchable memory. The companion sends each job to whichever provider you chose for it.

They arrive in three shapes.

The three shapes a provider takes

The companion

Sends each job to the provider you picked.

  • paste a keyCloud service

    OpenAI, Deepgram, Anthropic. The companion verifies the key, then stores it in the Mac keychain.

  • download a modelLocal engine

    Whisper, MLX, Parakeet. The companion fetches the model and runs it on your machine. No key, offline.

  • point at a daemonLocal runtime

    Ollama, LM Studio. A runtime already running on your Mac that the companion discovers.

Different plumbing, one result — a capability the companion can route to.

What ships

Providers
16
Run locally
8
Free
8
Metered
8
ProviderWhat it doesRunsCost
AnthropicanthropicllmClaude powers the assistant and chat, with strong reasoning and reliable tool use. One key, pay as you go.Cloudapi_keyMetered
Apple Intelligenceapple-fmllmUses the language model built into Apple Intelligence, running privately on your Mac. No key, no download.Localapple-fmFree
Mac voiceapple-speechttsReads the assistant's replies aloud with the Mac's own voice. No key, no download, nothing leaves your Mac.Localapple-speechFree
DeepgramdeepgramsttttsStreaming speech-to-text with live partial transcripts, plus Aura text-to-speech on the same key.Cloudapi_keyMetered
GooglegeminillmsttttsembeddingsGoogle's Gemini (and Gemma) models power the assistant and chat, turn speech into text for dictation, read replies aloud, and embed text for memory. One key covers all of it.Cloudapi_keyMetered
GroqgroqllmRuns open models like Llama and Qwen extremely fast, so the assistant answers with very little delay.Cloudapi_keyMetered
LM Studiolm-studiollmembeddingsUses models you've loaded in the LM Studio app, running privately on your own Mac. Nothing leaves your computer.Locallm-studioFree
Llamalocal-llmllmRuns open language models right on your Mac to summarize and enrich your notes, tuned for Apple Silicon. No key, and it works offline.Localllama-cppFree
MLX Whispermlx-whispersttTurns your speech into text right on your Mac, tuned for Apple Silicon. No key, and it works offline.Localmlx-whisperFree
OllamaollamallmembeddingsUses open models running in Ollama on your own Mac, so the assistant and memory work privately with nothing leaving your computer.LocalollamaFree
OpenAIopenaillmsttttsembeddingsPowers the assistant and chat, turns your speech into text for dictation, reads replies aloud in a natural voice, and embeds text for memory. One key covers all of it.Cloudapi_keyMetered
OpenRouteropenrouterllmOne key gives the assistant hundreds of models from many makers (Anthropic, OpenAI, Meta, Mistral, and more), so you can switch without a separate account for each.Cloudapi_keyMetered
ONNXsherpasttRuns open speech-to-text models on your Mac — pick SenseVoice for many languages or Parakeet TDT for high-accuracy European transcription. No key, and it works offline.LocalsherpaFree
Vercel AI Gatewayvercel-ai-gatewayllmembeddingsOne key gives the assistant hundreds of models across makers, with one bill for all of them and an automatic fallback if a model is unavailable.Cloudapi_keyMetered
WhisperwhispersttTurns your speech into text right on your Mac. No key, and it works offline.Localwhisper-cppFree
Zeph Cloudzeph-cloudllmsttUse transcription and AI on your Zeph subscription. No API keys, no per-provider accounts. Sign in and go. Zeph routes to strong models behind the scenes; you pick an effort tier and spend your monthly allowance.CloudMetered

Cost is what the provider’s manifest states. Most cloud vendors state nothing, and an unstated cost is not a free one — those rows say so rather than guess.

You ask for a capability, not a vendor

A provider advertises what it can do — transcribe, answer with a model, embed text, speak. You set one default per job, and the companion routes there.

So a workflow never names OpenAI or Whisper. It asks to transcribe, and your chosen provider serves it. Swap the provider and every workflow that transcribes follows, untouched. That vocabulary — the jobs, and how a default resolves — is Capabilities.

Verifying a key, and finding a daemon

A pasted key is checked before it is kept: the companion makes one probe call to the vendor and stores the key in the login keychain only if it answers. A downloaded model is fetched once and held on disk, so later runs need no network. A local runtime is one you already run on your Mac; the companion looks for it where it usually listens and uses it once it is up.

Next

On this page