Getting started
Models and API keys
Which keys Shelley reads, what each one turns on, and how to add any other model.
Shelley doesn’t come with access to a model. It gets models from three places: API keys in its environment, custom models you add in the UI, and, on exe.dev VMs, the VM’s LLM integration.
Keys from the environment
Shelley reads four API-key variables:
| Variable | Turns on |
|---|---|
ANTHROPIC_API_KEY |
Claude models, via api.anthropic.com |
OPENAI_API_KEY |
GPT models, via api.openai.com, plus transcription for recordings |
GEMINI_API_KEY |
Gemini models, via generativelanguage.googleapis.com |
FIREWORKS_API_KEY |
Open-weight models hosted by Fireworks (GLM, Kimi, DeepSeek), via api.fireworks.ai |
That’s the complete list. There’s no variable for other providers and no base-URL override; use a custom model for those. Each key turns on Shelley’s built-in list of models for that provider, which changes from release to release.
export ANTHROPIC_API_KEY=sk-ant-...
export OPENAI_API_KEY=sk-...
shelley --db ~/.shelley/shelley.db serve
Shelley reads these when it starts, so restart it after changing them. They’re also inherited by every command the agent runs; see Running it safely.
Recordings need transcription
The record button next to the message box captures your voice, or your voice
and screen, and Shelley transcribes it on the server with OpenAI’s
gpt-transcribe and whisper-1. That needs a route to both models:
OPENAI_API_KEY, an exe.dev LLM integration (or the older llm_gateway), or
custom models of an OpenAI type whose model names are exactly gpt-transcribe
and whisper-1 (they’ll also show up in the model picker). Without one,
Shelley tells you recording is unavailable instead of recording.
Browsers only allow microphone and screen capture on HTTPS pages or
localhost, so recording isn’t offered if you reach Shelley over plain HTTP at
some other address.
See what Shelley found
shelley models prints the models the server would offer with your current
environment and flags, without starting it. It takes the same global flags as
serve, before the command:
shelley models
You get one row per model: its ID, provider, API type, base URL, where it came
from (for example $ANTHROPIC_API_KEY), and a * on the default. Transcription
routes, if any, are listed after the models. Custom models aren’t included;
they live in the database. predictable is always listed, but the web UI only
shows it under --predictable-only.
The default model
The model picker remembers your last choice in the browser. When there isn’t
one, or it’s no longer available, new conversations get the server’s default:
--default-model if set, else default_model from
shelley.json, else the first available model.
shelley --default-model claude-sonnet-5.5 --db ~/.shelley/shelley.db serve
Use an ID from shelley models or the picker. An ID that isn’t available is
ignored without complaint.
Custom models and endpoints
For any other provider, a self-hosted model, or a compatible proxy, add a custom model in the UI. Open the model picker and click Manage… in its header (or press Ctrl+K / ⌘K and choose Add/Remove Models & Keys), then + Add Model.
Pick the API format, which pre-fills the endpoint:
| Format | Default endpoint |
|---|---|
| Anthropic | https://api.anthropic.com/v1/messages |
| OpenAI (Chat API) | https://api.openai.com/v1 |
| OpenAI (Responses API) | https://api.openai.com/v1 |
| Google Gemini | https://generativelanguage.googleapis.com/v1beta |
Switch to Custom endpoint to point it elsewhere, keeping the same shape:
the full messages URL for Anthropic, the base URL for OpenAI (Shelley appends
/chat/completions or /responses). For example:
- a local OpenAI-compatible server:
http://localhost:11434/v1 - a z.ai coding subscription (as opposed to its per-token API), with the OpenAI
Chat format:
https://api.z.ai/api/coding/paas/v4
Then fill in:
- Model: the model name the endpoint expects.
- Display Name: what the picker shows.
- API Key: required. If your server doesn’t check keys, enter a placeholder.
- Optionally, max output tokens, image support, reasoning support and how Shelley’s reasoning levels map onto the model’s, and tags.
Test sends a request with those settings before you commit. Saved models appear in the picker right away, no restart needed. They’re stored in Shelley’s database, API key included, in plain text.
On exe.dev
On an exe.dev VM, Shelley discovers the VM’s LLM integration at startup and
offers its models with no keys on the VM. If you also set keys, the
integration’s models win where the same model ID comes from both. The refresh
button in the model picker re-runs discovery, and --disable-llm-integration
turns it off. Outside exe.dev, there is nothing to discover.
Shelley on exe.dev covers fixing a missing model list,
bringing your own key there, and credits; exe.dev’s
LLM integration docs cover setting up
the integration itself. The older llm_gateway setting
lives in shelley.json; it’s skipped when an integration is
found, and --disable-gateway ignores it.
No key at all
shelley --predictable-only --db /tmp/shelley-try.db serve
--predictable-only hides every model except predictable, a built-in fake
model that needs no key and makes no API calls. Recording is off too. It’s
useful for trying the UI, or for testing hooks and
skills without spending tokens.