# Models and API keys

> Which keys Shelley reads, what each one turns on, and how to add any other model.

Shelley doesn't come with access to a model. It gets models from three
places: API keys in its environment, custom models you add in the UI, and, on
exe.dev VMs, the VM's LLM integration.

## Keys from the environment

Shelley reads four API-key variables:

| Variable | Turns on |
|---|---|
| `ANTHROPIC_API_KEY` | Claude models, via `api.anthropic.com` |
| `OPENAI_API_KEY` | GPT models, via `api.openai.com`, plus transcription for recordings |
| `GEMINI_API_KEY` | Gemini models, via `generativelanguage.googleapis.com` |
| `FIREWORKS_API_KEY` | Open-weight models hosted by Fireworks (GLM, Kimi, DeepSeek), via `api.fireworks.ai` |

That's the complete list. There's no variable for other providers and no
base-URL override; use a [custom model](#custom-models-and-endpoints) for
those. Each key turns on Shelley's built-in list of models for that provider,
which changes from release to release.

```sh
export ANTHROPIC_API_KEY=sk-ant-...
export OPENAI_API_KEY=sk-...
shelley --db ~/.shelley/shelley.db serve
```

Shelley reads these when it starts, so restart it after changing them. They're
also inherited by every command the agent runs; see
[Running it safely](/docs/security).

### Recordings need transcription

The record button next to the message box captures your voice, or your voice
and screen, and Shelley transcribes it on the server with OpenAI's
`gpt-transcribe` and `whisper-1`. That needs a route to both models:
`OPENAI_API_KEY`, an exe.dev LLM integration (or the older `llm_gateway`), or
custom models of an OpenAI type whose model names are exactly `gpt-transcribe`
and `whisper-1` (they'll also show up in the model picker). Without one,
Shelley tells you recording is unavailable instead of recording.

Browsers only allow microphone and screen capture on HTTPS pages or
`localhost`, so recording isn't offered if you reach Shelley over plain HTTP at
some other address.

## See what Shelley found

`shelley models` prints the models the server would offer with your current
environment and flags, without starting it. It takes the same global flags as
`serve`, before the command:

```sh
shelley models
```

You get one row per model: its ID, provider, API type, base URL, where it came
from (for example `$ANTHROPIC_API_KEY`), and a `*` on the default. Transcription
routes, if any, are listed after the models. Custom models aren't included;
they live in the database. `predictable` is always listed, but the web UI only
shows it under `--predictable-only`.

## The default model

The model picker remembers your last choice in the browser. When there isn't
one, or it's no longer available, new conversations get the server's default:
`--default-model` if set, else `default_model` from
[shelley.json](/docs/config), else the first available model.

```sh
shelley --default-model claude-sonnet-5.5 --db ~/.shelley/shelley.db serve
```

Use an ID from `shelley models` or the picker. An ID that isn't available is
ignored without complaint.

## Custom models and endpoints

For any other provider, a self-hosted model, or a compatible proxy, add a
custom model in the UI. Open the model picker and click **Manage…** in its
header (or press Ctrl+K / ⌘K and choose **Add/Remove Models & Keys**), then
**+ Add Model**.

Pick the API format, which pre-fills the endpoint:

| Format | Default endpoint |
|---|---|
| Anthropic | `https://api.anthropic.com/v1/messages` |
| OpenAI (Chat API) | `https://api.openai.com/v1` |
| OpenAI (Responses API) | `https://api.openai.com/v1` |
| Google Gemini | `https://generativelanguage.googleapis.com/v1beta` |

Switch to **Custom endpoint** to point it elsewhere, keeping the same shape:
the full messages URL for Anthropic, the base URL for OpenAI (Shelley appends
`/chat/completions` or `/responses`). For example:

- a local OpenAI-compatible server: `http://localhost:11434/v1`
- a z.ai coding subscription (as opposed to its per-token API), with the OpenAI
  Chat format: `https://api.z.ai/api/coding/paas/v4`

Then fill in:

- **Model**: the model name the endpoint expects.
- **Display Name**: what the picker shows.
- **API Key**: required. If your server doesn't check keys, enter a
  placeholder.
- Optionally, max output tokens, image support, reasoning support and how
  Shelley's reasoning levels map onto the model's, and tags.

**Test** sends a request with those settings before you commit. Saved models
appear in the picker right away, no restart needed. They're stored in Shelley's
database, API key included, in plain text.

## On exe.dev

On an exe.dev VM, Shelley discovers the VM's LLM integration at startup and
offers its models with no keys on the VM. If you also set keys, the
integration's models win where the same model ID comes from both. The refresh
button in the model picker re-runs discovery, and `--disable-llm-integration`
turns it off. Outside exe.dev, there is nothing to discover.

[Shelley on exe.dev](/docs/exe-dev#models) covers fixing a missing model list,
bringing your own key there, and credits; exe.dev's
[LLM integration](https://exe.dev/docs/integrations-llm) docs cover setting up
the integration itself. The older `llm_gateway` setting
lives in [shelley.json](/docs/config); it's skipped when an integration is
found, and `--disable-gateway` ignores it.

## No key at all

```sh
shelley --predictable-only --db /tmp/shelley-try.db serve
```

`--predictable-only` hides every model except `predictable`, a built-in fake
model that needs no key and makes no API calls. Recording is off too. It's
useful for trying the UI, or for testing [hooks](/docs/hooks) and
[skills](/docs/skills) without spending tokens.
