ncode Desktop Customize
Providers, models and web search
ncode works with the model providers you add. You need at least one, with your own API key: Anthropic, or any endpoint with an OpenAI-compatible API (OpenAI, OpenRouter, DeepSeek, a local Ollama or LM Studio, and many more).
Providers#
ncode talks to the model endpoints you add. It ships with no key of its own: you bring yours.
A provider has:
- a kind: Anthropic (the Messages API) or OpenAI-compatible (any endpoint that speaks the OpenAI chat API: OpenAI itself, OpenRouter, DeepSeek, or a local server such as Ollama or LM Studio);
- a base URL, for example
https://api.anthropic.comorhttps://api.openai.com/v1; - an API key (a local server may not need one);
- its models, typed in or fetched from the endpoint's model list, and a default model.
Models per role#
Different kinds of work can use different models. Each role has its own default, and a conversation can still pick its own chat and sub-agent model:
| Role | Used for |
|---|---|
| Chat | the assistant's turns |
| Sub-agent | the agents a swarm or the assistant starts |
| Scheduled | tasks that run on a schedule |
| Workflow | agents inside workflow runs |
| Implementer | the agent that carries out a judged plan |
| Judge | the reviewer in a judged (consensus) turn; it falls back to the sub-agent model |
| Research lead, worker and reporter | the three roles of a deep research |
Reasoning effort#
Effort is how hard a model thinks before it answers. The built-in levels are:
| Level | Meaning |
|---|---|
low | fastest |
medium | balanced |
high | deeper reasoning |
xhigh | coding and agentic work |
max | maximum thinking |
A provider, and each of its models, can carry its own list of levels, so the levels you see are the ones your provider offers.
Pricing and budget#
You can enter a price per million input and output tokens (and cache reads and writes) for each model, so every run shows what it cost. A monthly budget is a spend target you watch; nothing is blocked when it runs out.
Adding a provider#
- Open Settings → Providers & models and click Add provider.
- Fill in Name, Kind, Base URL and API key (click Show to check what you pasted).
- Optionally list Models (one per line) and a Default model; you can also fetch them in the next step.
- For an Anthropic provider, Server-side fallback on refusal is on by default: when the API declines a request, it answers with its recommended substitute model instead of an empty reply.
- Click Save provider.
| Service | Kind | Base URL |
|---|---|---|
| Anthropic | Anthropic | https://api.anthropic.com |
| OpenAI | OpenAI-compatible | https://api.openai.com/v1 |
| OpenRouter | OpenAI-compatible | https://openrouter.ai/api/v1 |
| DeepSeek | OpenAI-compatible | https://api.deepseek.com/v1 |
| Ollama on this Mac | OpenAI-compatible | http://localhost:11434/v1 |
| LM Studio on this Mac | OpenAI-compatible | http://localhost:1234/v1 |
For an OpenAI-compatible service, the base URL is the part before /chat/completions in its documentation.
Fetching models#
Click Fetch models on a provider's row, or Fetch all models from endpoints at the top, to ask each endpoint for its model list. It is also the quickest test of a key: if the list arrives, the key works. If it fails, the endpoint's error is shown, with your key removed from it.
Deleting a provider asks first; conversations that used it fall back to the default model.
Default models#
Settings → General → Defaults sets the model each role starts with: Chat model, Sub agent model, Default scheduled model, Default workflow model and Implementer model (consensus), plus a default effort for each. Click Save defaults. The research roles have their own pickers under Deep research.
A conversation can override its chat and sub-agent model from the composer's model chooser.
Effort levels per provider#
Click Efforts on a provider's row to edit which effort levels it offers and how each one reaches its API, for the whole provider or per model. Leave it alone to use the built-in levels for the provider's kind.
Prices, usage and budget#
- Settings → Pricing holds a price per million tokens for each model: input, output, cache read and cache write, plus its context window. With prices set, every run, agent and conversation shows what it cost.
- Usage history in the rail shows spend over the last 7, 30 or 90 days, run by run.
- Settings → Budget → Monthly budget (USD) is a target the usage bar fills against. Nothing stops when it is reached.
Web search#
Web search and page reading#
Agents search the web with web_search and read pages with web_fetch. Search needs at least one search engine with your own key; nothing is preconfigured.
- Engines answer searches: Tavily, Exa, Brave and Serper.
- Readers turn a web page into clean text: Jina and Firecrawl. Without a reader, pages are read directly.
Engines are tried in the order you arrange them. The first enabled engine that returns results answers; an engine that fails or finds nothing hands over to the next one. Each engine has its own key, an optional base URL and an on/off switch.
Reading a page on your own machine or network (localhost, private and link-local addresses) is allowed, but asks for approval unless the project is in Full access.
Setting up a search engine#
- Get an API key from one of the engines (Tavily, Exa, Brave or Serper).
- Open Settings → Deep research. The Search providers block lists every engine and reader.
- On the engine's row, paste the key, switch it on and save.
- Use the up and down arrows to set the order engines are tried in.
A reader (Jina or Firecrawl) is optional: agents read pages directly without one.