8. Models beyond free Gemini¶
Free tiers prove a hello. Paid or local models are what you run on.
8.0 Why¶
A free Gemini key is a demo budget. It carries the hellos in chapter 1 and a handful of skill calls, and then the hourly page rotation from chapter 3 or chapter 7 starts eating the daily cap. The failure does not look like a quota: Telegram goes quiet, a cron stops rewriting the page, and a chat turn hangs, so the install looks broken when it is not.
Common fixes make it worse. Re-running full onboarding to add one provider rebuilds Gateway settings you spent two chapters getting right. Pinning a model in the Control UI picker is strict: a pinned model never fails over. Guessing a model id from a provider’s marketing page gets “model not found”. A 20B-class local model as the backup turns a rate limit into what feels like a hang. What works: a key in the env file the daemon reads, an id the catalog printed, thinking off, and a short fallback list of comparable models.
Cloud keys and a small same-LAN Ollama failover belong in this chapter. Publishing a large local or remote LLM at its own hostname, then pointing n8n and OpenClaw at that URL, is LLM on Edgible. That hostname carries an api-key secret, never None.
Where you run this: provider consoles in the host browser; keys, models commands and the hello on the Ubuntu guest; the Mac-side Ollama in 8.5.2 on the host; the fallback notice in a Telegram DM.
8.1 The job¶
1. OpenClaw on the VM onboarded free Gemini Flash. That is enough for the series. Free-tier 429s, Think: medium hangs, and a huge local failover are why Telegram then feels broken. Here you add another provider without reinstalling the Gateway: DeepSeek, OpenAI, Groq, Claude, or a small Ollama on the same LAN. Optional: a fallback list so a 429 still answers. A published model URL is LLM on Edgible.
A Cursor subscription is not a chat model. That is 7. Cursor Agent (ACP). Do not paste a Cursor key into openclaw models set.
Done when
openclaw models list --provider <that-provider>prints the id you will use.openclaw agent --agent main --thinking off --model <provider/id> --message "Say hello in one sentence."replies on the VM (identity ritual counts).openclaw models statusshows the new primary; cooldown empty or understood.- Control UI picker is Default.
/statusin Telegram matches. /think off(oropenclaw config set agents.defaults.thinkingDefault off).- If you set fallbacks:
openclaw config get agents.defaults.modellists them; you did not use a 20B/27B as the snappy backup. - Hello World and openclaw-ui still load. You did not publish
11434or18789on the router.
Need first: 1. OpenClaw on the VM (loopback Gateway) (Gateway up, Gemini hello already worked). The rest of the series can stay on Flash until you do this.
Not this chapter: installing OpenClaw, publishing Control UI, pairing Telegram, hiring Cursor, or publishing an Ollama / vLLM Edgible app (LLM on Edgible).
8.2 Rules that stay¶
Do this on the VM (Gateway host). Put keys in ~/.openclaw/.env so systemd sees them. Do not re-run full openclaw onboard unless you want to redo Gateway setup.
| Rule | Why |
|---|---|
| Control UI / Telegram picker Default | A pinned /model is strict, with no failover. |
/think off (or thinkingDefault off) |
Flash and DeepSeek V4 both hang if thinking stays on medium. |
models set = primary; fallbacks = backup |
Fallback is turn-local. Next message starts on primary again. |
Use an id list actually prints |
Guessing deepseek/deepseek-v4-flash when the catalog says something else is “model not found”. |
Do not publish Ollama or the Gateway on Edgible None |
Control UI stays org. A published inference URL is LLM on Edgible, never None. |
Check cooldown before blaming the new key:
openclaw models status
openclaw status --usage
8.3 DeepSeek V4 Flash (cheap paid default)¶
Best first paid try for this tutorial load (hellos, /skill, small cron): not Gemini Flash, not Grok Fast, not DeepSeek R1 / “thinking” SKUs.
- Create a key at platform.deepseek.com/api_keys and put credit on the account.
- On the VM, add it to
~/.openclaw/.env:
# one line, no quotes
echo 'DEEPSEEK_API_KEY=sk-YOURKEY' >> ~/.openclaw/.env
openclaw gateway restart
- List, pin Flash (onboarding wizards often default to Pro):
openclaw models list --provider deepseek
openclaw models set deepseek/deepseek-v4-flash
openclaw gateway restart
Smoke test. Use the Flash id list printed. Then:
openclaw agent --agent main --thinking off \
--model deepseek/deepseek-v4-flash \
--message "Say hello in one sentence."
Identity ritual still counts. If that fails, prove the key without OpenClaw:
set -a && source "$HOME/.openclaw/.env" && set +a
curl -sS https://api.deepseek.com/chat/completions \
-H "Authorization: Bearer $DEEPSEEK_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"Reply with exactly: ds-ok"}],"max_tokens":16}'
JSON with a short completion means the guest can reach DeepSeek. If curl works and openclaw agent does not, the Gateway is not loading .env. Restart again.
Optional wizard (only if you never added the provider): openclaw onboard --auth-choice deepseek-api-key. Skip it if the Gateway is already how you like it.
8.4 Other cloud keys¶
Same pattern: env var in ~/.openclaw/.env, gateway restart, models list --provider …, models set, then the --model hello from 8.3.
| You want | Typical env | List / set (confirm with list) |
|---|---|---|
| DeepSeek V4 Flash | DEEPSEEK_API_KEY |
deepseek/deepseek-v4-flash |
| Groq (snappy, cheap) | GROQ_API_KEY |
groq/llama-3.3-70b-versatile |
| OpenAI Platform | OPENAI_API_KEY |
openai/… from list, not ChatGPT Free |
| Claude Sonnet | ANTHROPIC_API_KEY |
anthropic/claude-sonnet-… from list |
| OpenRouter (many models, one bill) | OPENROUTER_API_KEY |
openrouter/… |
| xAI Grok | see xAI | not Groq |
Exact --auth-choice flags: CLI automation and model providers. Set a spend cap on the provider console.
Claude vs Cursor: a Cursor product sub does not fill ANTHROPIC_API_KEY. Sonnet as Telegram’s model needs an Anthropic key (or Claude Code on the box, which is not this chapter).
Cost for these tutorials (fat OpenClaw prompt, short replies): DeepSeek Flash is cents. Grok mid-tier is dollars. Sonnet is tens of times DeepSeek on output, and hourly On this day cron is the only thing that can add up. Keep cron on Flash or DeepSeek Flash.
8.5 Local Ollama¶
Prompts stay on hardware you own. If OpenClaw and Ollama share a LAN (Gateway in the Mac’s UTM guest), point at the LAN URL (8.5.2). Do not hairpin through Edgible. If the Gateway is on a different home VM (the layout in LLM on Edgible), skip 8.5.2–8.5.3 and use 4. OpenClaw uses that URL.
8.5.1 Same machine as OpenClaw (enough RAM)¶
| RAM on the VM / mini-PC | What to expect |
|---|---|
| 4 GB (this guide’s VM default) | Too small. Stay on a cloud key. |
| 8 GB | Floor. Tiny model only (~1B). Weak at tools. |
| 16 GB+ | Usable local chat. |
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.2:1b
ollama run llama3.2:1b "Say hello in one word"
Onboard-style (only if you are not already on Gemini): --auth-choice ollama --custom-model-id llama3.2:1b. Otherwise register the provider as in 8.5.3.
8.5.2 Ollama on the Mac, OpenClaw in the VM (32 GB Mac)¶
Do not put the weights in the 8 GB guest. Install Ollama for Mac. After the VM’s 8 GB, you have on the order of 16 GB left for a model.
On the Mac:
ollama pull qwen2.5:7b
ollama run qwen2.5:7b "Say hello in one word"
A fallback should be a 7B-class chat model (qwen2.5:7b or llama3.1:8b). gpt-oss:20b (13 GB) and qwen3.5:27b (17 GB) are slow failovers next to the guest, and they make a Gemini 429 feel like a hang. Use a 20B only as an explicit /model when you want a slow local turn.
Ollama defaults to Mac localhost only. The VM’s 127.0.0.1 is the guest. Listen on the LAN (home only; do not port-forward 11434 on the router):
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"
killall Ollama; open -a Ollama
On the VM:
HOST=$(ip route | awk '/default/ {print $3; exit}')
curl -sS "http://${HOST}:11434/api/tags"
You want JSON with your tag, not connection refused. UTM NAT is often 192.168.64.1 if $HOST is wrong. Do not curl 127.0.0.1 on the VM.
8.5.3 Point the Gateway at Mac Ollama¶
There is no real Ollama key; ollama-local is a dummy. An explicit models.providers.ollama block turns off auto-discovery. Register the tag from /api/tags:
HOST=$(ip route | awk '/default/ {print $3; exit}')
openclaw config set models.providers.ollama.baseUrl "http://${HOST}:11434"
openclaw config set models.providers.ollama.api ollama
openclaw config set models.providers.ollama.apiKey ollama-local
openclaw config set models.providers.ollama.models \
'[{"id":"qwen2.5:7b","name":"qwen2.5:7b"}]' --strict-json
openclaw gateway restart
openclaw models list --provider ollama
list must print ollama/qwen2.5:7b (or your tag). Then either openclaw models set ollama/qwen2.5:7b (local primary) or leave DeepSeek/Gemini as primary and put Ollama in fallbacks (8.7).
Ollama often hides models that /api/show does not mark as tool-capable with ≥16K context. Fallback can still use the config id. To pin local in chat: /model ollama/qwen2.5:7b.
8.6 A published model (another home VM)¶
If OpenClaw is not next to the Mac, do not use $HOST:11434. Register https://ollama.<org>.edgible.com with the api-key secret: 4. OpenClaw uses that URL. Do not set that app to None.
8.7 Fallback chain¶
Example: DeepSeek Flash primary, Gemini Flash then local 7B when DeepSeek is down:
openclaw models list --provider deepseek
openclaw models list --provider google
openclaw models list --provider ollama
openclaw models set deepseek/deepseek-v4-flash
openclaw config set agents.defaults.model.fallbacks \
'["google/<the-flash-id-from-list>","ollama/qwen2.5:7b"]' --strict-json
openclaw gateway restart
openclaw config get agents.defaults.model
You want primary = DeepSeek Flash and fallbacks listing ids that list printed.
Failover applies to the configured default and to cron (chapter 3). It does not apply if you pick a model in the Control UI or /model. Leave the picker on Default.
In a Telegram DM, a switch can show once per state change:
↪️ Model Fallback: ollama/qwen2.5:7b (selected deepseek/deepseek-v4-flash; rate_limit)
Groups suppress that notice; /status still has Fallback. The notice is not a live ticker at the instant of the 429. Expect typing until the fallback has tokens.
8.8 Verify¶
-
openclaw models list --provider <that-provider>prints the id you will use. -
openclaw agent --agent main --thinking off --model <provider/id> --message "Say hello in one sentence."replies on the VM (identity ritual counts). -
openclaw models statusshows the new primary; cooldown empty or understood. - Control UI picker is Default.
/statusin Telegram matches. -
/think off(oropenclaw config set agents.defaults.thinkingDefault off). - If you set fallbacks:
openclaw config get agents.defaults.modellists them; you did not use a 20B/27B as the snappy backup. - Hello World and openclaw-ui still load. You did not publish
11434or18789on the router.
Next¶
That’s the series for models. 9. Tear down OpenClaw when you want the agent and openclaw-ui gone. Index. Cursor ACP (not a chat key) is 7. Cursor Agent from OpenClaw on the Edgible site. A published LLM is LLM on Edgible.