Ollama defaults to 4,096 tokens of context — this is too low for BrowserOS. Below 15K tokens, the context overflows and the agent gets stuck in a loop constantly trying to recover. Only Chat Mode will work at low context lengths. Set at least 15,000–20,000 tokens for local models to function properly.
Set context length when starting Ollama:
OLLAMA_CONTEXT_LENGTH=20000 ollama serve
Increasing context length uses more VRAM. Run ollama ps to check your current allocation. See the Ollama context length docs for more details.