A coding agent that actually works with the model you already run.
Most agent CLIs bet on native function-calling and fall apart the moment a model gets it slightly wrong. Polyglot parses and repairs tool calls out of raw model text instead - the same reliable loop whether it's Claude, GPT, or Qwen running on your own GPU.
Then polyglot to start. Or build from source.
Running agents for a team on self-hosted models? See Polyglot for teams →
I'll check that file.
<tool_call name="read_file">
{path: 'src/app.ts',}{"path": "src/app.ts"}⏺ read_file ✓ executed
Why this exists
What breaks
Native function-calling assumes the model outputs near-perfect JSON against a provider's exact schema. Open-weight models frequently don't - malformed arguments, missing fields, near-miss tool names - and the whole agent loop stops.
What Polyglot does
Tools are described in the system prompt, and a fault-tolerant streaming parser extracts and repairs the call from whatever text comes back - trailing commas, single quotes, a model that defaults to OpenAI-style JSON instead of the taught format. Same executor, every provider.
Every model we tested, not just the flattering ones
The same six coding tasks, 30 runs per agent per model, on seven models running locally. Polyglot is the only agent that doesn't fail on any of them.
Scroll sideways for all six agents.
| Model | Polyglot | pi | goose | goose + toolshim | Hermes | opencode |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | Works reliably (27 of 30) | Works reliably (28 of 30) | Works reliably (30 of 30) | Works reliably (27 of 30) | Works reliably (28 of 30) | Works reliably (30 of 30) |
| gpt-oss 20B | Works reliably (30 of 30) | Works reliably (29 of 30) | Works sometimes (18 of 30) | Fails (5 of 30) | Works sometimes (18 of 30) | Works sometimes (15 of 30) |
| Devstral Small 2 24B | Works reliably (29 of 30) | Works reliably (28 of 30) | Works sometimes (24 of 30) | Works sometimes (18 of 30) | Works reliably (30 of 30) | Works sometimes (23 of 30) |
| qwen3-coder 30B | Works reliably (30 of 30) | Works reliably (26 of 30) | Works reliably (29 of 30) | Works sometimes (25 of 30) | Works reliably (29 of 30) | Works sometimes (23 of 30) |
| qwen2.5-coder 32B | Works sometimes (24 of 30) | Fails (0 of 30) | Fails (0 of 30) | Works reliably (28 of 30) | Fails (0 of 30) | Fails (0 of 30) |
| qwen2.5-coder 14B | Works reliably (27 of 30) | Fails (0 of 30) | Fails (0 of 30) | Works sometimes (21 of 30) | Fails (0 of 30) | Fails (0 of 30) |
| qwen2.5-coder 7B | Works sometimes (12 of 30) | Fails (0 of 30) | Fails (0 of 30) | Fails (4 of 30) | Fails (0 of 30) | Fails (0 of 30) |
On the newest models every agent does well. On models that write tool calls as text instead of through native function calling, most agents make no tool calls at all and just describe the work. Tiers were set before the runs; within a tier, a run or two either way is noise. qwen2.5-coder 7B sits on the line at 12 of 30. Polyglot was a pre-release build of 0.13.2. Every run, the method and the harness are public on GitHub; read the full write-up.
What you get
Any provider, same behavior
Anthropic, OpenAI, or any OpenAI-compatible endpoint - Ollama, vLLM, LM Studio, llama.cpp server, TGI. The tool-call path is identical underneath all of them.
Permission modes
manual asks before every write or command. auto runs freely with allow/deny rules. plan explores read-only until you approve what it's about to do.
MCP servers
Connect any MCP server over stdio. Its tools go through the same text-parsed grammar as the built-ins - usable by a local model, not just Claude or GPT.
Multi-agent
The task tool delegates to a fresh sub-agent with its own tools and context, up to a bounded depth. Parallel task calls in one turn run concurrently.
Full TUI
Streaming responses, inline tool-call cards, approval prompts, Shift+Tab to cycle permission mode mid-session - not a bare line-by-line prompt.
Source-available
Read it, modify it, self-host it, run it internally, for free, forever. FSL-licensed - converts to Apache 2.0 two years after each release.
Try it against your own model.
No account, no API key required if you're running local.
For teams
A self-hosted layer for teams: qualify open-weight models against your own tools, run coding agents reliably on them, and keep an audit trail of what they did.