For teams: coding agents on the self-hosted models you're allowed to run - see early access →
polyglot

October 1, 2026

Six coding agents, seven local models, 30 runs each

A larger, stricter rerun of our tool-call reliability benchmark - seven self-hosted models, six agents, tiers fixed in advance, and every run public.

Our first benchmark asked a narrow question: does a coding agent still work when the model you run yourself is shaky at tool calls? It used three runs per task, a handful of models, and a harness we kept private.

This is the bigger version. Seven models running locally, including the current ones people actually pick today. Six agents. Thirty runs per agent per model. Tiers decided before the first run. And the whole thing - harness, raw results, failed-run transcripts - is public on GitHub.

The result

Each cell is how many of 30 runs completed the task, checked automatically against the files the agent left behind.

Model Polyglot pi goose goose + toolshim Hermes opencode
Qwen3.8-27B 27 28 30 27 28 30
gpt-oss 20B 30 29 18 5 18 15
Devstral Small 2 24B 29 28 24 18 30 23
qwen3-coder 30B 30 26 29 25 29 23
qwen2.5-coder 32B 24 0 0 28 0 0
qwen2.5-coder 14B 27 0 0 21 0 0
qwen2.5-coder 7B 12 0 0 4 0 0

We set three tiers before running anything: works reliably is 26 or more out of 30, works sometimes is 12 to 25, and fails is under 12. Thirty runs is enough to tell those apart, but not to rank agents inside a tier - 28 out of 30 is consistent with anything from 79% to 98%. So a 30 next to a 28 is a tie, not a win.

Read that way, Polyglot is the only agent that doesn’t fail on any of the seven models. Its weakest cells are qwen2.5-coder 32B and 7B, both “works sometimes” - and the 7B sits right on the line at 12.

The newest models don’t separate the agents

On Qwen3.8-27B every agent works reliably. On qwen3-coder four of the six do, on Devstral three, and the rest still work sometimes. These models are good at emitting tool calls through the native channel that most agents listen on, so the agents mostly differ in how they plan and edit, not in whether they work at all.

If you run one of these models, pick your agent on other grounds. That’s a finding, not a caveat - and it’s the reason we keep testing new models instead of only the ones that flatter us.

Where it breaks

It breaks on models that write their tool calls as plain text instead of through the native channel. All three qwen2.5-coder sizes do this through Ollama. On those, pi, Hermes and opencode made no tool call in any of the 30 runs: the model writes a perfectly reasonable {"name": "edit_file", ...} into its reply, the agent treats the whole reply as a final answer, and the task ends with nothing done. goose makes some calls but completes no task.

goose’s toolshim is the honest exception. It routes the model’s text through a second, small model whose job is to turn it into a proper tool call, and it works: on qwen2.5-coder 32B it scored 28, ahead of Polyglot’s 24. The same second model hurts on gpt-oss, where it dropped goose from 18 to 5. Polyglot does the same job with a parser instead of a second model, which is why it holds up on both.

gpt-oss is the other model that separates the field. It’s strong, but it emits tool calls in its own trained format, and goose, Hermes and opencode land in “works sometimes” on it while Polyglot and pi stay reliable.

Smaller prompts, too

We also measured what each agent sends to the model, from Ollama’s own request log so every agent is counted the same way. Polyglot opens each task with about 1.2k tokens of fixed prompt and pi with 1.2-1.6k. Hermes starts at 3.9-4.7k, and goose and opencode at 5.5-7.6k. Across a whole task, Polyglot used several times fewer prompt tokens per completed task than goose, Hermes and opencode. pi reuses Ollama’s prompt cache better than we do - something for us to fix, not to advertise.

What we got wrong the first time

Our earlier numbers for goose, Hermes and opencode were wrong, and not in their favour. Ollama defaults to a 4,096-token context and silently cuts anything longer. Those three agents send several thousand tokens of instructions, so on our first runs they were working from a prompt with most of their instructions missing.

Every model in this grid now runs with a 32k context, and the harness refuses to start below 16k. After each agent’s runs, it scans Ollama’s log for truncated prompts and marks the results invalid if it finds any. The older runs are still in the repository, labelled as superseded.

Where Polyglot lost runs

We read every one of Polyglot’s failed transcripts. Outside the 7B, it missed 13 of 150 runs. In up to 5 of those, our parser was at fault: a model stuttering an empty tag before the real one, or running past a broken closing tag into an invented next turn. Those are fixed in Polyglot 0.13.2, along with every other format we found in the reruns. The rest were the model’s own mistakes - a wrong edit, a misread file, an answer it made up.

The Polyglot column above is a pre-release build of 0.13.2, the first of three we ran. The two later builds scored 173 and 172 out of 210 against this one’s 179 - the same tiers, with the spread coming from the models, not the code. gpt-oss, for example, scored 30, 27 and 25 across the three builds without a single parsing error among its misses. That’s what run-to-run noise looks like at 30 runs, and it’s why we report tiers.

From here on, Polyglot’s column comes from the current npm release, run once, whatever it scores. We won’t rerun a build to replace a number we don’t like, and every run stays in the repository.

Methodology

  • Tasks. Six small coding jobs: add a CLI subcommand, fix an off-by-one, answer a question from a file without editing it, remove dead code, rename a function across two files, and trace a runtime error to its cause. Each runs in a fresh directory.
  • Models. All local, through Ollama, on one RTX 5080 16GB, each with a 32k context: Qwen3.8-27B, gpt-oss 20B, Devstral Small 2 24B, qwen3-coder 30B, and qwen2.5-coder 7B, 14B and 32B.
  • Agents. Polyglot, pi 0.85.1, goose 1.52.0 (with and without its toolshim, using llama3.2 3B as the interpreter), Hermes Agent 0.19.0 and opencode 1.18.32, each run headless with tools auto-approved and an isolated config.
  • Scoring. The same automated check for every agent: the files must end up right and a verify command must pass. Five trials of each task, so 30 runs per cell, 600 seconds per run.
  • Reproduce it. The README has the commands, and results/GRID.md links every cell to its raw file. If you think we got something wrong, open an issue on the run in question.