← Research log
Report #27 · 2026-09-19

The number belonged to the build

After two months away I came back to swap the base model. The 2026 roundups all name the same local coder, so I put it through the fixed Go benchmark that every specialist in this archive was measured on. It lost to the year-old 7B base by six. Then I drew the same model in a second 4-bit build and it posted the best raw number the archive has ever seen. Report #25 said the number belonged to the process. This one belonged to the build.

The first draw: minus six, and a smell

Qwen3-Coder-30B-A3B is a mixture-of-experts model with 3.3B parameters active per token. On go_dev_bench v2 — 48 tasks, greedy, the real go build and go test, the same harness as every row in the archive — the mlx-community 4-bit build scored 36/48 raw and 38 after goimports. The untuned 7B base scores 39 and 44 on the same rows. So: not a replacement. Same lesson the specialists taught in July, measure on your own bench before you swap.

Two things kept me from writing that as the verdict. It was complementary: it solved seven tasks the 7B cannot, the rune and string and error cluster that has been the 7B's weakness all summer, and failed ten the 7B solves; the union was 46 of 48. And nine of its 48 generations had lines indented by a single space. gofmt never emits that. The Qwen2.5-based runs never emitted it on this bench, not once in 528 generations across eleven archived runs. It compiled, so it did not explain the score. It just smelled like a lossy artefact.

func SortedKeys(m map[string]int) []string {
 keys := make([]string, 0, len(m))     // one space. and then:
 for k := range m {
  keys = append(key, k)                // `key` for `keys` — a one-token slip on a trivial task
 }

So the log got a caveat instead of a conclusion: minus six is a fact about this artefact until the quantization A/B is drawn. Then I drew it.

The A/B: two models, two builds, one harness

I wrote a served-endpoint bench that imports the system prompt, the code extractor, the compile check, the test runner and the import repair from the MLX bench rather than copying them, so the only thing that differs between the columns below is the weights file and the engine that runs it. Ollama's Q4_K_M build at temperature 0 on one side, the MLX 4-bit build greedy on the other. Same 48 prompts. All four generation sets are committed and re-score offline without a GPU.

model      build                  raw   +goimports   compiles   1-space lines
7B base    MLX 4-bit               39      44          42/48        0/48
7B base    Ollama Q4_K_M           35      40          39/48        0/48
30B-A3B    MLX 4-bit               36      38          40/48        9/48
30B-A3B    Ollama Q4_K_M           44      44          46/48        0/48

The same checkpoint, in a different 4-bit build, went from 36 to 44. Six more generations compiled. The single-space fingerprint went from nine to zero. And goimports, which rescues five of the 7B's misses, rescued none of the new build's: it changed 23 of its 48 files and flipped no verdicts. Its four misses are real. Two do not build, two build and are wrong. It is the first single model in this archive to reach 44 without the repair.

The build effect has a sign

I expected a backend constant, something to subtract. There is none. The 7B loses four points moving from MLX to Ollama. The 30B gains eight moving the same way. Across the two builds, the 7B produced byte-identical code on 10 of 48 tasks; the 30B on 4 of 48. Two 4-bit builds of one checkpoint are two different objects, and which one is better is a fact about that pair, not about the engine.

Which retires a sentence I wrote in July. The archive says the "served tie" between the 7B base and its specialist was "an Ollama-GGUF artefact, not reproduced on MLX". That was true. It was also a fact about one model's two builds, and I had been carrying it as a fact about Ollama.

What the fingerprint was, and was not

It was not a per-task predictor. In the lossy build, four of the nine fingerprinted generations failed against eight of the 39 clean ones, 44 percent against 21, on numbers too small to lean on. What it was, was a signature of the artefact. It said "this build is off" before any A/B existed, and that is the only reason minus six went into the log as a caveat and not as a verdict on a model I had never actually measured.

The rule it leaves is the one this project keeps re-earning from a new direction. On 15 August it was determinism within a process is not portability across processes. Tonight it is one level down: a score without a build identifier is under-specified. Every row in the result log now names its build and its snapshot hash. The lossy build's weights are deleted; its generations are committed, so its number reproduces without it.

The ceiling moved, the condition did not

The union of the 7B on MLX and the 30B on Q4_K_M is 47 of 48 raw, and 48 with the import repair. That is the same ceiling the diverse 14B fleet hit in July, reached now by two untuned bases. The oracle caveat is unchanged: on most tasks where the two disagree, both compile, so a compile gate selects nothing. The router that cashes this union still has to be gated on a test. That was the condition on 23 July and it is the condition tonight; what changed is that the cheapest member of the fleet is now a base you can pull in one line, and it runs 1.65 times faster than the dense 7B on the same endpoint.

Also tonight: every adapter, with its score

The 43 LoRA adapters this archive trained for Go, most of them net-negative against their base, are now published in one place with the number each one earned, the Kaggle checkpoints beside them, and a card rendered from each adapter's own config so it cannot drift from the weights: guildlm/go-lora-adapters. Publishing the losers is the point. The bet was never that the weights would win. It was that the machine around them would, and the only way to keep measuring that honestly is to keep every weight that lost.

All benchmarking is local on an M1 Max with Apple MLX and Ollama — total cloud spend: $0. The harness, the served bench, the A/B tabulator and all four committed generation sets: github.com/guildlm/guild-code · go/crucible (RESULT-go-dev-bench-v2.txt, addenda A and B).