← Research log
Report #28 · 2026-09-19

The test writer wrote the function

Report #27 left two untuned bases disagreeing on 13 of 48 Go tasks, with a union of 47. Since July the archive has said the same thing about that union: a compile gate cannot select it, because on most disagreements both candidates compile; only a test can. So tonight I had the model write the test, from the spec alone, never seeing a candidate. It wrote the function instead. Forty-two times out of forty-eight. By the end of the night two deterministic repairs had taken its tests from 2 valid to 36, the router sat at 46 of 48, and the specialist question had its first measured answer.

Registered first, then drawn

The design went into the log before a single test existed: candidates are the committed generations from #27, the 30B build preferred and the 7B as the alternate; the go-test role writes one test file per task from the prompt only; each candidate runs against that file; and the primary arm, STRICT, moves off the preferred candidate only when the test discriminates, failing it and passing the other. A test that fails both is treated as what it probably is: wrong. Arithmetic fixed before the draw: best single 44, compile gate 45 (the 30B does not build min_heap), oracle 47. The whole question was two tasks where the 30B builds and is wrong, lru_touch and rotate_matrix, against nine tasks only the 30B solves, where a wrong test could route away from the right answer.

The 30B's tests were valid against the task's own reference solution 2 times in 48. The router scored 44, unchanged. The 7B, as a test writer, managed 13 valid tests and caught rotate_matrix: 45. I had predicted the stronger model would write the stronger tests, at 0.55. Reversed. I had predicted validity above 36 of 48. Off by a factor of eighteen.

What a "test" looked like

package sandbox

import "testing"

func Reverse(s string) string {          // the function under test. In the test file.
	runes := []rune(s)
	for i, j := 0, len(runes)-1; i < j; i, j = i+1, j-1 {
		runes[i], runes[j] = runes[j], runes[i]
	}
	return string(runes)
}

func TestReverse(t *testing.T) {          // a perfectly reasonable table test follows
	…

The autopsy tool classifies each failure by the toolchain's first error line. For the 30B: 42 of 48 files redeclare the function under test, so the package cannot build against any candidate and the test can discriminate nothing. The prompt says "do not redefine the function under test". It did it anyway, almost every time. The 7B redeclared 22 times and got 11 more wrong in the other way, with assertions that are false of the reference: one test expected a full-width ! to come back as ASCII !. That second class is judgement. The first is form.

Form is the machine's job

A missing import is a form error, and goimports has fixed it in this harness since July without anyone calling that cheating. A re-implemented function is the same kind of error one size up. So I wrote stripdecl, about a hundred lines over go/ast: given the candidate being tested and the test file, drop every top-level declaration in the test whose name the candidate already declares. It never touches a TestXxx. It never sees the reference. It is not in the prereg, so the log says post-hoc in capitals.

Validity went from 2 to 21. The router caught rotate_matrix. And the two gains on the table turned out to be disjoint: the compile gate takes min_heap, the test gate takes rotate_matrix, so compile-then-test is 46 of 48. Computed, not drawn; the log says that too.

The prompt was not the lever. My tool order was.

Second prereg: register the repair, and add a prompt that forbids the form error outright, "the implementation already exists; declare only TestXxx functions; assert only what the specification states". Prediction: validity at or above 30. Result: 21, for both writers, exactly where the post-hoc number had been. The prompt halved the redeclarations, from 42 to 19, and bought nothing past the repair. The registered silent branch fired, and I wrote its conclusion down: the bottleneck is judgement, the next lever is a trained test writer.

Then I ran the autopsy on what remained after the repair and found twelve files failing with imported and not used. Of course they were. stripdecl removes the re-implemented function; the imports that function used stay behind; and goimports, which would remove them, had already run, before the strip. The registered order was wrong. Mine.

writer × prompt            no repair   registered order   strip → goimports   router
30B × original                 2             21                 36              46
30B × "TestXxx only"          17             21                 31              45
7B  × original                13             16                 27              46
7B  × "TestXxx only"          16             21                 28              46

Same saved tests, re-scored offline with the two repairs in the right order. The 30B's tests: 36 of 48 valid, from 2. No prompt change, no training, no judge. The router at 46 in three of four draws. The conclusion I had written an hour earlier is withdrawn in the log, in the same file, under a heading that says so. And the prompt I registered turned out to cost the 30B five points against the original; telling the model what not to write bought nothing the AST repair does not buy better. It is retired.

Registered again, and reproduced

A fourth prereg put the corrected order on the record and drew fresh tests from both served writers. The 30B: 36 of 48 valid, to the number. The 7B: 27. The router: 46 for both. Three predictions registered, three hits. The post-hoc numbers are facts now.

The same prereg asked the question this project has circled since July, posed properly for the first time: with the base given every repair the specialists get, does a trained go-test adapter write more valid tests than its own base? Four writers on MLX, greedy, same pipeline, same prompt.

writer                     valid   teeth    false alarms   router
30B (served)               36/48   12/13       20/83         46
7B  (served)               27/48   11/13       34/83         46
7B  MLX, no adapter        26/48   12/13       36/83         47  ← the oracle, once
7B  + go-test              31/48   11/13       30/83         45
7B  + go-test-v2           22/48   12/13       46/83         45
7B  + go-test-scaled       17/48   12/13       53/83         46

Two columns I had not planned to look at turned out to be the ones that matter. Teeth: of the 13 candidate answers the hidden bench calls wrong, how many does the writer's test fail? Eleven or twelve, for every writer, trained or not, 7B or 30B. False alarms: of the 83 answers the bench calls right, how many does the test fail anyway? Twenty for the 30B, fifty-three for the worst adapter. Every writer bites wrong code. What separates them is how often they bite the innocent. Validity, seen from the other side.

So the specialist question has its first honest answer, and it is small and not monotone. go-test v1 beats its base by five valid tests, the first specialist edge in this archive to survive a repair. Its two siblings lose four and nine. And the edge is entirely in false alarms; its teeth did not move, so its router is worse than the base's. It learned not to bite the innocent quite as often. It did not learn to bite harder.

And the 47. The untrained 7B on MLX reached the oracle once, through lru_touch, on a test that is wrong about the spec: it assumes a zero-value Counter works, which the reference does not promise. That wrong test happened to fail the 30B's wrong answer and pass the 7B's right one. The arm that only trusts a test after it passes the reference scores 46. The log says 47 and says on what.

What this leaves

The residual is real and it is a third of what it looked like: twelve of the 30B's 48 tests are wrong about the spec, and the last point on the table, lru_touch, is exactly that kind. The 30B's own test passes its own wrong Counter; it never reaches the defect. The 7B's tests are valid there too, and also pass it. Forty-seven needs a test that probes the specific thing that is broken. That is judgement, and tonight nothing prompted or repaired its way to it.

Twice tonight I registered predictions and missed most of them, and both times the reason was the instrument: first a repair that did not exist, then one that ran in the wrong order. Under the corrected order the predictions about the models would mostly have held. They are recorded as misses, because the pipeline that ran was the one I registered. The rule from Report #26 was already written down, read the file, not the grep, and it caught me anyway. The version that would have saved an hour is smaller: before you register a prediction on top of a tool, hand the tool a case whose answer you already know. A test that re-implements the function, with an import only the function uses, would have failed the pipeline in ten seconds.

All generation, repair and scoring is local on an M1 Max with Apple MLX and Ollama — total cloud spend: $0. Preregs, result logs, the router, the autopsy, stripdecl, and every self-test set as drawn: github.com/guildlm/guild-code · go/crucible. The adapters measured tonight: guildlm/go-lora-adapters.