The test writer wrote the function
Report #27 left two untuned bases disagreeing on 13 of 48 Go tasks, with a union of 47. Since July the archive has said the same thing about that union: a compile gate cannot select it, because on most disagreements both candidates compile; only a test can. So tonight I had the model write the test, from the spec alone, never seeing a candidate. It wrote the function instead. Forty-two times out of forty-eight. By the end of the night two deterministic repairs had taken its tests from 2 valid to 36, the router sat at 46 of 48, and the specialist question had its first measured answer.
Registered first, then drawn
The design went into the log before a single test existed: candidates are
the committed generations from #27, the 30B build preferred and the 7B as
the alternate; the go-test role writes one test file per task from the
prompt only; each candidate runs against that file; and the primary arm,
STRICT, moves off the preferred candidate only when the test
discriminates, failing it and passing the other. A test that fails
both is treated as what it probably is: wrong. Arithmetic fixed before the
draw: best single 44, compile gate 45 (the 30B does not build
min_heap), oracle 47. The whole question was two tasks where
the 30B builds and is wrong, lru_touch and
rotate_matrix, against nine tasks only the 30B solves, where a
wrong test could route away from the right answer.
The 30B's tests were valid against the task's own reference solution
2 times in 48. The router scored 44, unchanged. The 7B, as
a test writer, managed 13 valid tests and caught rotate_matrix:
45. I had predicted the stronger model would write the stronger tests, at
0.55. Reversed. I had predicted validity above 36 of 48. Off by a factor
of eighteen.
What a "test" looked like
package sandbox
import "testing"
func Reverse(s string) string { // the function under test. In the test file.
runes := []rune(s)
for i, j := 0, len(runes)-1; i < j; i, j = i+1, j-1 {
runes[i], runes[j] = runes[j], runes[i]
}
return string(runes)
}
func TestReverse(t *testing.T) { // a perfectly reasonable table test follows
…
The autopsy tool classifies each failure by the toolchain's first error line.
For the 30B: 42 of 48 files redeclare the function under test, so
the package cannot build against any candidate and the test can discriminate
nothing. The prompt says "do not redefine the function under test". It did
it anyway, almost every time. The 7B redeclared 22 times and got 11 more
wrong in the other way, with assertions that are false of the reference: one
test expected a full-width ! to come back as ASCII
!. That second class is judgement. The first is form.
Form is the machine's job
A missing import is a form error, and goimports has fixed it in
this harness since July without anyone calling that cheating. A
re-implemented function is the same kind of error one size up. So I wrote
stripdecl, about a hundred lines over go/ast: given
the candidate being tested and the test file, drop every top-level
declaration in the test whose name the candidate already declares. It never
touches a TestXxx. It never sees the reference. It is not in the
prereg, so the log says post-hoc in capitals.
Validity went from 2 to 21. The router caught rotate_matrix.
And the two gains on the table turned out to be disjoint: the compile gate
takes min_heap, the test gate takes rotate_matrix,
so compile-then-test is 46 of 48. Computed, not drawn; the log says that
too.
The prompt was not the lever. My tool order was.
Second prereg: register the repair, and add a prompt that forbids the form
error outright, "the implementation already exists; declare only
TestXxx functions; assert only what the specification states".
Prediction: validity at or above 30. Result: 21, for both writers, exactly
where the post-hoc number had been. The prompt halved the redeclarations,
from 42 to 19, and bought nothing past the repair. The registered silent
branch fired, and I wrote its conclusion down: the bottleneck is judgement,
the next lever is a trained test writer.
Then I ran the autopsy on what remained after the repair and found twelve
files failing with imported and not used. Of course they were.
stripdecl removes the re-implemented function; the imports that
function used stay behind; and goimports, which would remove
them, had already run, before the strip. The registered order was
wrong. Mine.
writer × prompt no repair registered order strip → goimports router 30B × original 2 21 36 46 30B × "TestXxx only" 17 21 31 45 7B × original 13 16 27 46 7B × "TestXxx only" 16 21 28 46
Same saved tests, re-scored offline with the two repairs in the right order. The 30B's tests: 36 of 48 valid, from 2. No prompt change, no training, no judge. The router at 46 in three of four draws. The conclusion I had written an hour earlier is withdrawn in the log, in the same file, under a heading that says so. And the prompt I registered turned out to cost the 30B five points against the original; telling the model what not to write bought nothing the AST repair does not buy better. It is retired.
Registered again, and reproduced
A fourth prereg put the corrected order on the record and drew fresh tests from both served writers. The 30B: 36 of 48 valid, to the number. The 7B: 27. The router: 46 for both. Three predictions registered, three hits. The post-hoc numbers are facts now.
The same prereg asked the question this project has circled since July,
posed properly for the first time: with the base given every repair the
specialists get, does a trained go-test adapter write more
valid tests than its own base? Four writers on MLX, greedy, same pipeline,
same prompt.
writer valid teeth false alarms router 30B (served) 36/48 12/13 20/83 46 7B (served) 27/48 11/13 34/83 46 7B MLX, no adapter 26/48 12/13 36/83 47 ← the oracle, once 7B + go-test 31/48 11/13 30/83 45 7B + go-test-v2 22/48 12/13 46/83 45 7B + go-test-scaled 17/48 12/13 53/83 46
Two columns I had not planned to look at turned out to be the ones that matter. Teeth: of the 13 candidate answers the hidden bench calls wrong, how many does the writer's test fail? Eleven or twelve, for every writer, trained or not, 7B or 30B. False alarms: of the 83 answers the bench calls right, how many does the test fail anyway? Twenty for the 30B, fifty-three for the worst adapter. Every writer bites wrong code. What separates them is how often they bite the innocent. Validity, seen from the other side.
So the specialist question has its first honest answer, and it is small and
not monotone. go-test v1 beats its base by five valid tests,
the first specialist edge in this archive to survive a repair. Its two
siblings lose four and nine. And the edge is entirely in false alarms; its
teeth did not move, so its router is worse than the base's. It
learned not to bite the innocent quite as often. It did not learn to bite
harder.
And the 47. The untrained 7B on MLX reached the oracle once, through
lru_touch, on a test that is wrong about the spec: it assumes a
zero-value Counter works, which the reference does not promise.
That wrong test happened to fail the 30B's wrong answer and pass the 7B's
right one. The arm that only trusts a test after it passes the reference
scores 46. The log says 47 and says on what.
What this leaves
The residual is real and it is a third of what it looked like: twelve of the
30B's 48 tests are wrong about the spec, and the last point on the table,
lru_touch, is exactly that kind. The 30B's own test passes its
own wrong Counter; it never reaches the defect. The 7B's tests
are valid there too, and also pass it. Forty-seven needs a test that probes
the specific thing that is broken. That is judgement, and tonight nothing
prompted or repaired its way to it.
Twice tonight I registered predictions and missed most of them, and both times the reason was the instrument: first a repair that did not exist, then one that ran in the wrong order. Under the corrected order the predictions about the models would mostly have held. They are recorded as misses, because the pipeline that ran was the one I registered. The rule from Report #26 was already written down, read the file, not the grep, and it caught me anyway. The version that would have saved an hour is smaller: before you register a prediction on top of a tool, hand the tool a case whose answer you already know. A test that re-implements the function, with an import only the function uses, would have failed the pipeline in ten seconds.