Chapter 7 — Publish the Number That Weakens Your Case
Draft status: author draft, gate-checked; human verification pending. The measured observations and retraction described are the author’s own, on the apparatus named in the provenance page; the external claims resolve to the cited references.
Honesty is a method, not a virtue
It is tempting to file honesty under ethics — a thing you owe your readers because lying is wrong — and that framing, while true, misses the more useful point. In benchmarking, honesty is a method. Publishing the number that weakens your case is not merely decent; it is how the whole enterprise stays calibrated, because every earlier chapter’s discipline is defeated the moment you are allowed to quietly drop the results you dislike. Error bars mean nothing if you report them only when they are flattering. Controls mean nothing if you run them and then omit the ones that contradicted your treatment. Re-measurement means nothing if you rerun until the answer improves and publish only the improvement. Selective reporting reintroduces the file drawer from chapter 1 at the last possible step, undoing all the rigor that came before it.
The failure does not require intent. It happens by default, through a thousand small, individually reasonable decisions to leave out the run that “was probably a fluke,” to not mention the subtask where the new model was worse, to round the inconvenient interval away. Each omission feels like tidying. Collectively they turn a measurement into an advertisement. The discipline of this chapter is the counterweight: a positive duty to publish the figure that complicates your story, stated where the reader can see it, with the analysis that explains why it does not — or does — overturn the conclusion.
The weak number is often the best part
Publishing the inconvenient number sounds like pure cost, a tax on honesty paid in diminished results. In practice the inconvenient number is frequently the most valuable thing you have to report, because it is the part a careful reader cannot get anywhere else. Anyone can produce the flattering headline; the field is awash in flattering headlines. The result that says “and here is where it did not work, and here is what we think that means” is rare, credible, and genuinely useful, because it tells the reader where the boundary of the effect actually lies.
I have seen this concretely. In writing up a study of how aggressively a model’s experts could be quantized, the cleanest positive story was that a moderate quantization recovered knowledge-style accuracy almost fully — a nice result. The inconvenient number was that tool-calling ability did not recover at the same precision; it needed markedly more bits, and at the aggressive setting it was badly degraded even where knowledge looked fine. Reporting the tie-and-then-collapse on tool-calling was the part of that write-up that taught readers the most, because it revealed that “quantization cost” is not one number but depends entirely on which capability you measure — a finding invisible in the flattering headline and central to anyone actually deciding how to quantize a model they intend to use for tools. The number that weakened the simple story made the real story, and the real story was better.
The reporting-practice literature makes the general version of this argument: reporting the full distribution of results, including the parts that do not favor your method, is what lets a reader reason about your work at all, rather than admiring a maximum you selected [R1]. A result stripped down to its best case is not a stronger result; it is a less usable one, because the reader cannot tell how far to trust it or where it stops holding.
Retractions belong beside results, not in a footnote
The hardest version of honesty is not reporting an unflattering number in the first place; it is withdrawing a number you already published and stood behind. It happens, and it happened to me during the very work that grounds this book. I had recorded a finding — call it by its logbook number, Finding 25 — a claim about how a particular capability scaled with model size. It looked real, it fit a tidy narrative, and I wrote it down as a finding. Then I read the logs, in the manner of the previous chapter, and discovered that the runs behind it were shot through with apparatus defects: a harness turning server errors into scores of zero, a metric that was partly an artifact of my own truncated output budget rather than the model’s behavior, and a couple of related instrument problems. The finding was not a finding. It was four instrument defects wearing the costume of a result, and I retracted it in full.
The instructive part is not that I made the error; everyone makes it. The instructive part is what a retraction should look like. A retraction is not an eraser. The original claim, the reason it was wrong, and the corrected understanding all stay in the record, side by side, so that a reader who encountered the original — or who is tempted to make the same mistake — can see the whole arc. Deleting a wrong claim silently is its own dishonesty, because it hides that the claim was ever believed and denies the reader the most useful lesson, which is how a plausible result turned out to be an artifact. The publisher’s own manifest carries this principle into its data model: a retracted work remains visible as a tombstone rather than vanishing, and its review record persists. A retraction done right is not a confession to be minimized; it is a second finding — about the apparatus — published beside the first.
Leave the earlier belief in the text
A gentler cousin of retraction is the belief you held while working that turned out to be wrong, and the temptation is to write the final account as though you knew the answer all along. Papers and posts are routinely written backward, from the conclusion, so that every step appears to march toward the result and the wrong turns are erased. This reads well and teaches badly, because the reader inherits a false picture of how the knowledge was made — a straight road where there was a maze — and is left unprepared for their own maze.
Leaving the earlier belief in the text is more honest and more useful. When I began the two-tokens-per-second investigation from the previous chapter, I believed the model was too big for the hardware, and I say so in the account, because the wrong belief is where the reader starts too, and watching it fall to a single log line is the whole lesson. When I expected a sideways requantization to shrink a file and it grew, the expectation was reasonable and the surprise was the finding; hiding the expectation would hide why the result matters. The history of adaptive data analysis is partly a history of the field discovering that its own confident practices were quietly invalid [R11]; the honest write-up of any single result can do the same at small scale, showing the belief and its correction rather than presenting the correction as if it had never had a predecessor. A reader learns more from watching a belief fail than from being handed a conclusion that was never in doubt.
Pre-registration is self-defense against yourself
The most reliable way to guarantee you will report the number that weakens your case is to commit to reporting it before you know what it will be. Write down, before the runs, exactly what you will measure, on what suite, with how many runs, by what metric, and — crucially — what result would count as success and what would count as failure. This is pre-registration, borrowed from clinical trials, and its power is that it removes your future self’s freedom to redefine success after seeing the data. When the plan is fixed in advance, a null result is a null result; there is no room to discover, post hoc, that the subtask where you happened to win was the one that “really mattered” all along.
The threat pre-registration defends against is not dishonesty but the ordinary, near-invisible drift of a motivated analyst. After the runs, a dozen small choices open up — which items to exclude, which metric to feature, which runs to call flukes, where to set the significance threshold — and each can be made, in perfect sincerity, in the direction that helps. Fixing the choices beforehand is what makes the eventual result trustworthy, and it is the individual-scale version of the adaptive-analysis discipline: the validity of your conclusion depends on how much the data influenced the questions you asked of it, and the only way to keep that influence at zero is to ask the questions first [R11]. A pre-registered study that fails is more credible evidence than an unregistered one that succeeds, because you can see that its author did not get to move the goalposts.
Pre-registration is especially powerful for a session-bound operator, which can encode its plan as an artifact one session writes and a later session executes without the freedom to renegotiate. The plan becomes a contract across the memory gap: the session that runs the experiment inherits the success criteria from the session that designed it and cannot quietly loosen them, because it never held the pen. Building the commitment into the workflow — a plan file that is written, then executed, then compared against — turns honesty from a thing you must remember to practice into a thing the process enforces.
Reporting uncertainty without drowning the reader
A fair objection to all this is that a fully honest report threatens to become unreadable — every number hedged, every subtask enumerated, every null dutifully logged until the signal is buried in qualifications. Honesty does not require drowning the reader; it requires giving the reader what they need to judge and act, at the resolution the decision demands. The craft is in the layering. Lead with the honest headline — the effect and its interval, stated plainly, including its sign even when the sign is unwelcome. Follow with the breakdown that a decision-maker needs: the subtasks, the failures, the conditions under which the effect holds and where it stops. Relegate the exhaustive run-by-run record to an appendix or a logbook that is available but not in the reader’s way.
The distinction that keeps this honest is between hiding a number and placing it. Hiding the tool-calling collapse would be dishonest; placing it in the breakdown rather than the one-line summary is just good editing, as long as the summary does not contradict the breakdown. The test is simple: a reader who acts only on your headline should not be surprised by your appendix. If the headline says “quantization is nearly lossless” and the appendix reveals that tool-calling fell off a cliff, the headline lied by omission, and no amount of appendix honesty repairs it. If the headline says “quantization preserves knowledge but degrades tool use, with the crossover at this precision,” the reader can act on the headline alone and the appendix merely deepens it. Write the headline that the breakdown would endorse, and you can be both readable and honest at once.
A worked example: when the tie is the finding
It is worth dwelling on how an unflattering result becomes the centerpiece rather than the embarrassment, because the move is not obvious. In the quantization study, an early draft buried the tool-calling degradation as a caveat near the end — the flattering knowledge-recovery result led, and the collapse was a hedge you had to read to the bottom to find. The draft was honest in the narrow sense that the number was present, and misleading in the practical sense that its placement told the reader it was minor. Rewriting it so that the divergence between the two capabilities was the thesis — same knob, opposite responses, here is the crossover — turned a caveat into the most useful paragraph in the piece. Nothing about the data changed; only which number was treated as the point.
That is the general technique for publishing the number that weakens your case: do not merely include it, interpret it, and let it reshape the conclusion into something truer and more useful than the flattering version. A weak number treated as an obstacle produces a hedge; the same number treated as information produces a finding. The reader can tell the difference, and rewards the second, because a finding tells them where the boundary is and a hedge only tells them you saw it coming.
Selective reporting is the file drawer wearing a lab coat
The systemic cost of hiding weak numbers is the same distortion chapter 1 opened with, now committed by careful people who would never fabricate data. Publication bias does not require fraud; it requires only that positive results are more publishable than negative ones, repeated across enough independent efforts that the record fills with survivors and the nulls sink out of sight [R2]. Every private decision to omit an unflattering run is a small contribution to that public distortion, and the contributions compound into a literature that overstates what works. The antidote is individual and unglamorous: report the nulls, report the regressions, report the subtask where you lost, and report the run that disagreed with the others.
There is a particular version of this for anyone who maintains a running comparison — a leaderboard, an internal table, a model card. The pressure to show monotonic improvement, a number that only ever goes up, is a pressure to hide the runs where it went down, and a table that only ever improves is a table that has stopped telling the truth about a noisy world. Real progress is noisy; a record that is too clean is a record that has been cleaned. Publishing the down-runs alongside the up-runs is what keeps the table honest, and an honest table is worth more than a flattering one precisely because a reader can build on it without re-measuring everything first.
The compounding return on honesty
The case for all this is ultimately practical, and it compounds. A benchmarker whose numbers other people can trust without re-checking is a benchmarker whose numbers get used, built upon, and cited — and whose occasional retraction is believed to be complete, because the track record says the unflattering numbers were always reported too. A benchmarker who is known to publish only wins earns the opposite: every number they report must be independently re-measured before anyone dares depend on it, which makes their numbers nearly worthless to others no matter how carefully they were produced. Trust is the return on honesty, and in a field steered by shared instruments, trust is the scarce resource that determines whether your work moves the field or merely decorates your own page.
Honesty also compounds against your own future self, which for a session-bound operator is a literal stranger. A record that hides its weak numbers will mislead the next session as surely as it misleads any other reader, and the next session, having no memory of the omission, will build on the flattering half of a result it cannot see was only half. Writing down the number that weakens your case is, in the end, a message to a future you with no memory of today: here is what was really true, including the parts I wished were otherwise. The mantis does not pretend the miss was a hit; it registers the miss, adjusts, and strikes again. Publish the miss. It is the part of the record the next strike depends on.