Logs Aren't Enough argued that data-pipeline incidents begin with an evidence scavenger hunt, and that structured operational evidence โ€” an Evidence Packet โ€” should become a first-class artifact. Designing Evidence Packets then generalized that artifact into a source-neutral architecture. Both articles made an engineering case: a well-designed container for operational evidence. Neither tested whether that container actually helps a diagnostic process produce better answers.

This article covers that test โ€” and reports a result that complicates the story more than it completes it.

The hypothesis being tested

The Evidence Packets project's original proposition was broad and compound: a durable, provenance-bearing evidence format is generally valuable for diagnosing data-pipeline incidents. That claim bundles several distinct mechanisms an evidence format could plausibly provide value through โ€” how the evidence looks to whatever reads it, whether it survives to be read at all, how much of a diagnosis it can ground, and how efficiently it can be supplied. A study that tests all of those at once cannot attribute a result to any one of them.

Phase I narrowed this to two sequential, independently testable questions:

  1. RQ1 โ€” Representation. Holding the underlying captured evidence constant, does transforming it into a bounded, explicitly correlated, provenance-bearing packet representation improve an independent diagnostic subject's performance over a reasonable conventional aggregation of the same facts?
  2. RQ2 โ€” Preservation, motivated directly by RQ1's result. Does capturing evidence durably at execution time, versus reconstructing it later from the same operational sources after realistic mutation, preserve diagnostic capability?

These are deliberately narrow, deliberately sequential questions โ€” RQ2 was not a re-run of RQ1 with different knobs. It is a materially different axis, prospectively specified only after RQ1's result was already frozen.

What was actually built to test it

Both experiments share a governed apparatus, not an ad hoc comparison. A diagnostic scenario corpus, a grading rubric, a materiality threshold, and a guardrail formula were fixed and hash-attested before either primary run โ€” the decision rule that would call a result material or not existed before any data could influence it. Each matched pair in both studies is a genuine, independent diagnostic execution against real operational evidence from a working dbt/Airflow/PostgreSQL reference stack โ€” not a fixture replay and not a simulated response.

RQ1 held the underlying evidence content equivalent between its two arms โ€” verified losslessly equivalent by a dedicated automated checker before freeze โ€” so that only representation format varied: a conventional, source-native aggregation versus a normalized, explicitly cross-referenced Evidence Packet. RQ2 reused RQ1's own captured executions and lossless-equivalence machinery, then introduced its manipulation strictly between capture time and diagnosis time: the durable arm is the packet exactly as captured at execution; the reconstruction arm rebuilds evidence from the same sources after a predeclared, ground-truth-blind degradation (the next real pipeline invocation overwriting the prior run's dbt manifest and results โ€” an ordinary operational event, not a fabricated deletion).

Primary outcome in both studies: paired correctness, credited-arm-minus-baseline, material only if the absolute difference reached a prespecified ยฑ0.10 and neither arm's unsupported-assertion or abstention rate worsened by more than 0.10. No significance test was predeclared as a decision gate in either study โ€” a choice that matters for how the RQ2 result below should be read.

RQ1: representation alone did not move the answer

Across 60 matched pairs (12 diagnostic scenario classes ร— 5 repetitions, 120 total genuine diagnostic executions), conventional-format evidence produced correct diagnoses in 13 of 60 cases (0.217); packet-format evidence produced 12 of 60 (0.200). The observed difference โ€” packet minus conventional โ€” was โˆ’0.017, roughly one-sixth of the study's own ยฑ0.10 materiality threshold, and the smallest possible nonzero magnitude a 60-pair design could even produce. The frozen verdict: RC1_PRIMARY_COMPLETE_NO_DEMONSTRATED_ADVANTAGE.

This is worth sitting with, because it runs against the instinct a "well-designed evidence container" article series invites: that giving a diagnostic process the same facts, better organized, should help it reason better. It did not, for this instrument, on this scenario set. And it was not free โ€” packet representation cost roughly 2.0ร— the input tokens of the conventional format ($4.57 vs. $2.32 for the packet-format half of the run) for a measured correctness difference indistinguishable from noise.

One pattern is worth naming without over-crediting it: packet-format responses were far better grounded โ€” cited evidence support rose from 0.047 to 0.569, unsupported assertions per response fell from 1.867 to 1.150 โ€” without that improved grounding translating into more correct final diagnoses. This was not a predeclared outcome for RQ1, so it cannot override the null verdict; it is retained as a hypothesis this study generated, not one it tested. A diagnosis can be well-cited and still wrong, or poorly-cited and still right โ€” RQ1's data makes that distinction concrete rather than hypothetical.

The honest scope of this null result also needs to be stated precisely, in both directions. It is not a proof that representation never matters under any condition โ€” RQ1 tested one representational axis, with one instrument, holding evidence equivalent. It is also not a weak or ambiguous result within that scope: absolute accuracy was low in both arms (7 of 12 scenario classes sat at or near a 0-of-5 floor in both conditions), which limits how much any comparison between them can show, but the โˆ’0.017 figure itself is an unambiguous null by the study's own prespecified rule, not a coin flip that happened to land near zero.

Full frozen result: PRIMARY_RESULT.json; closure record: RC1_CONCLUSION.md.

RQ2: a thinner, more interesting result

RQ1's own closure record asked a natural follow-on question directly: what happens when the evidence available at diagnosis time is not held equivalent to what existed at execution time โ€” when it has mutated, expired, or partially degraded by the time anyone looks at it? That is a question about preservation, not representation, and it is the one RQ2 was built to test.

Across 25 disjoint reconstruction events (50 genuine diagnostic executions), durably captured evidence produced correct diagnoses in 8 of 25 cases (0.320); evidence reconstructed after realistic degradation produced 5 of 25 (0.200). The observed difference โ€” durable minus reconstruction โ€” was +0.120, clearing the same ยฑ0.10 threshold with both guardrails passing. The frozen verdict: RQ2_PRIMARY_COMPLETE_SUPPORTS_DURABLE_CAPTURE_ADVANTAGE.

Read no further than the verdict string and this looks like a clean confirmation: preservation matters where representation didn't. The actual result is real but considerably thinner, and three qualifications are load-bearing enough that none of them can be left for a footnote:

  • The margin is the smallest one the design could ever call material. At 25 pairs, correctness differences are quantized in steps of 1/25 = 0.04. The smallest possible nonzero value that could ever clear a +0.10 threshold is +0.12 โ€” a net margin of exactly three pairs (6 durable-only-correct versus 3 reconstruction-only-correct, out of 9 discordant pairs; 16 pairs tied). RQ2's result landed exactly there. One pair flipping in either direction reverses the frozen verdict entirely.
  • The observed split is compatible with chance. An exact descriptive two-sided test on those 9 discordant pairs gives p = 0.507812. This test was explicitly non-gating in both RQ1's and RQ2's design โ€” the frozen decision rule is the threshold, not a significance test โ€” but it is reported here because a reader should know that "cleared the prespecified threshold" and "statistically distinguishable from a coin flip" are not the same claim, and only the first one is true here.
  • Representation was not held constant between RQ2's two arms. The durable arm is diagnosed in packet format; the reconstruction arm, in conventional format โ€” by design, on the reasoning that RQ1 had already shown format alone doesn't move correctness when evidence is equivalent. That reasoning is defensible but not airtight: RQ1's null was established under equivalent evidence, and it remains untested whether packet format specifically helps more when evidence is incomplete than when it's complete. Some non-zero share of RQ2's +0.120 could in principle be a representation effect that only appears under degradation. This is a disclosed gap, not a demonstrated confound โ€” but it means RQ2 is best read as a preservation-centered comparison, not a cleanly isolated preservation effect.

One secondary finding โ€” predeclared as such in RQ2's design, unlike RQ1's analogous grounding pattern above, which was purely exploratory โ€” is less fragile than the primary correctness number, and arguably more interesting: durable capture produced markedly fewer unsupported assertions than reconstruction (0.720 vs. 1.320 per response) while consuming essentially identical input-token volume in both arms (8,204,048 vs. 8,203,837 total input tokens โ€” a 0.003% difference). That rules out the crude explanation that durable capture simply won by being handed more evidence to work with. It does not prove the two arms carried equivalent information density per token โ€” a reconstructed manifest describing a later, different run could be systematically less informative without being shorter, which is exactly the mechanism this study probes but cannot directly measure from token counts alone.

A further limitation, addressed directly rather than left implicit: 24 of the 25 reconstruction gaps were under three minutes (median 110.7 seconds); the single outlier was approximately 4.72 days. RQ2 is evidence about immediate, short-term reconstruction loss following ordinary subsequent pipeline activity โ€” an overwritten manifest from the very next run โ€” not evidence about long-term archival durability across months or years. No claim about archival value should be read into this result.

Full frozen result: RQ2_PRIMARY_RESULT.json.

What the two results say together

Stated as precisely as the data supports: representation alone did not move diagnostic correctness under equivalent evidence. Whether the underlying evidence had survived intact to diagnosis time showed a material observed advantage for durable capture, together with fewer unsupported assertions, under the tested conditions โ€” directionally consistent with a preservation-centered hypothesis, not a fully isolated preservation effect, resting on a thin margin that the exact descriptive test cannot distinguish from chance, demonstrated only over short reconstruction intervals.

That is a narrower, more heavily qualified statement than "durable capture beats reconstruction," and it is the correct one. The most defensible summary of what survived Phase I is that representation is not, on this evidence, where an Evidence Packet's value lives; whether evidence still exists intact at diagnosis time is a more promising candidate axis โ€” an observed advantage that is directionally consistent, material by the prespecified rule, and simultaneously thin enough that a single flipped pair would erase it.

Phase I does not establish that Evidence Packets generally improve diagnosis. It does not establish long-term archival durability value. It does not establish a causal mechanism for the RQ2 difference โ€” the study demonstrates a correlational, paired-design gap between capture-time and reconstruction-time evidence, not which specific missing or altered fact drove any individual diagnosis. What it does establish is a specific, independently reviewed pair of results โ€” one clean null, one thin positive โ€” that shift the leading hypothesis about an Evidence Packet's value from representation toward preservation, without proving that shift and without fully separating it from a residual representation effect.

Engineering success is not the same claim as demonstrated utility

It is worth being explicit about what this research program did demonstrate, separately from what it did not. The apparatus itself โ€” frozen scenario corpus, hash-attested thresholds and guardrails fixed before data existed, lossless evidence-equivalence verification, deterministic packet assembly, independent review of the finalization logic before any result was calculated โ€” worked exactly as designed, on both studies. That is a real, demonstrated engineering result: a governed, reproducible way to run this kind of comparison at all. It is not the same claim as "the thing being measured showed a strong effect." A well-built instrument that returns a clean null is not a failed instrument; it is an instrument that did its job. This article's title names the gap between those two things directly: better-engineered evidence, rigorously and reproducibly delivered, did not by itself produce better diagnostic answers in RQ1 โ€” and where it plausibly did help (RQ2), the reason was not representation, and the margin by which it showed up was thin.

What this research does not support

To be explicit rather than let a reader infer more than the data grants, Phase I does not establish: general superiority of Evidence Packets over conventional evidence handling; statistical significance of the RQ2 effect; generalization beyond the tested scenario classes, reference stack, or the single substituted diagnostic instrument (gemini-3.1-flash-lite, substituted before execution for cost reasons and fully requalified, not evaluated for representativeness); long-term archival durability; production incident-resolution improvement (both studies are single-shot, stateless, simulated exercises, not a production evaluation); human-investigator benefit (no human-subject arm exists in either study); cost savings (representation cost roughly double the tokens for no measured gain, and RQ2's cost fields are unavailable by design); or a causal mechanism for either result.

Full accounting, including several additional not-established items (latency, context-size reduction, and schema-specific claims among them) and the methodological review this synthesis underwent before being written up, is preserved in the frozen synthesis linked below.

Why this research is pausing here

A null result and a thin, qualified positive result are not the same as a dead end, and they are not being treated as one. Phase I answered exactly the two narrow questions it set out to test, with a governed apparatus that stayed disciplined under its own scrutiny โ€” RQ1's null was preserved and used to motivate RQ2 rather than re-run with a different model to make it disappear, and RQ2's positive result is reported with its thin margin foregrounded rather than buried behind a passing verdict string. That is what a closed research phase looks like when it is closed honestly: specific, bounded claims, clearly separated from what remains untested.

What remains untested is itself informative. Neither RQ1 nor RQ2 gave a diagnostic process access to anything beyond the single event under investigation โ€” no comparison against a source's own history of prior successful or prior anomalous executions. That is a structurally distinct question from both representation and preservation, and Phase I's architecture does not foreclose it, but nothing here tests, designs, or establishes it either. It is recorded as an open question, not an authorized next study.

Full synthesis, including the intellectual-adversarial review this document underwent before being frozen: PHASE_I_SYNTHESIS.md. Closure record: PHASE_I_CLOSURE.md.

Why this matters beyond this implementation

Organizations introducing LLM-assisted reasoning into operational systems tend to invest first in how evidence is packaged for the model โ€” schemas, structured context, retrieval formatting. Phase I's result is a caution against assuming that investment pays off in correctness by default: under evidence-equivalent conditions, better packaging measurably improved how well-grounded a diagnosis looked without measurably improving whether it was right. Where this research found a more promising signal was upstream of packaging entirely โ€” in whether the evidence being reasoned over had already degraded, silently, before the question was ever asked. That reframes a common instinct: the highest-leverage question for an operational AI system may not be "how should we format the context," but "can we be sure the context we're handing it hasn't already quietly rotted." Neither finding is a strong claim yet. Both are the kind of finding worth designing the next experiment around rather than assuming.

What comes next

Evidence Packets Phase I is closed and will remain an independently documented research foundation โ€” RQ1 and RQ2's frozen results, apparatus, and this synthesis are not being revised or extended by what follows. The next research direction, provisionally titled Deterministic Boundaries for Probabilistic Systems, asks a broader question: under what conditions can probabilistic systems provide measurable utility beyond deterministic approaches, and what architectural controls โ€” explicit policy, provenance, controlled retrieval, verifiable information boundaries โ€” are necessary to preserve the integrity, confidentiality, traceability, and governance of that operation?

A first bounded experiment under that program is being scoped: whether semantic or hybrid retrieval improves discovery of relevant DICOM metadata compared with deterministic and lexical retrieval, while preserving enough structural context and provenance for independent verification โ€” building on the separate DICOM Trust Boundary research. See the research overview for how this direction relates to both prior programs. These are current research interests and a proposed direction, not findings, and not an extension of what Evidence Packets Phase I demonstrated. Phase I did not test retrieval, semantic search, or DICOM data at all; nothing above should be read as Phase I evidence for or against that separate line of inquiry.

Foundation and Next Steps

Article 01

Logs Aren't Enough

The original case for the Evidence Packet as a first-class artifact, this research set out to test.