# Supplementary Note S-N4: disjoint-variant comparison and uncertainty challenges

## Decision

The new results support a narrower, concrete contribution: a browser workflow
with explicit source-to-final action records. They do not support superiority
in scientific accuracy, speed or usability. No scientific production code was
changed or tuned during these experiments.

## Disjoint-variant comparison

Frozen Harmonizer implementation: `15e8878` (full commit in execution manifest).
GWASLab 4.2.1 and tidyGWAS 1.0.0 completed identical 256-row inputs.
The records are the first 256 eligible non-palindromic SNPs from the
GCST90018642 harmonized source that are absent from the August truth set.
Selection scanned 329 source rows and independently checked the reference
base. This is a same-study, same-region, disjoint-variant extension, not an
independent-study validation. Four perturbation classes contain 64 rows each.

| Endpoint | Harmonizer | GWASLab | tidyGWAS |
|---|---:|---:|---:|
| Rows retained | 256/256 | 256/256 | 256/256 |
| Allele-equivalent records | 256/256 | 256/256 | 256/256 |
| Beta, SE and p within relative 1e-6 | 256/256 | 256/256 | 256/256 |
| EAF within relative 1e-6 | 256/256 | 229/256 | 256/256 |
| Canonical ALT-effect orientation | 256/256 | 256/256 | 64/256 |

The scorer reorients equivalent allele pairs only for comparison; it does not
alter released outputs. tidyGWAS preserved valid alternative effect-allele
orientation. Its 192 noncanonical rows are not scientific errors. All 256
GWASLab frequencies agree with the independently calculated float32
conversion and swap operation. Maximum absolute EAF error is
4.206286075e-8; the 27 relative-tolerance failures are precision differences,
not allele-recovery errors. Harmonizer had 179/256 exact decimal matches
across the canonical numeric fields; its new endpoint is tolerance-based,
not exact numerical recovery. Small-run timings are recorded for diagnosis
but are not a performance benchmark.

## Native evidence artifacts

Harmonizer's production download pipeline wrote 256 source-indexed audit
records: 64 unchanged, 64 swaps, 64 complements and 64 complements with
swaps. All 256 original/final allele, beta and EAF pairs were checked against
the independent input and truth files; 192 records underwent transformations.

GWASLab saved harmonized rows with STATUS and a native execution log; this
configured output did not include original-value columns. The original input
was retained separately by the benchmark. This is not a claim that other
GWASLab configurations cannot supply equivalent evidence.

tidyGWAS natively saved a complete 256-row raw Parquet snapshot and row IDs,
which support reconstruction of original-to-result changes. Thus, traceability
is not unique to Harmonizer. Harmonizer's narrower convenience is an explicit
action record supplied with its browser-oriented download workflow. This
artifact comparison does not measure human usability or scripting effort.

## Independently sourced FinnGen extension

Before running a second-source comparison, an addendum retained the same
scientific implementation, four perturbation classes and scoring rules.
The first 1,024 eligible FinnGen R13 T2D SNPs outside the earlier 99,999-row
FinnGen subset and August GCST truth set were selected after scanning 101,312
source rows. Eight later records at already selected coordinates were skipped
to keep assay scoring unambiguous. The initial preparation stopped on
duplicate coordinates before scientific execution; the sampling clarification
was documented before all three tool runs. No production algorithm was tuned.

The full FinnGen MD5 matched the verified upstream value
`f21ddd163f3dee38f18d130eb13ad5fd`; its SHA-256 is
`76424434f5c966ad4da6d471eff0f6fcd01731004568aa6f3922d5ae52fad305`.
The [FinnGen download documentation](https://github.com/FINNGEN/finngen-documentation/blob/master/data-download.md)
documents GRCh38, REF/ALT and ALT-effect conventions. Binary log-odds
semantics and the R13 ALT-frequency mapping are documented in the paper's
source register. There was no coordinate overlap between the two new assays.

| FinnGen endpoint | Harmonizer | GWASLab | tidyGWAS |
|---|---:|---:|---:|
| Retained and allele-equivalent rows | 1,024/1,024 | 1,024/1,024 | 1,024/1,024 |
| Beta, SE and p within relative 1e-6 | 1,024/1,024 | 1,024/1,024 | 1,024/1,024 |
| EAF within relative 1e-6 | 1,024/1,024 | 731/1,024 | 1,024/1,024 |
| Canonical numeric fields exactly recovered | 1,024/1,024 | 0/1,024 | 256/1,024 |

All 1,024 GWASLab frequencies again agreed with the explicit float32 model;
maximum absolute EAF difference was 4.254913333e-8. tidyGWAS kept valid
alternative effect orientation in 768 rows, so its noncanonical exact-match
count is not an accuracy ranking. Harmonizer's 1,024 source-to-final audit
records verified, including 768 transformed rows. tidyGWAS again saved raw
records and row IDs. Across both studies, all three tools retained 1,280
records with equivalent alleles and beta/SE/p. Harmonizer verified 1,280 audit
records, including 960 transformations. This remains a small, chromosome-1
transformation assay, not blinded, never-before-processed data or association
validity evidence. The FinnGen full source had previously been processed in a
historical run; disjointness refers to the specified subset/truth registers.

## Synthetic production-download and component challenges

| Check | Result |
|---|---|
| SNP transformation scenarios | 4/4 passed |
| Palindrome scenarios | 6/6 passed |
| Indel release/quarantine and leftmost normalization assertions | 3/3 passed |
| Original/final and named-quarantine audit assertions | 7/7 passed |
| Ratio/CI scenarios | 1/2 passed |
| Written-map build detection/abstention cases | 4/4 passed |

Overall: 25/26 prespecified assertions passed. Of 14 production-download
scenarios, 13 had the expected disposition and fields. All five expected
quarantines occurred; eight of nine expected releases occurred. No incorrect
release was observed in these small scenarios. One valid ratio/CI swap was
unexpectedly quarantined with `orientation_invariant_failed` and
`confidence_interval_orientation_mismatch`. Its input was OR=0.5,
CI=[1/3, 2/3], and expected reference-oriented output OR=2, CI=[1.5, 3].
The unchanged OR/CI control passed. This is a current local integration
limitation, not positive empirical OR/CI validation; do not describe all
ratio/CI transformations as working end to end.

The first attempt used an invalid option enum before scientific execution;
the adapter was corrected from `flag` to `warn`. The next attempt scored the
audit against an incorrect column alias; the scorer was corrected to actual
`original_EA`/`original_NEA` fields and rerun. Original attempts remain saved.
No production implementation changed. The ratio failure remains in the
machine-readable results.

The fixed 24-point palindrome sweep tested eight EAF values at tolerances
0.01, 0.05 and 0.10 with reference EAF=0.1. Direct/inverse matches were kept
and the other six values were quarantined at every tolerance. This sparse
grid is a sensitivity check only; no empirical threshold calibration follows.
No full-workflow liftover, empirical ancestry-panel or independent-study
accuracy claim follows from these synthetic checks.

## Historical control exclusion accounting

The archived 100k GCST90018642 audit contains all 100,000 source records and
accounts for its 16,618 exclusions: 12,491 unresolved palindromes and 4,127
ambiguous indels. The audit agrees with the summary. Of the palindromes,
11,679 are outside the 0.42-0.58 frequency band: absence of compatible
orientation evidence, rather than midpoint frequency alone, caused their
quarantine. This explains policy-driven loss; it does not adjudicate whether
each excluded variant could be correctly recovered with additional evidence.

## Browser evidence recheck

The nine-file September 30 browser evidence manifest matches all saved
SHA-256 values. Both Firefox and Safari audit ZIPs pass CRC checks and all
gzip members decompress successfully. Their browser report records complete
anonymous execution and data/audit downloads on distinct 500-row Pan-UKBB
subsets. These existing runs, not new browser executions in this experiment,
support updating the stale live-draft browser limitation. They are smoke
checks, not scientific accuracy or signed-in-session validation.

## Reproduction and remaining boundaries

`STRENGTHENING_PROTOCOL_20260930.md` fixes endpoints; the adjacent results JSON
and execution manifest preserve outcomes, input/reference checksums,
dependency versions and frozen implementation file hashes. Runners are in
`scripts/paper_strengthening_*.py`, `scripts/paper_strengthening_tidygwas.R`
and `scripts/score_paper_strengthening.py`. Raw run evidence is outside Git
at `results/benchmarks/paper_strengthening_20260930/` in the parent workspace.
Third-party data and reference redistribution require their own terms.

Further work remains for independent-study truth, empirical population
calibration, representative excluded-row adjudication, measured user tasks,
matched scale performance, the OR/CI integration issue, and row-level
identifier-only auditing. Author sign-offs, repository/release availability,
provider-boundary verification and Zenodo deposition remain separate gates.
