# Prospective paper-strengthening protocol

Frozen scientific implementation: Git HEAD at protocol creation (recorded in
the execution manifest). Do not tune scientific code after seeing outcomes.
New runners are measurement adapters, not scientific transformations.

## Endpoints fixed before execution

1. Select the first 256 eligible, reference-consistent non-palindromic SNPs
   from the external GCST90018642 source whose coordinate keys do not occur
   in the August controlled benchmark. Assign cyclic unchanged, swap,
   complement and complement-plus-swap perturbations. Preserve p-values.
   Score coordinates, allele orientation, beta, SE, EAF and reported p-value
   against the pre-perturbation record; relative tolerance 1e-6, zero exact.
   Report exact matches separately. This is a disjoint-variant test from the
   same study, not an independent-study or blinded external validation.
2. Run Harmonizer's production download pipeline, GWASLab and tidyGWAS on
   those identical rows with explicit GRCh38 and documented semantics.
   Report available output/audit artifacts, original-value retention and
   named row dispositions. Do not infer usability or superior speed from
   this small run. Reference and identifier policies differ and must be
   reported; benchmark adapters may not manufacture native audit evidence.
3. Independently hand-specify synthetic full-download cases for canonical
   SNP transformations, palindromes with missing/compatible/incompatible
   population evidence, declared/undeclared indels and ratio/CI swaps.
   Record correct release, expected quarantine, incorrect release and
   unexpected quarantine separately. Synthetic oracles do not validate
   empirical population accuracy or hosted execution.
4. Sweep palindrome frequency tolerances on a fixed boundary grid. Report
   decision changes; no threshold optimization or calibration claim.
5. Test build detection using written synthetic build maps with direct,
   mixed and unavailable evidence. These test abstention at component scale.

Every run records dependency versions, implementation commit, inputs,
options, reference identity and output checksums. Failed cases remain in
results. Preserve third-party licensing; package small derived records only
when redistribution is permitted. Existing browser evidence is verified
against its saved files before updating the live draft.

## Interpretation

Retention is not correctness. Comparator differences are adjudicated rather
than automatically scored as errors. No universal accuracy, independence of
repeated scenarios, empirical palindrome calibration, usability improvement
or submission readiness is inferred. Author approvals and an immutable DOI
remain separate gates.

## Prospective second-source extension

After the GCST experiment, retain the same frozen scientific implementation,
four perturbation classes and numeric scoring rules. Select 1,024 eligible
FinnGen R13 T2D SNPs whose coordinate keys are absent from the earlier
99,999-row FinnGen subset and the August GCST truth set. Confirm the full
FinnGen source MD5 against the verified upstream value
`f21ddd163f3dee38f18d130eb13ad5fd`. FinnGen documents GRCh38, REF/ALT with ALT
as effect allele, and beta/SE conventions; binary log-odds semantics are
already documented in the paper's source register. Keep the source ALT
frequency unchanged except for known allele swaps. Skip multiallelic rsID
strings rather than choose an identifier. Require a reference-base match.
Use the first eligible record per coordinate for assay sampling; skip later
records at that coordinate. This makes coordinate-based scoring unambiguous
and is not a production rule for handling multiallelic records. The initial
preparation detected duplicate coordinates and stopped before any scientific
execution; this sampling clarification was made before the three tool runs.

Run the three configured tools on identical records. This independently
sourced study tests transformation recovery relative to the submitted
statistics, not association validity or empirical palindrome handling.
Report any failures. Retain the earlier runner files under
`measurement-source-v1/` and record the extended adapter separately; do not
revise the earlier 256-row endpoint or reinterpret its results.
