QC and validation#
Every processing stage in the pipeline emits its own quality-control status records. A status record marks one run, or one locus set, as invalid; the pipeline filters that record’s inputs out before they reach the next stage, so a known-bad input is never silently carried forward into fine-mapping.
The dashed Validation & QC line above is exactly this mechanism: every
stage that can invalidate an input feeds it, and all four converge on the
same published Reports output.
Validation stages#
Four stages can emit a status record:
|
Reason |
Trigger |
Implementation |
|---|---|---|---|
|
|
A manifest row’s ancestry has no matching entry in |
Inline Groovy check in |
|
|
The per-study locus-breaker output has zero logical rows. |
|
|
|
The per-run collected-loci output has zero logical rows. |
|
|
|
Any ancestry in a locus set’s pairwise LD statistics has
|
|
MANIFEST, LOCUS_BREAKER, and LOCUS_COLLECTION records are keyed
by runId. LD_ANNOTATION records are keyed by
fineMappingLocusSetId instead — see Why the filtering key changes.
Status record schema#
Each stage writes newline-delimited JSON. A typical record looks like:
{"runId": "run-1", "path": "collected_loci/fine_mapping_locus_sets/run-1", "validationStage": "LOCUS_COLLECTION", "reason": "EMPTY_DATASET"}
The LD_ANNOTATION variant adds fineMappingLocusSetId in place of a
bare runId-scoped path. fineMappingLocusSetId is an MD5 digest of
the set’s sorted studyLocusId values, not a coordinate-derived string:
{"runId": "run-1", "fineMappingLocusSetId": "3f2a9c7e1b8d4f6a0c5e9b2d7a1f4c8e", "path": "stats.jsonl", "validationStage": "LD_ANNOTATION", "reason": "EMPTY_LD_PAIRS"}
Status records are published alongside the pipeline’s other outputs:
Stage |
Published path |
|---|---|
|
|
|
|
|
|
|
|
Why the filtering key changes#
The first three stages operate one-study-per-runId, so invalidation
accumulates forward by runId: a run marked invalid at MANIFEST is
excluded from LOCUS_BREAKER’s input, and a run marked invalid at either
MANIFEST or LOCUS_BREAKER is excluded from LOCUS_COLLECTION’s
input.
LOCUS_COLLECTION merges loci from multiple studies into shared
fineMappingLocusSetId groups, so a single runId no longer identifies
one output row from that point on. LD_ANNOTATION therefore filters its
own output by fineMappingLocusSetId instead of by runId, removing
any locus set with zero LD pairs in any ancestry before it reaches
FINE_MAPPING. No further validation stage runs after LD annotation.
Duplicate summary-statistics limitation#
Summary statistics must contain at most one row for each (studyId,
variantId) pair before LocusBreaker runs. Duplicate rows are not a harmless
formatting detail: they can contain different p-values, effect sizes, or
alleles, so there is no generally safe row to keep.
The two available LocusBreaker backends do not currently handle this case in the same way:
the collector removes every row for a duplicated variant before clumping;
Gentropy can rank duplicate rows during clumping before later processing removes ambiguous variants.
Consequently, a collector-versus-Gentropy comparison is not meaningful when the input contains duplicated variant IDs. The lead variant, locus boundary, and downstream fine-mapping input can differ even when both tasks complete successfully. This is a known limitation, not evidence that one backend is numerically wrong.
Until a shared preflight validator is available, reject or repair duplicated summary statistics before starting the pipeline. Do not silently keep the most significant duplicate. A simple DuckDB check is:
SELECT studyId, COUNT(*) AS n_rows,
COUNT(DISTINCT variantId) AS n_variant_ids,
COUNT(*) - COUNT(DISTINCT variantId) AS n_duplicate_rows
FROM read_parquet('summary_statistics.parquet')
GROUP BY studyId
HAVING COUNT(*) > COUNT(DISTINCT variantId);
The collector canonical-region command silently drops all rows for any
duplicated variantId before processing — both copies are removed, not
just the weaker one. This is consistent with the collector LocusBreaker’s
QUALIFY count(*) OVER (PARTITION BY studyId, variantId) = 1 filter.
Duplicate counts in the run report and a shared Gentropy preflight are
tracked in GitHub issue #11.
Triggering validation paths in stub tests#
The collector empty_status and collector check_ld_pair_stats stub
blocks are silent by default, since most stub fixtures are not meant to
exercise the invalid path. Force them on with:
params {
empty_status_stub_emit = true
ld_pair_stats_stub_emit = true
ld_pair_stats_stub_empty_locus_set_ids = ['3f2a9c7e1b8d4f6a0c5e9b2d7a1f4c8e']
}
empty_status_stub_emit covers both the LOCUS_BREAKER and
LOCUS_COLLECTION stages, since they share the same
COLLECTOR_EMPTY_STATUS process. MANIFEST validation needs no stub
flag: its UNREGISTERED_ANCESTRY check is a plain Groovy comparison
against params.ld_registry, so it behaves identically in stub and
non-stub runs.