Workflows ========= Locus breaker ------------- ``LOCUS_BREAKER`` accepts manifest-derived study metadata and summary-statistics paths. The backend is selected with ``params.locus_breaker_method``: * ``collector`` — DuckDB-based collector implementation, the default. * ``gentropy`` — Gentropy Spark implementation for parity and comparison. The workflow passes native Nextflow tuples containing a metadata ``Map`` and a ``Path``. This keeps task hashing stable and preserves ``-resume`` behavior. Duplicate variant limitation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before choosing either LocusBreaker backend, validate that each summary statistics input has one row per ``(studyId, variantId)``. The collector and Gentropy implementations currently resolve duplicate rows differently, so their outputs are not a valid parity comparison for affected studies. See :ref:`duplicate-summary-statistics-limitation` for the failure mode and a DuckDB preflight query. This limitation is tracked in GitHub issue #11. Locus collection ---------------- ``LOCUS_COLLECTION`` groups locus outputs by ``runId`` and runs one ``COLLECT_CANONICAL_REGIONS`` task per run, sweeping each run's loci into disjoint canonical regions and materializing one published ``fineMappingLocusSetId`` per region that has a qualifying variant for every input study — see :doc:`locus-processing` for how regions are built and published. The workflow emits ``ch_full_overlap_loci`` (the *Collected Loci* candidate sets passed to later fine-mapping stages) alongside ``ch_partial_overlap_loci`` and ``ch_non_overlap_loci``, which are always empty under the current canonical-region algorithm — they are vestigial channel names from an earlier, since-removed classification step and carry no data. Locus annotation ---------------- ``LOCUS_ANNOTATION`` runs LD annotation once per ``fineMappingLocusSetId``. It emits the locus-set Parquet file, the MultiAncestryPairwiseLD dataset, and the study metadata carried by the channel. The LD backend is selected with ``params.ld_annotation_method``: * ``gentropy`` uses the Spark/Hail implementation and each registry entry's ``vi_path`` and ``bm_path``; * ``hailing_ducks`` uses the collector image with Hailing Ducks v1.1.0 and each registry entry's native hg38 ``ht_path`` and ``bm_path``. Both backends emit the same flat ``MultiAncestryPairwiseLD`` contract. The Hailing Ducks adapter converts pipeline variant IDs such as ``1_100_A_AT`` to the native ``chr1_100_A_AT`` lookup convention, maps them back on output, and uses the study metadata to query each study-locus only against its registered ancestry. Hailing Ducks is the supported default and the only backend configured by the repository profiles. The Gentropy backend remains available through its local modules for comparison or specialized runs, but is discouraged for routine execution because its Spark-based processing is slower. Fine-mapping ------------ ``FINE_MAPPING`` consumes valid records from ``LOCUS_ANNOTATION`` after empty LD results have been filtered. MultiSuSiE is the currently integrated and supported method. The repository also contains SuSiEx and SuShiE modules, but their full pipeline integration is pending and they should not be used for production runs yet: * ``multisusie`` — integrated and supported; * ``susiex`` — pending full integration; * ``sushie`` — pending full integration. Select methods with ``params.fine_mapping_methods``. The default is ``['multisusie']``. Do not add SuSiEx or SuShiE to production runs until their integration work is complete. For development or stub-level comparison: .. code-block:: groovy params.fine_mapping_methods = ['multisusie'] Each method runs once per ``fineMappingLocusSetId`` with the same annotated locus, pairwise LD, and study metadata. Metadata is serialized as JSONL inside each task, so no separate metadata process is required. Every method emits a metadata-preserving record containing: * ``study_locus_path`` — the Gentropy-compatible StudyLocus Parquet output; * ``extended_results_path`` — the AnnData H5AD output with component-level posterior results and provenance; * ``stats_path`` — the JSON status record for the locus-set fit; * ``runId``, ``fine_mapping_locus_set_id``, and the input ``metas``. Images are configured independently with ``params.multisusie_container``, ``params.susiex_container``, and ``params.sushie_container``. The MultiSuSiE default is used by the supported route; the SuSiEx and SuShiE images are pinned for pending integration work. Repository defaults are defined in ``nextflow_schema.json``. The current pinned defaults are: .. code-block:: groovy params { multisusie_container = 'ghcr.io/project-defiant/multisusie:1982e8138f83778aaf0617a1ffd8753f2561775a' multisusie_purity_min_r2 = 0.01 susiex_container = 'ghcr.io/project-defiant/susiex/susiex:600051c01bdd14a2783d34066f8aaaa0e3dececf' sushie_container = 'ghcr.io/project-defiant/sushie/sushie:dd7e4ae3cb96cc7c6f5718e0bf4fbe77ab07bf93' } The Google Cloud profiles likewise set ``params.collector_container`` to the pinned Collector image. Local development overrides are supported, but should be explicit and should not replace the repository defaults unintentionally. Method options use the process ``task.ext.args`` interface. MultiSuSiE purity and low-memory settings are managed by the pipeline and must not be overridden through ``task.ext.args``. For example: .. code-block:: text withName: MULTISUSIE_FINE_MAPPING { ext.args = '--L 10 --max-iter 100' } withName: SUSIEX_FINE_MAPPING { ext.args = '--n-sig 10 --max-iter 200' } withName: SUSHIE_FINE_MAPPING { ext.args = '--L 10 --rho 0.1 --max-iter 500' } Non-converged or otherwise non-reportable fits write ``stats.json`` and omit the result files according to each method's application contract. Hard input or output errors still fail the corresponding process.