Variant-Specific Scaffold Alignment for Complex Variant Phasing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing (NGS) struggles with accurate detection and phasing of complex genetic variants, particularly in repetitive or highly homologous sequences, leading to false positives and negatives due to reference bias, and existing methods require large cohorts or family data, limiting their applicability to rare mutations.
Innovation Solution
A custom scaffold approach using variant-specific reference sequences for read alignment, incorporating additional markers to enhance phasing and separation of sequence reads, improving detection sensitivity and specificity for complex genetic variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single primary reference sequence is used for read alignment, then the alignment process is simple and fast, but detection sensitivity and specificity deteriorate due to reference bias in repetitive or highly homologous sequences
Solution Approach 1:
The reference sequence is segmented into multiple alternative loci representing different haplotypes. Each locus contains variant-specific sequences that serve as alignment targets for reads containing corresponding variants. This segmentation allows reads with complex variants to align to the most appropriate reference sequence, reducing reference bias and improving detection accuracy without requiring a single complex reference structure
Solution Approach 2:
Different regions of the reference sequence are assigned different qualities based on their representativeness for specific haplotypes. Alternative loci are constructed with locally optimized sequences that match specific variant combinations. This local quality enhancement ensures that reads from diverse haplotypes can find their optimal alignment target, improving both sensitivity and specificity in variant detection
2Measurement precision
If long-read sequencing is used to link multiple variants in one sequence read for direct phasing, then phasing accuracy is improved, but cost increases and throughput decreases
Solution Approach 1:
Instead of using expensive long-read sequencing to directly observe variant linkage, the patent creates computational copies of variant combinations in alternative reference loci. Short reads are aligned to these copied variant configurations, and phasing is inferred from alignment patterns. This copying approach achieves accurate phasing using inexpensive short-read sequencing with high throughput
3Adaptability or versatility
If statistical inference methods are used for phasing based on population or pedigree data, then phasing can be performed without long reads, but large cohorts or family members are required which limits applicability to rare mutations
Solution Approach 1:
Alternative reference loci are pre-configured with all possible variant combinations before sequencing. This preliminary action allows individual sample reads to be directly compared against predefined variant patterns without requiring population statistics or pedigree information. Rare mutations can be detected and phased by matching reads to corresponding pre-configured loci, eliminating the need for large sample sizes
4Measurement precision
If NGS algorithms for read-based phasing are used to computationally assemble overlapping reads, then haplotype blocks can be constructed, but sensitivity for small structural variants is limited and genome reference bias confounds phasing
Solution Approach 1:
Alternative reference loci serve as intermediaries between raw sequencing reads and final phasing results. These loci mediate the alignment process by providing variant-specific target sequences that guide reads to their correct haplotype assignments. This intermediary approach improves phasing sensitivity for structural variants while simplifying the alignment process compared to de novo assembly methods
Data Source
AI summary
Disclosed are methods, systems and computer-program products for the determination of complex genetic variants. The disclosed methods, systems and computer-program products may include obtaining a mutant scaffold nucleotide sequence that comprises a sequence that includes mutations characteristic of the complex genetic variant; obtaining a wild-type scaffold nucleotide sequence having a wild-type sequence; generating an alignment of at least one sequence from the sample to the mutant scaffold and to the wild-type scaffold; and determining that the sample contains a mutation characteristic of the complex genetic variant based on alignment to the mutant scaffold and not the wild-type scaffold.


