DAG Reference Construct for Genotyping Structural Variations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genotyping methods using next-generation sequencing are computationally expensive and inefficient due to the complexity of aligning millions of short sequence reads, especially when dealing with structural variations, requiring massive computing power and often resulting in incomplete or inaccurate genotyping.
Innovation Solution
The use of a directed acyclic graph (DAG) reference sequence construct that accounts for multiple alleles and structural variations, allowing for multi-dimensional alignment of sequence reads, reducing computational requirements and improving accuracy by aligning reads directly to the construct rather than comparing them to multiple reference sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional pairwise alignment methods are used to align sequence reads to a reference, then alignment accuracy can be maintained, but computational cost and processing time increase exponentially with the number of reads and structural variations
Solution Approach 1:
The patent segments the alignment problem by dividing the set of reads into smaller batches or groups that can be processed independently. Instead of performing exhaustive pairwise comparisons between all reads and all possible reference sequences, the method processes reads in manageable segments, reducing the computational burden while maintaining alignment accuracy through systematic coverage of all read-reference pairs.
Solution Approach 2:
The patent applies preliminary action by pre-processing the reference sequences to create an optimized alignment structure or index before actual read alignment. This may involve pre-computing alignment scores, creating reference graphs, or organizing reference sequences in a manner that facilitates faster lookup during the actual alignment process, thereby reducing the computational power needed during the main alignment operation.
2Reliability
If multiple reference sequences are used to account for structural variations, then genotyping accuracy improves, but the complexity of the alignment process increases significantly
Solution Approach 1:
The patent merges multiple reference sequences into a unified alignment framework or composite reference structure. Instead of treating each reference sequence separately and performing independent alignments, the method combines them into a single integrated structure that allows reads to be aligned against multiple references simultaneously or in a coordinated manner, reducing process complexity while maintaining the ability to detect structural variations.
Solution Approach 2:
The patent creates a universal alignment approach that can handle multiple reference sequences and various types of structural variations through a single standardized process. This multi-functional alignment method is designed to work with diverse reference sequences and variation types without requiring separate specialized procedures for each case, thereby reducing overall process complexity while maintaining genotyping accuracy.
3Loss of information
If exhaustive comparison of all reads against all reference sequences is performed, then complete genotyping information is obtained, but processing time becomes prohibitively long
Solution Approach 1:
The patent applies partial action by performing alignment operations on a selective or sampled basis rather than exhaustively comparing every read against every reference sequence. The method may use sampling strategies, filtering criteria, or priority-based processing to identify and align only the most relevant read-reference pairs, obtaining sufficient genotyping information without the prohibitive time cost of complete exhaustive comparison.
Data Source
AI summary
The invention provides methods and system for making specific base calls at specific loci using a reference sequence construct, e.g., a directed acyclic graph (DAG) that represents known variants at each locus of the genome. Because the sequence reads are aligned to the DAG during alignment, the subsequent step of comparing a mutation, vis-a-vis the reference genome, to a table of known mutations can be eliminated. The disclosed methods and systems are notably efficient in dealing with structural variations within a genome or mutations that are within a structural variation.


