Variant-Specific Read Matching for Low-Frequency Variant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for detecting variants in a data set often result in messy data sets with true and false positive variants, missed relevant instances, and false negatives due to aligning reads to a reference data set and characterizing mismatches, rather than systematically checking for catalogued variants.
Innovation Solution
A method that utilizes a variant-specific unique data set, where the reference data set does not include the variant, involves determining matching criteria for each read, and identifying the variant based on a quantity and quality of reads that satisfy these criteria, allowing for highly sensitive and specific detection of variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods align reads to reference data set and characterize mismatches, then variant detection is performed, but messy data sets with true and false positive variants and false negatives are produced
Solution Approach 1:
Instead of aligning reads to reference and characterizing mismatches as variants, the patent inverts the approach by directly querying reads against a catalog of known variants. This reversal eliminates the need to interpret messy mismatch data and directly identifies known variants with high precision, resolving the contradiction between measurement precision and data set quality.
Solution Approach 2:
The patent performs preliminary action by pre-cataloging known variants before the actual detection process. This pre-prepared variant catalog allows direct querying during detection, avoiding the need to re-discover variants from scratch during each alignment process, thereby improving both accuracy and data quality.
2Measurement precision
If conventional methods align reads to reference data set, then variant detection is performed, but relevant variant instances may be missed
Solution Approach 1:
The patent inverts the conventional approach by querying reads against a pre-built catalog of known variants rather than deriving variants from alignment mismatches. This ensures that all known variant instances are systematically checked for, preventing missed detections while maintaining high precision through direct matching.
3Measurement precision
If variant-specific unique data set is used with multiple matching criteria, then sensitivity and specificity are improved, but computational complexity increases
Solution Approach 1:
The patent segments the variant detection process into distinct components: a pre-built catalog of known variants, multiple independent matching criteria (first matching criterion for exact matches, second matching criterion for approximate matches), and a decision logic that combines results. This segmentation allows the complex multi-criteria system to be managed as modular, independent components, reducing overall system complexity while maintaining high sensitivity and specificity.
Data Source
AI summary
Embodiments disclosed herein generally relate to detecting variants in a data set. A variant-specific unique data set, which includes a variant-inclusive portion that includes a particular variant and one or more other portions is accessed. The variant-specific unique data set corresponds to a particular region of a reference data set. A plurality of reads is received, with each read of the plurality of reads having been generated by processing a material collected from a subject. For each read of a subset of the plurality of reads, a first matching criterion and a second matching criterion can be determined to be satisfied, with the first matching criterion being more stringent than the second matching criterion. A data set associated with the subject can be determined to include the particular variant based on a quantity of reads in the subset. A result identified based on the particular variant can be output.


