Sequence Variant Detection in Enriched Genomic Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting low-frequency sequence variants in enriched samples are cumbersome and prone to errors, often resulting in false positives or missed calls, especially in cases of mosaic samples with complex genomic regions and indels, where distinguishing genuine mutations from sequencing errors is challenging.
Innovation Solution
A method involving obtaining sequence reads from an enriched sample, assembling them into discrete sequence assemblies, determining true variants by examining the sequence reads, and optionally verifying against known mutations, with a computer system and program for performing this analysis to output a report on sequence variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing analysis methods are used on enriched samples, then high uniformity and read depths are achieved, but detection of low frequency variants becomes erroneous with false positives and missed calls
Solution Approach 1:
The method segments the analysis process into distinct phases: (i) assembling sequence reads into discrete sequence assemblies representing potential variants, (ii) examining the constituent sequence reads of each assembly to determine authenticity, and (iii) optional verification against known mutations. This segmentation allows for targeted application of different analytical strategies to different aspects of variant detection, improving overall accuracy while reducing false positives and missed calls.
2Reliability
If multiple separate methods are used to address artifacts, then comprehensive coverage is achieved, but the solution becomes cumbersome and discordant between methods
Solution Approach 1:
The method merges multiple analytical functions into a single integrated workflow. The assembly process simultaneously handles variant identification, while the examination of constituent reads addresses artifact detection, and optional comparison with known mutations provides verification. This unified approach eliminates the need for multiple separate methods, reducing complexity and ensuring consistent results across all analysis stages.
Solution Approach 2:
The discrete sequence assembly structure serves multiple functions: it represents potential variants, provides a framework for examining constituent reads to distinguish true variants from artifacts, and enables optional comparison with known mutations. This multi-functionality allows a single method to comprehensively address various challenges including low frequency variant detection, artifact filtering, and complex event handling without requiring separate specialized tools.
3Measurement precision
If statistical evaluation at individual variant sites is performed, then heterozygote calls can be supported, but low frequency variants in mosaic samples become indistinguishable from sequencing errors
Solution Approach 1:
The method transitions from evaluating variants at single nucleotide positions to examining variants in the context of discrete sequence assemblies composed of multiple reads. This dimensional shift from point-based to assembly-based evaluation provides additional context and information. By examining the constituent reads of each assembly and their alignment patterns, the method can distinguish true low frequency variants from sequencing errors even when they occur at very low frequencies in mosaic samples.
Data Source
AI summary
Provided herein is a method for identifying a sequence variant in an enriched sample. In certain embodiments, this method may comprise: (a) obtaining: (i) a plurality of sequence reads from a sample that has been enriched for a genomic region and (ii) a reference sequence for the genomic region; (b) assembling the sequence reads to obtain a plurality of discrete sequence assemblies that correspond to potential variants; (c) determining which of the potential variants are true and which are artifacts by examining the sequence reads that make up each of the discrete sequence assemblies; (d) optionally determining whether each of the true potential variants contains a mutation that is known to be associated with the reference sequence; and (e) outputting a report indicating whether the sample comprises a sequence variant.

