Base Mutation Detection via Iterative Frequency Convergence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing base mutation detection algorithms in next-generation sequencing technology struggle to detect tri-allelic and tetra-allelic mutation sites and are not suitable for scenarios with low sequencing depth and large sample data, such as non-invasive prenatal testing.
Innovation Solution
A base mutation detection method that determines an initial frequency of specific bases at interested loci, calculates expected values, updates frequencies iteratively until convergence, and determines base mutation types and confidence levels, enabling detection in low-depth sequencing data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing base mutation detection algorithms are used, then detection capability for common mutations is maintained, but detection of tri-allelic and tetra-allelic mutation sites fails
Solution Approach 1:
The patent changes the fundamental parameters of the detection algorithm by introducing a new statistical model that accounts for multiple alleles. Instead of using traditional bi-allelic assumptions, the patent implements a multi-allelic framework that calculates genotype likelihoods across multiple possible alleles at each position, enabling detection of tri-allelic and tetra-allelic mutations while maintaining accuracy for common mutation types.
2Measurement precision
If algorithms are designed for high sequencing depth, then detection precision is improved, but applicability to low-depth sequencing data deteriorates
Solution Approach 1:
The patent implements a dynamic algorithm that automatically adapts to different sequencing depths. The statistical model adjusts its parameters and confidence thresholds based on the actual sequencing depth of the input data, allowing it to maintain detection precision across a wide range of sequencing depths from ultra-low (0.06x) to high coverage scenarios.
Solution Approach 2:
The algorithm dynamically changes its operational parameters based on the input data characteristics. When processing low-depth sequencing data, the patent adjusts the likelihood calculation thresholds and confidence intervals to account for higher uncertainty, whereas for high-depth data it uses stricter criteria, thereby optimizing detection precision for each scenario.
3Measurement precision
If algorithms process large sample data, then population-specific mutation spectrum accuracy is improved, but computational memory requirements increase
Solution Approach 1:
The patent segments the large-scale data processing into manageable chunks that can be processed independently. Instead of loading all sample data into memory simultaneously, the algorithm processes samples in batches while maintaining cumulative statistical results, thereby achieving accurate population-specific mutation spectra with significantly reduced memory requirements.
4Productivity
If algorithms are optimized for speed, then processing efficiency is improved, but stability in analyzing large sample data deteriorates
Solution Approach 1:
The patent performs preliminary actions by pre-calculating and storing reference data, mutation spectra, and statistical parameters before main analysis. This preprocessing step enables the algorithm to process large sample data efficiently and stably during the actual analysis phase, as the computationally intensive setup work has already been completed and cached for rapid retrieval.
Data Source
AI summary
Provided is a base mutation detection method, which includes: determining an initial frequency of sequencing data of samples being a specific base at an interested locus; calculating, based on the initial frequency, an expected value of each sample being the specific base at the interested locus; updating the initial frequency of the sequencing data of the samples being the specific base at the interested locus; further calculating the expected value of each sample being the specific base at the interested locus, further updating the initial frequency of the sequencing data of the samples being the specific base at the interested locus, and repeating the foregoing iteration until the expected value of each sample being the specific base at the interested locus converges; and determining, based on each converging expected value, a base mutation type and a mutation confidence at the interested locus of each sample.


