Chromatogram Peak Resolution Model for Nucleotide Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current chromatographic methods for nucleic acid sequencing face challenges in accurately identifying and resolving peaks, particularly in low-resolution regions, due to factors like peak overlap, stochastic signal emission, and background noise, leading to errors in base-calling and sequence determination.
Innovation Solution
The method involves determining resolution values of well-defined peaks and extrapolating these values to low-resolution regions to predict the number of constituent peaks, using a peak resolution model to resolve overlapping peaks and improve accuracy in base-calling, especially in the end regions of chromatograms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If chromatographic separation is used to separate nucleic acid fragments, then the sequence can be determined, but peak overlap and resolution degradation occur towards the end of the sequence
Solution Approach 1:
The patent applies preliminary action by establishing a peak resolution model using well-defined peaks from the beginning of the sequence, where resolution is high. This model is then used to predict and resolve convoluted peaks in the end region before final base-calling is performed. The resolution model is built in advance using reliable peak data and then applied to problematic regions.
Solution Approach 2:
The patent uses an intermediary approach by introducing a peak resolution model as a mediator between the observed chromatogram data and the final sequence determination. This model, derived from well-defined peaks, serves as a reference to interpret convoluted peaks in the end region, allowing accurate base-calling despite peak overlap.
2Reliability
If empirical base-calling methods are used, then the process is simple, but accuracy decreases in low-resolution regions
Solution Approach 1:
The patent performs preliminary action by pre-establishing a peak resolution model using well-defined peaks from the sequence beginning. This model is built in advance and stored for later use. When analyzing the end region with convoluted peaks, the pre-built model is retrieved and applied to predict the number of constituent peaks, enabling accurate base-calling without complex real-time calculations.
Solution Approach 2:
The patent applies copying by creating a simplified representation (copy) of peak resolution characteristics from well-defined peaks and applying this copy to convoluted peaks. The resolution model captures the essential resolution patterns and replicates them to interpret overlapping peaks, reducing the need for complex deconvolution algorithms.
3Measurement precision
If deconvolution methods are used to resolve peaks, then peak resolution improves, but computational complexity and sensitivity to noise increase
Solution Approach 1:
The patent applies parameter changes by transforming the problem from complex deconvolution to parameter prediction. Instead of attempting to mathematically deconvolute overlapping peaks, the method changes the approach to predicting peak resolution parameters (number of constituent peaks) based on the position and characteristics of well-defined peaks. This parameter-based approach is computationally simpler and less sensitive to noise.
Solution Approach 2:
The patent uses copying by creating a simplified model of peak resolution patterns from well-defined peaks and applying this model to convoluted peaks. Rather than performing complex deconvolution calculations, the method copies the resolution characteristics from reliable regions and applies them to interpret overlapping peaks, significantly reducing computational complexity.
4Adaptability or versatility
If the sieving medium resolution decreases with fragment length, then all fragments can be separated, but peak spacing and resolution become insufficient for accurate base-calling
Solution Approach 1:
The patent applies preliminary action by establishing a peak resolution model in the beginning of the sequence where peak spacing is adequate. This model captures the relationship between fragment length, peak spacing, and resolution. The model is then used to predict expected peak spacing in the end region, allowing compensation for the reduced resolution provided by the sieving medium at longer fragment lengths.
Solution Approach 2:
The patent uses feedback by continuously comparing the observed peak characteristics in the end region against the predictions from the peak resolution model. If the observed peak spacing or resolution deviates from the predicted values, the system can adjust its base-calling algorithm accordingly, using the model's expectations to guide the interpretation of overlapping peaks.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances the accuracy of nucleotide sequence identification by effectively resolving convoluted peaks and increasing the read-length of sequences, reducing errors associated with peak overlap and noise, thereby improving the overall reliability of chromatographic data analysis.
Implementation Method 1
The chain termination fragments are electrophoretically separated in a gel medium according to the fragment size
Implementation Method 2
chromatographic methods of nucleic acid sequencing utilize an electrophoretic sieving medium to separate DNA fragments on the basis of size
Implementation Method 3
the constituent compounds be labeled with a molecule that emits electromagnetic radiation, such as a fluorescent dye. This radiation can be detected by an optical detector sensitive in the spectral range of emitted radiation
Data Source
Figure 1(a)~1(b)
Figure 2
Figure 3
AI summary
The present invention relates to methods for resolving convoluted peaks in a chromatogram into one or more constituent peaks using peak resolution values. The methods of the invention determine empirical peak resolution values of "well- defined" or "isolated" peaks in the data, then extrapolate these empirical resolution values to peaks in neighboring regions to predict the number of constituent peaks at a given peak position. Predicted peak resolution values are compared to observed peak resolution values of low-resolution or convoluted peaks to determine the number of constituent peaks in the convoluted peaks. These methods enable extension of the region of data that can used for identifying nucleotide sequences, and increase base-calling accuracy in the low-resolution region (end region) of data.