Additive Smoothing for Low-Coverage Sequencing Data Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Low coverage sequencing data from nucleic acid samples is prone to noise, leading to false positive cancer calls and reduced classification performance, as existing noise reduction methods either exclude valuable data or are time-consuming and costly.
Innovation Solution
An additive smoothing technique is applied by introducing a pseudocount number to bincount values associated with genomic bins, reducing noise and improving classification performance without excluding valuable data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If low coverage sequencing data is used to increase productivity and reduce cost, then sequencing throughput increases and cost decreases, but noise increases leading to false positive cancer calls and reduced measurement precision
Solution Approach 1:
The patent applies additive smoothing by introducing a pseudocount parameter to modify the bincount values in the sequencing data. This parameter change transforms the raw count data into a smoothed dataset that reduces noise while preserving the underlying signal, thereby improving measurement precision without requiring higher sequencing coverage
Solution Approach 2:
The patent introduces a pseudocount as an intermediary value that mediates between the raw sequencing counts and the final cancer detection results. This intermediary smoothing step acts as a buffer that reduces the impact of stochastic noise in low coverage data while maintaining the ability to detect true cancer signals
2Measurement precision
If existing noise reduction methods are applied to improve measurement precision, then false positives decrease, but valuable sequencing data is excluded or processing time increases
Solution Approach 1:
Instead of excluding low coverage bins or samples, the patent modifies the parameter space by adding pseudocounts to all bins including those with zero or low counts. This approach transforms the entire dataset rather than filtering it, preserving all valuable sequencing information while reducing noise through the mathematical smoothing effect
Solution Approach 2:
The patent extracts and removes only the noise component from the sequencing data through additive smoothing, while preserving the true biological signal. By separating the noise reduction function from data exclusion, the method maintains data integrity while improving measurement precision
Data Source
AI summary
Systems and methods for reducing noise for the analysis of low coverage sequencing data from a nucleic acid sample using a method, including: receiving, at an input component of the system, a set of sequence reads associated with the nucleic acid sample; allocating, using a processor component of the system, the set of sequence reads into a plurality of genomic bins; and introducing, subsequent to the allocating, a pseudocount number to bincount values to produce a smoothed dataset, wherein each of the bincount values is associated with one of the plurality of genomic bins. Other aspects are described and claimed.


