Fine-Grained Count Bias Correction in DNA Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sequencing technologies face challenges in accurately determining and correcting for count bias in nucleic acid sequencing data, particularly in small sample fractions like fetal or tumor DNA, leading to systematic and random errors that obscure underlying biological relationships.
Innovation Solution
A method that partitions the genome into strata based on nucleotide content in small windows relative to each locus, allowing for fine-grained detection and correction of count bias by attributing each locus to a stratum and calculating expected counts to adjust for biases, thereby improving the accuracy of copy number estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current sequencing technologies are used to determine count bias in small sample fractions, then sequencing can be performed, but systematic and random errors obscure underlying biological relationships
Solution Approach 1:
The genome is partitioned into multiple strata based on nucleotide content characteristics (e.g., GC content). Each stratum is analyzed separately to determine count bias, allowing for fine-grained correction that accounts for local sequence composition variations. This segmentation enables more precise measurement of count bias in small sample fractions by breaking down the complex genomic data into manageable, composition-specific groups.
2Measurement precision
If alignment methods are used to compare sequenced DNA to reference sequences, then variations can be identified, but considerable computing power is consumed
Solution Approach 1:
Count bias is determined and corrected in advance by analyzing the relationship between read counts and nucleotide content characteristics before performing full alignment procedures. By pre-calculating bias correction factors based on sequence composition, the method reduces the computational burden of subsequent alignment operations while maintaining accurate variation identification.
3Measurement precision
If fine grained partitioning based on nucleotide content is implemented, then count bias correction is improved, but device complexity increases
Solution Approach 1:
The method uses nucleotide content parameters (such as GC content percentages) as the basis for partitioning the genome into strata. By changing the partitioning parameter from simple genomic coordinates to nucleotide composition-based groups, the system achieves fine-grained count bias correction that reflects actual sequence characteristics. This parameter change simplifies the conceptual framework while improving measurement precision.
Data Source
AI summary
Techniques for automated determination or correction of count bias are based on nucleic acid base content on a finer grained scale than a bin of interest in a target sequence. The techniques include obtaining a target sequence with bins where relative abundances indicate a condition and raw counts Hj of reads, from a subject, which start at each locus j. A partition indicates a fine-grained window at a position relative to a current locus and multiple strata indicating different base contents. Each locus is attributed to one stratum k(j). An expected count of each stratum, E(k), is determined based on Hj for j belonging to the stratum and a number of loci in the target belonging to the stratum. A copy number of a bin is based on a sum of E(k(j)) in the bin. Output data indicates condition of the subject based at least partly on the copy number.


