Fine-Grained Count Bias Correction in DNA Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sequencing technologies face challenges in accurately determining and correcting for count bias in nucleic acid sequencing data, particularly in small sample fractions like fetal or tumor DNA, leading to systematic and random errors that obscure underlying biological relationships.

Innovation Solution

A method that partitions the genome into strata based on nucleotide content in small windows relative to each locus, allowing for fine-grained detection and correction of count bias by attributing each locus to a stratum and calculating expected counts to adjust for biases, thereby improving the accuracy of copy number estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current sequencing technologies are used to determine count bias in small sample fractions, then sequencing can be performed, but systematic and random errors obscure underlying biological relationships

Engineering Contradiction:
Improveaccuracy of count bias determinationVSAvoiderrors in sample preparation and sequencing
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The genome is partitioned into multiple strata based on nucleotide content characteristics (e.g., GC content). Each stratum is analyzed separately to determine count bias, allowing for fine-grained correction that accounts for local sequence composition variations. This segmentation enables more precise measurement of count bias in small sample fractions by breaking down the complex genomic data into manageable, composition-specific groups.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If alignment methods are used to compare sequenced DNA to reference sequences, then variations can be identified, but considerable computing power is consumed

Engineering Contradiction:
Improveidentification of variationsVSAvoidcomputing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

Count bias is determined and corrected in advance by analyzing the relationship between read counts and nucleotide content characteristics before performing full alignment procedures. By pre-calculating bias correction factors based on sequence composition, the method reduces the computational burden of subsequent alignment operations while maintaining accurate variation identification.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If fine grained partitioning based on nucleotide content is implemented, then count bias correction is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of copy number estimationVSAvoidcomplexity of partitioning and stratum attribution
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method uses nucleotide content parameters (such as GC content percentages) as the basis for partitioning the genome into strata. By changing the partitioning parameter from simple genomic coordinates to nucleotide composition-based groups, the system achieves fine-grained count bias correction that reflects actual sequence characteristics. This parameter change simplifies the conceptual framework while improving measurement precision.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11127485B2Techniques for fine grained correction of count bias in massively parallel DNA sequencing
Publication Date: 2021.09.21 ECHELON DIAGNOSTICS INC
  • US11127485B2 patent drawing
  • US11127485B2 patent drawing
  • US11127485B2 patent drawing

AI summary

Techniques for automated determination or correction of count bias are based on nucleic acid base content on a finer grained scale than a bin of interest in a target sequence. The techniques include obtaining a target sequence with bins where relative abundances indicate a condition and raw counts Hj of reads, from a subject, which start at each locus j. A partition indicates a fine-grained window at a position relative to a current locus and multiple strata indicating different base contents. Each locus is attributed to one stratum k(j). An expected count of each stratum, E(k), is determined based on Hj for j belonging to the stratum and a number of loci in the target belonging to the stratum. A copy number of a bin is based on a sum of E(k(j)) in the bin. Output data indicates condition of the subject based at least partly on the copy number.