Locus-Dependent Bias Correction for Non-Invasive Prenatal Diagnostics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sequencing technologies face challenges in accurately determining and correcting for count bias in nucleic acid sequencing, particularly at a finer grained scale, which can obscure underlying biological relationships and lead to errors in diagnosing medical conditions such as fetal anomalies and cancerous tumors due to systematic and random errors in sample preparation and alignment processes.
Innovation Solution
A method and system for measuring nucleic acid sequence abundances by obtaining data on target sequences at multiple loci, determining alignment and observed variations, and calculating locus-dependent observed variations to correct for count bias, allowing for more precise determination of copy numbers and presentation of condition data based on these calculations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional bin-level bias correction is used, then the process is simpler, but the accuracy of copy number determination is insufficient due to coarse-grained correction
Solution Approach 1:
The patent segments the bias correction process into two distinct levels: bin-level correction for broad regional biases and locus-level correction for fine-grained local biases. This segmentation allows the system to achieve high precision copy number determination by applying appropriate correction granularity to different genomic regions, resolving the contradiction between accuracy and complexity.
Solution Approach 2:
The patent implements local quality by applying locus-specific bias correction factors that are tailored to the unique characteristics of each genomic location. Instead of uniform correction, the system calculates and applies individualized correction factors based on local sequence context, GC content, and other position-specific variables, thereby improving measurement precision without requiring complete process redesign.
2Measurement precision
If locus-level bias correction is implemented, then measurement precision improves, but computational resources and processing time increase
Solution Approach 1:
The patent performs preliminary action by pre-calculating and storing locus-level bias correction factors in lookup tables before actual copy number analysis. These pre-computed correction factors capture the computational intensity of locus-level analysis while enabling rapid application during clinical testing, thereby reducing real-time computational power requirements while maintaining high measurement precision.
Solution Approach 2:
The patent uses copying by creating and storing correction factor profiles for representative genomic regions that can be reused across multiple samples. Instead of recalculating locus-level biases for every sample, the system copies and applies established correction patterns, significantly reducing computational power requirements while maintaining accurate bias correction.
3Reliability
If more correction factors are applied, then reliability of genetic condition identification improves, but the complexity of data processing increases
Solution Approach 1:
The patent extracts and isolates the most critical bias factors (GC content, sequence context, mappability) that have the greatest impact on copy number accuracy. By extracting only the essential correction factors rather than attempting to correct for all possible variables, the system improves diagnostic reliability while keeping data processing complexity manageable through focused correction on high-impact parameters.
Data Source
AI summary
Techniques for measuring abundances of sequences includes obtaining first data that indicates a target sequence at a plurality of loci, wherein the target sequence comprises a plurality of bins of loci for which a relative abundance is indicative of a condition of interest. Second data is determined that indicates alignment with the target sequence of reads of DNA fragments in a sample from the subject. Third data is determined that indicates locus dependent observed variations in abundance. A raw count Hj of reads is determined that start at each locus j; and, a copy number of a first bin is determined based on a sum over all loci in the first bin of expected counts for each partition weighted by the locus dependent observed variations. Output data that indicates condition of the subject based at least in part on the copy number of the first bin is presented.


