Genomic Sequencing Bias Normalization via GC Content Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sequencing technologies face challenges in accurately analyzing genetic variations due to sequencing bias, which can lead to non-uniform distribution of reads across a genome and impair effective data analysis.
Innovation Solution
A system comprising memory and microprocessors is configured to perform processes that reduce sequencing bias by generating relationships between local genome bias estimates and bias frequencies, comparing these relationships between test and reference samples, and normalizing sequence read counts to minimize bias.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sequencing is performed across the entire genome, then comprehensive genetic information is obtained, but sequencing bias causes non-uniform distribution of reads leading to inaccurate analysis
Solution Approach 1:
The patent applies parameter changes by normalizing read counts based on GC content parameters. It calculates expected read counts for each genomic region considering GC content, and adjusts actual read counts to account for bias. This transforms the raw sequencing data into normalized data that corrects for systematic biases, thereby improving measurement precision without requiring more uniform physical read distribution
Solution Approach 2:
The patent introduces an intermediary normalization process between raw sequencing and final analysis. It uses GC content as an intermediary parameter to model and correct bias. The normalization factor acts as a mediator that translates biased read counts into corrected values, enabling accurate genetic variation analysis despite non-uniform read distribution
2Measurement precision
If bias correction methods are applied to improve data quality, then analysis accuracy improves, but computational complexity and processing time increase
Solution Approach 1:
The patent segments the genome into regions with similar GC content characteristics and processes each region independently. It divides the normalization task into manageable segments based on GC content bins, calculating normalization factors for each segment rather than treating the entire genome uniformly. This segmentation reduces computational complexity while maintaining correction accuracy
Solution Approach 2:
The patent applies partial action by focusing normalization efforts on regions where bias most significantly impacts analysis. It identifies and prioritizes correction for genomic regions with extreme GC content that cause the most severe bias, rather than uniformly applying complex corrections across all regions. This selective approach improves accuracy where needed while limiting unnecessary computational overhead
Data Source
AI summary
Provided herein are methods, processes, systems and machines for non-invasive assessment of genetic variations. In particular, provided herein are methods, processes, systems and machines for non-invasive assessment of copy number variations. In some aspects, copy number variations include aneuploidies (e.g., trisomy 13, 18, or 21). In some aspects, copy number variations include microdeletions or microduplications.


