Copy Number Noise Measurement for Targeted Panel Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Copy number analysis in targeted panel sequencing is challenging due to small covered genomic regions, inability to use segmentation algorithms, lack of matched normal analysis, variability of DNA quality, and wide range of copy number alterations, leading to noise and false positive/negative calls.
Innovation Solution
A method to determine statistical noise level by obtaining sequencing data, aligning reads, calculating copy number values, determining standard error/standard deviation ratios, and normalizing by genomic locus size, allowing for noise reduction and improved copy number calling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If targeted panel sequencing is used for copy number analysis, then sequencing cost and time are reduced, but noise in copy number data increases leading to false positive and false negative calls
Solution Approach 1:
The patent introduces an in-silico normal cohort as an intermediary reference to bridge the gap between targeted panel sequencing data and reliable copy number analysis. This virtual normal cohort, generated by simulating sequencing data from normal genomes, serves as a mediator that enables accurate copy number calling without requiring physical matched normal samples, thereby maintaining both sequencing efficiency and analytical reliability
Solution Approach 2:
The patent applies preliminary action by pre-generating and storing in-silico normal sequencing data and copy number profiles before actual sample analysis. This pre-computed reference data is then used during clinical analysis to rapidly determine copy number alterations, eliminating the need for time-consuming matched normal sequencing while ensuring consistent and reliable results
2Measurement precision
If matched normal analysis is performed, then copy number accuracy is improved, but cost and availability requirements increase
Solution Approach 1:
The patent uses copying by creating in-silico copies of normal genome sequencing data to generate a virtual normal cohort. These synthesized normal samples replicate the characteristics of actual normal sequencing data, providing a reference for copy number analysis without requiring physical normal tissue samples. This copying approach maintains measurement precision while eliminating sample availability constraints
Solution Approach 2:
The patent applies parameter changes by transforming the fundamental assumption of copy number analysis from requiring physical normal samples to using computationally generated normal data. By changing the reference from biological material to in-silico simulated data, the method maintains copy number accuracy while dramatically simplifying sample requirements and reducing costs
3Reliability
If segmentation algorithms are used for copy number analysis, then analysis robustness is improved, but applicability to targeted panel data is reduced due to gaps between covered regions
Solution Approach 1:
The patent applies local quality by adapting the analysis approach to work with the specific characteristics of targeted panel data, where coverage is localized to specific genomic regions rather than continuous. The method processes copy number information from discrete covered regions independently and aggregates results, rather than requiring continuous coverage for segmentation algorithms to function
Solution Approach 2:
The patent inverts the traditional approach by not attempting to force targeted panel data into segmentation algorithm frameworks. Instead, it reverses the logic by using the in-silico normal cohort to directly compare against target sample data, eliminating the need for segmentation and making the method universally applicable to any targeted panel design regardless of coverage gaps
Data Source
AI summary
The present invention relates to a method for determining the statistical noise level in the calculation of a subject's genetic copy number value in massively parallel nucleic acid sequencing data derived from a sample, as well as a method for determining a subject's genetic copy number value in massively parallel nucleic acid sequencing data derived from a sample and a method to determine a subject's genetic copy number value for stratifying the subject for cancer therapy.


