Synthetic Control Generation for Sequencing Data Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for normalizing sequencing coverage data in genomic profiling techniques often rely on paired normal samples or process-matched controls, which are not always available or may introduce sequencing noise specific to each sample.
Innovation Solution
The method involves generating a synthetic set of sequence read count data, known as a synthetic 'control sample', using a 'panel of normals' approach. This involves selecting suitable normal samples, performing multivariate analysis to characterize noise, projecting this decomposition onto individual sample data, removing noise components, and normalizing the data using the synthetic control sample.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If paired normal samples or process-matched controls are used for normalization, then the normalization process can be performed, but the availability is limited and sequencing noise specific to each sample is introduced
Solution Approach 1:
The patent creates a synthetic control sample by copying and aggregating data from multiple normal samples in a panel. Instead of relying on a single paired normal sample or process-matched control, the system synthesizes a composite control that represents the collective normal variation, thereby improving both availability and reliability of normalization
Solution Approach 2:
The synthetic control sample is constructed as a composite of multiple normal samples, combining their sequencing data to create a composite reference that captures the diversity of normal variation. This composite approach eliminates the limitations of single-sample controls while maintaining normalization reliability
2Measurement precision
If individual process-matched controls are used, then normalization can be performed, but sequencing noise specific to each sample is contributed
Solution Approach 1:
The patent merges multiple normal samples into a single synthetic control by aggregating their sequencing data. This combining approach dilutes sample-specific noise while preserving the collective normal signal, thereby improving measurement precision and reducing harmful sequencing noise
Solution Approach 2:
The system extracts and removes noise components from the synthetic control sample through iterative refinement processes. By identifying and eliminating noise elements while retaining the normal signal, the method achieves high normalization accuracy with minimal sequencing noise
3Adaptability or versatility
If a panel of normals approach is used to generate synthetic control, then the need for paired normal samples is eliminated, but computational complexity increases
Solution Approach 1:
The patent segments the complex task of normalization into distinct computational steps: constructing the panel of normals, synthesizing the control sample, identifying noise components, and iteratively refining the synthetic control. This segmentation makes the complex process more manageable and systematic
Solution Approach 2:
The system uses the normal samples themselves to create the synthetic control and identify noise patterns, rather than requiring external reference materials or complex experimental designs. The data speaks for itself, reducing the need for additional resources while maintaining flexibility
Data Source
AI summary
Methods and systems for generating a set of synthetic sequence read count data for use in normalizing sequence coverage data derived from patient sample are described. The disclosed methods may comprise receiving sequence read count data for each of a plurality of non-subject normal samples; generating a non-subject profile for the plurality of non-subject normal samples; receiving sequence read count data for a sample from a subject; generating a synthetic normal set of sequence read count data based on the non-subject profile; and normalizing the sequence read count data for the sample from the subject using the synthetic normal set of sequence read count data to generate normalized sequence read count data for one or more subgenomic intervals in the sample from the subject.


