Genomic Copy Number Variation Detection via Bayesian Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting genome copy number variations in cancer and inherited diseases face challenges due to additive noise and the need for sophisticated statistical algorithms, which can render data unreliable without proper noise correction.

Innovation Solution

A Bayesian approach using a probabilistic generative model with parameters pr and pb to minimize a score function, allowing for efficient dynamic programming implementation and noise correction, thereby determining regional changes in the genome and their associated copy numbers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If microarray methods are used to sample the genome uniformly and create oligonucleotides for ratiometric measurement, then high-throughput copy number measurement is achieved, but a large amount of uncharacterized additive noise remains that may render the data worthless

Engineering Contradiction:
Improvehigh-throughput measurementVSAvoiddata reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the genome into distinct regions (normal, amplified, deleted) based on copy number variations. By dividing the genomic data into segments and analyzing each segment separately with appropriate statistical models, the method can identify true biological signals while filtering out additive noise that would otherwise corrupt the entire dataset.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary statistical model that acts as a mediator between the raw microarray data and the final copy number interpretation. This model includes parameters for normal variation, amplification, and deletion, allowing the system to distinguish true biological variations from technical noise through probabilistic reasoning.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data from multiple sources with varying protocols and technologies are used, then comprehensive genomic analysis is achieved, but the procedure becomes complex and requires general algorithms based on minimal prior assumptions

Engineering Contradiction:
Improvedata integration capabilityVSAvoidanalysis procedure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent develops a universal statistical model that can handle multiple data sources with varying protocols and technologies. The model is designed to be multi-functional, accommodating different microarray platforms, normalization methods, and experimental designs through a unified framework that requires minimal source-specific assumptions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter changes to adapt the analysis procedure to different data sources. By allowing key parameters (such as noise variance, copy number estimates, and segment boundaries) to vary based on the specific characteristics of each data source while maintaining the overall model structure, the system can integrate diverse data without requiring completely separate analysis pipelines for each source.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If sophisticated statistical algorithms are used to correct noise in genomic data, then measurement accuracy is improved, but the computational complexity and time required for analysis increases

Engineering Contradiction:
Improvecopy number detection accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-defining the statistical model structure, including the distributional assumptions for normal, amplified, and deleted regions, before analyzing the actual genomic data. This preliminary setup allows for more efficient computation during the actual analysis phase, as the complex statistical framework is already in place and can be applied systematically across the dataset.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs dynamic programming and iterative optimization techniques that allow the statistical model to adapt dynamically to the data being analyzed. The algorithm progressively refines copy number estimates and segment boundaries through multiple passes, balancing computational efficiency with measurement precision by stopping when convergence criteria are met rather than requiring a fixed number of iterations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7937225B2Systems, methods and software arrangements for detection of genome copy number variation
Publication Date: 2011.05.03 NEW YORK UNIV
  • US7937225B2 patent drawing
  • US7937225B2 patent drawing
  • US7937225B2 patent drawing

AI summary

The present invention relates to systems, methods and software arrangements for the detection of variations in the copy number of a gene in a genome. These systems, methods and software arrangements are based on a simple prior model that uses a first process generating amplifications and deletions in the genome, and a second process modifying the signal obtained to account for the corrupting noise inherent in the technical methodology used to scan the genome. A Bayesian approach according to the present invention determines, e.g., the most plausible hypothesis of regional changes in the genome and their associated copy number. The systems, methods, and software arrangements can be are framed as optimization problems, in which a score function is minimized. The system, methods and software arrangements may be useful to assist the scientific study, diagnosis and/or treatment of any disease which has a genetic component, including but not limited to cancers and inherited diseases.