BRLMM-P Genotype Calling Model for Probe Array Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing data from biological probe arrays, such as Affymetrix GENECHIP® arrays, face challenges in accurately determining genotype information due to variations in probe intensities and the need for mismatch probes, which can complicate the estimation of cluster centers and genotypes.
Innovation Solution
The BRLMM-P model extends the RLMM model by using Bayesian probability to improve genotype calling, requiring only perfect-match probes and performing multiple chip analysis to estimate probe effects and allele signals, thereby reducing variance and improving accuracy without relying on mismatch probes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If mismatch probes are used to improve genotype determination accuracy, then reliability improves, but device complexity increases
Solution Approach 1:
The patent extracts and removes the mismatch probe component from the genotyping system, relying solely on perfect match probes to determine genotypes. This simplifies the probe array design while maintaining accuracy through improved computational methods for estimating cluster centers and accounting for probe effects.
Solution Approach 2:
The patent changes the analytical parameters by implementing improved cluster center estimation methods and variance modeling that specifically account for probe effects. This allows accurate genotype calling using only perfect match probes by compensating for probe-specific variations through statistical modeling.
2Measurement precision
If multiple chip analysis is performed to reduce variance and improve accuracy, then measurement precision improves, but loss of time increases
Solution Approach 1:
The patent merges the analysis of multiple chips into a unified statistical framework that simultaneously processes data from multiple arrays. This combines the variance-reduction benefits of multiple chip analysis with computational optimizations that reduce the time penalty, by integrating probe effect estimation and cluster center determination across all chips in a single coordinated analysis.
Data Source
AI summary
An embodiment of a method of analyzing data from processed images of biological probe arrays is described that comprises receiving a plurality of files comprising a plurality of intensity values associated with a probe on a biological probe array; normalizing the intensity values in each of the data files; determining an initial assignment for a plurality of genotypes using one or more of the intensity values from each file for each assignment; estimating a distribution of cluster centers using the plurality of initial assignments; combining the normalized intensity values with the cluster centers to determine a posterior estimate for each cluster center; and assigning a plurality of genotype calls using a distance of the one or more intensity values from the posterior estimate.


