Machine Learning Probe Intensity Prediction for Genotyping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genotyping technologies face challenges in accurately predicting probe intensity values, particularly due to variations in signal intensities caused by DNA sample preparation methods, individual differences, and reliance on external reference datasets for normalization, which can lead to inaccurate genotype calling and copy number variant (CNV) calling.
Innovation Solution
The use of machine learning models, such as linear regression, random forest, and neural networks, trained on sample-specific image data to predict probe intensity values, allowing for more accurate and personalized normalization without relying on external reference datasets, and enabling effective on-device processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If external reference datasets are used for normalization, then normalization can be performed, but accuracy of probe intensity prediction deteriorates due to individual differences and sample preparation variations
Solution Approach 1:
The patent implements self-service by training machine learning models on sample-specific image data from the same sample being analyzed. The model learns probe intensity patterns directly from the individual sample's own imaging data, eliminating the need for external reference datasets. This self-service approach accounts for individual-specific variations in DNA sample preparation and imaging conditions, thereby improving both probe intensity prediction accuracy and normalization reliability simultaneously.
2Measurement precision
If machine learning models are trained on sample-specific image data, then probe intensity prediction accuracy improves, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by training the machine learning model once on sample-specific image data before performing genotype calling and CNV analysis. The model training is performed as a preliminary step that captures probe intensity patterns for the entire sample, and then these trained models are reused for multiple subsequent analyses. This preliminary training approach improves probe intensity prediction accuracy while limiting device complexity by avoiding repeated training operations.
3Adaptability or versatility
If normalization relies on external reference datasets, then standardization is achieved, but genotype calling accuracy deteriorates due to inability to account for individual variations
Solution Approach 1:
The patent implements local quality by transitioning from global normalization using external reference datasets to local normalization using sample-specific machine learning models. Each sample receives customized normalization based on its own imaging characteristics and probe intensity patterns. This local approach maintains adaptability while improving genotype calling accuracy by accounting for individual-specific variations in DNA sample preparation, hybridization efficiency, and imaging conditions.
Data Source
AI summary
Systems, methods, and apparatus are described herein for training machine learning models to predict probe intensity values using sample-specific image data and/or applying the predicted probe intensity values. As described herein, sample-specific image may include a signal associated with a sample for a process probe in a microarray relating to a single individual. The machine learning model may be trained, using the sample-specific image data, to predict a probe intensity value. The probe intensity value may be a raw probe intensity value or a normalized probe intensity value. After being trained, the machine learning model may receive as input a probe sequence or probe features. The machine learning model may be used to predict a total probe intensity value based on the probe sequence or the one or more probe features.


