Machine Learning Probe Intensity Prediction for Genotyping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genotyping technologies face challenges in accurately predicting probe intensity values, particularly due to variations in signal intensities caused by DNA sample preparation methods, individual differences, and reliance on external reference datasets for normalization, which can lead to inaccurate genotype calling and copy number variant (CNV) calling.

Innovation Solution

The use of machine learning models, such as linear regression, random forest, and neural networks, trained on sample-specific image data to predict probe intensity values, allowing for more accurate and personalized normalization without relying on external reference datasets, and enabling effective on-device processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If external reference datasets are used for normalization, then normalization can be performed, but accuracy of probe intensity prediction deteriorates due to individual differences and sample preparation variations

Engineering Contradiction:
Improveprobe intensity prediction accuracyVSAvoidnormalization accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements self-service by training machine learning models on sample-specific image data from the same sample being analyzed. The model learns probe intensity patterns directly from the individual sample's own imaging data, eliminating the need for external reference datasets. This self-service approach accounts for individual-specific variations in DNA sample preparation and imaging conditions, thereby improving both probe intensity prediction accuracy and normalization reliability simultaneously.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If machine learning models are trained on sample-specific image data, then probe intensity prediction accuracy improves, but device complexity increases

Engineering Contradiction:
Improveprobe intensity prediction accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by training the machine learning model once on sample-specific image data before performing genotype calling and CNV analysis. The model training is performed as a preliminary step that captures probe intensity patterns for the entire sample, and then these trained models are reused for multiple subsequent analyses. This preliminary training approach improves probe intensity prediction accuracy while limiting device complexity by avoiding repeated training operations.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If normalization relies on external reference datasets, then standardization is achieved, but genotype calling accuracy deteriorates due to inability to account for individual variations

Engineering Contradiction:
Improvenormalization adaptabilityVSAvoidgenotype calling accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements local quality by transitioning from global normalization using external reference datasets to local normalization using sample-specific machine learning models. Each sample receives customized normalization based on its own imaging characteristics and probe intensity patterns. This local approach maintains adaptability while improving genotype calling accuracy by accounting for individual-specific variations in DNA sample preparation, hybridization efficiency, and imaging conditions.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20230316054A1Machine learning modeling of probe intensity
Publication Date: 2023.10.05 ILLUMINA INC
  • US20230316054A1 patent drawing
  • US20230316054A1 patent drawing
  • US20230316054A1 patent drawing

AI summary

Systems, methods, and apparatus are described herein for training machine learning models to predict probe intensity values using sample-specific image data and/or applying the predicted probe intensity values. As described herein, sample-specific image may include a signal associated with a sample for a process probe in a microarray relating to a single individual. The machine learning model may be trained, using the sample-specific image data, to predict a probe intensity value. The probe intensity value may be a raw probe intensity value or a normalized probe intensity value. After being trained, the machine learning model may receive as input a probe sequence or probe features. The machine learning model may be used to predict a total probe intensity value based on the probe sequence or the one or more probe features.