ML Peptide Retention-Time Prediction for Mass-Spectrometry Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Proteomic analyses face challenges due to the high complexity and dynamic range of peptide abundances, limiting the accuracy and efficiency of peptide detection in high-throughput systems, particularly in samples with multiple proteins, as existing techniques are constrained by database size and signal interactions.

Innovation Solution

A machine-learning model, utilizing an encoder-decoder network with LSTM cells, is trained on peptide amino-acid characteristics and retention times to estimate retention times, enabling the construction of a comprehensive retention-time library for peptide detection, which is then used in conjunction with mass spectrometry to identify peptides within samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional mass spectrometry techniques are used for peptide detection, then peptide identification can be performed, but accuracy and efficiency are limited due to high complexity and dynamic range of peptide abundances

Engineering Contradiction:
Improvepeptide detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a retention time prediction model as an intermediary component that bridges chromatography separation and mass spectrometry detection. This model predicts retention times of peptides based on their amino acid sequences, enabling more accurate peptide identification by filtering candidate peptides before mass spectrometry analysis and reducing false positives through retention time matching.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive protein databases are used to identify peptides, then more peptides can be detected, but the finite size of databases limits detection accuracy

Engineering Contradiction:
Improvepeptide identification accuracyVSAvoiddatabase size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the approach from relying on database completeness to using predictive parameters (retention time based on amino acid characteristics) to improve identification accuracy. By incorporating retention time prediction into the identification process, the system can accurately identify peptides even with limited database coverage, as the retention time parameter provides an additional filtering criterion independent of database size.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If high-throughput peptide analysis is performed, then more samples can be processed, but signal interactions between multiple proteins reduce detection quality

Engineering Contradiction:
ImprovethroughputVSAvoiddetection quality
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the peptide identification process into distinct stages: retention time prediction based on amino acid sequences, filtering of candidate peptides using predicted retention times, and targeted mass spectrometry analysis. This segmentation allows high-throughput processing by efficiently filtering out non-candidate peptides before detailed analysis, while maintaining detection quality through retention time matching that reduces signal interference from co-eluting peptides.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12362041B1Methods and systems for using machine-learning models to estimate peptide-retention time
Publication Date: 2025.07.15 VERILY LIFE SCIENCES LLC
  • US12362041B1 patent drawing
  • US12362041B1 patent drawing
  • US12362041B1 patent drawing

AI summary

The present disclosure relates to a machine-learning computing system for training and running a machine-learning model to estimate peptide-retention time for a sample. The machine-learning model can be configured to process inputs that characterize an individual peptide and/or amino acids in the peptide and to output an estimated retention time within a liquid-chromatography column for the peptide. The machine-learning model can include an encoder-decoder model. The encoder and/or the decoder can include a neural network. A subset of peptides can then be identified that are associated with estimated retention times within a specific elution time period during which portion of the sample was eluted from a chromatography column, and mass-spectrometry data can be analyzed to determine which of the subset of peptides are present within the sample.