MS1 Peptide Quantification Using Transfer Learning Across LC-MS Runs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing MS-based bottom-up proteomics methods, particularly Match-Between-Run (MBR) algorithms, struggle to scale accurately for large-scale datasets due to combinatorial complexity, leading to incomplete peptide identifications and quantification across multiple runs.
Innovation Solution
A novel MS1-XIC-based algorithm using transfer learning and ion mobility separation, combined with semi-supervised machine learning, adapts global retention time and ion mobility prediction models to local conditions for each run, enabling accurate and scalable quantification of peptides and proteins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Match-Between-Run (MBR) algorithms are used for quantification across multiple runs, then quantitative consistency is improved, but scalability to large cohorts deteriorates due to combinatorial complexity
Solution Approach 1:
The patent segments the large-scale quantification problem into independent run-specific analysis units. Each run is processed separately with its own peptide identification and quantification, eliminating the need for pairwise comparisons across all runs. This segmentation allows linear scalability while maintaining quantitative consistency through centralized peptide atlas construction and imputation strategies.
Solution Approach 2:
The patent introduces a peptide atlas as an intermediary structure that aggregates peptide information across all runs. This atlas serves as a mediator that enables consistent peptide identification and quantification without requiring direct pairwise comparisons between runs. The atlas-based approach decouples the relationship between runs, allowing independent processing while maintaining global consistency.
2Loss of information
If MBR algorithms align LC gradients and transfer peptide identifications, then missing values are reduced, but computational complexity increases dramatically for hundreds or thousands of samples
Solution Approach 1:
The patent performs preliminary peptide identification and quantification for each run independently before combining results. By pre-processing each run separately and constructing a peptide atlas in advance, the method avoids the computationally intensive pairwise alignment and identification transfer operations required by MBR algorithms, reducing complexity from exponential to linear scaling.
Solution Approach 2:
The patent creates a peptide atlas that copies and aggregates peptide information from individual runs. This atlas serves as a reference structure that enables consistent peptide matching across runs without requiring actual alignment operations. The copying approach allows efficient retrieval and comparison of peptide data while avoiding the computational burden of gradient alignment and identification transfer.
3Measurement precision
If DDA acquisition is used for peptide identification, then MS/MS spectra are obtained, but the stochastic selection of precursors leads to high proportions of missing values in quantitative comparisons
Solution Approach 1:
The patent uses a peptide atlas as an intermediary to bridge the gap caused by stochastic precursor selection in DDA. The atlas aggregates peptide information from all runs, enabling consistent identification and quantification even when individual runs have missing values. This mediator structure allows the method to recover and impute missing peptide data while maintaining identification accuracy.
Solution Approach 2:
The patent changes the approach from run-specific precursor selection to atlas-based peptide identification. By shifting from local stochastic selection parameters to global atlas parameters, the method ensures consistent peptide identification across all runs regardless of the stochastic variations in individual DDA experiments, thereby reducing missing values while maintaining measurement precision.
Data Source
Figure 1a
Figure 1b
Figure 1c
AI summary
Method for the identification of analytes in a mixture using a plurality of LC-MS/MS data sets from different runs, including at least an ion mobility (IM) and a retention time (RT) dimension, wherein in a first step, a plurality of analytes, as a subset are confidently identified individually for each run of datasets, and for each run separately, a machine learning model which learns and predicts RT and IM of said analytes, is adapted to run conditions, using transfer learning for said sampled subset of confidently identified analytes; wherein in a second step, analytes, identified in said first step are confidently attributed in a global context over more than one run, and wherein in a third step a global model of said machine learning model for retention time (RT) and ion mobility (IM) prediction is adapted to local conditions of each run using transfer learning, providing an query range in RT and IM dimensions for the signal processing, scoring and validation modules in a final step.