Mass Spectrometry Upsampling via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The sampling rate requirements in mass spectrometry, particularly in GC-MS and LC-MS experiments, limit throughput and data quality due to the need for high sampling rates to avoid aliasing artifacts and ensure accurate representation of analyte concentrations, which can be insufficient for non-Gaussian elution peaks and varying peak widths.
Innovation Solution
A system and method that utilize a trained machine learning model to upsample mass chromatogram data from a low sampling rate to a higher sampling rate, allowing for increased throughput and improved data quality by generating a higher sampling rate representation of mass spectra, enabling more target analytes to be analyzed within a given time frame and using smaller isolation widths in DIA experiments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a high sampling rate is used to avoid aliasing artifacts and ensure accurate representation of analyte concentrations, then measurement precision is improved, but productivity deteriorates due to reduced throughput
Solution Approach 1:
The system performs preliminary action by acquiring data at a lower sampling rate before analysis, then applies a machine learning model to reconstruct the high sampling rate signal retrospectively. This allows the system to avoid the real-time computational burden of high sampling rate acquisition while still achieving accurate analyte concentration measurements through post-processing reconstruction of the elution peaks.
2Productivity
If a low sampling rate is used to increase throughput, then productivity is improved, but measurement precision deteriorates due to aliasing artifacts and inaccurate peak representation
Solution Approach 1:
The machine learning model serves as an intermediary between the low sampling rate acquired data and the desired high sampling rate representation. The model learns the relationship between sparse samples and continuous elution peaks from training data, then uses this learned relationship to reconstruct accurate high-resolution chromatograms from the low sampling rate input, effectively mediating the contradiction between throughput and precision.
3Measurement precision
If a high sampling rate is used to capture non-Gaussian elution peaks and varying peak widths accurately, then measurement precision is improved, but loss of time increases due to longer analysis duration
Solution Approach 1:
The system performs preliminary action by acquiring data at a lower sampling rate that requires less time, then applies a machine learning model to reconstruct the high sampling rate signal retrospectively. This allows the system to avoid the real-time computational burden of high sampling rate acquisition while still achieving accurate analyte concentration measurements through post-processing reconstruction of the elution peaks.
4Loss of time
If a low sampling rate is used during acquisition, then loss of time is reduced, but measurement precision deteriorates unless corrected by post-processing
Solution Approach 1:
The machine learning model serves as an intermediary between the low sampling rate acquired data and the desired high sampling rate representation. The model learns the relationship between sparse samples and continuous elution peaks from training data, then uses this learned relationship to reconstruct accurate high-resolution chromatograms from the low sampling rate input, effectively mediating the contradiction between throughput and precision.
Data Source
AI summary
A method of performing mass spectrometry includes obtaining, based on a series of mass spectra acquired over time with a first sampling rate as analytes elute from a separation system during an experiment, a first mass chromatogram dataset. The first mass chromatogram dataset represents a detected intensity of ions derived from the analytes and having a selected m/z as a function of time over a time period. The method further includes generating, based on the first mass chromatogram dataset and an upsampling model trained to upsample mass chromatogram data, a second mass chromatogram dataset representing an estimated intensity of the ions as a function of time over the time period. The second mass chromatogram dataset has a second sampling rate that is greater than the first sampling rate.


