Chromatographic Alignment via Machine Learning Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing chromatographic alignment methods for LC-MS and GC-MS systems are computationally intensive, consume vast resources, and are not scalable for analyzing large numbers of samples, making them inefficient for comparative analytical methods in biology and biomarker discovery.
Innovation Solution
A method and system that utilize a machine learning model, specifically a trained neural network, to identify and predict retention time offset values for chromatographic features, allowing for efficient alignment of chromatographic profiles by dividing features into subsets based on intensity thresholds and applying peak matching algorithms, thereby reducing processing load and increasing scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional chromatographic alignment methods are used to ensure accurate retention time alignment, then alignment precision is improved, but computational resource consumption increases significantly
Solution Approach 1:
The patent segments the chromatographic features into two subsets: a first subset used for training the machine learning model and a second subset for prediction. This segmentation allows the system to process only a portion of features computationally intensive methods while using the ML model for the remainder, thus reducing overall computational resource consumption while maintaining alignment precision.
Solution Approach 2:
The patent performs preliminary action by training a machine learning model on a first subset of chromatographic features to learn retention time offset patterns. This pre-trained model is then applied to predict retention time offsets for a second subset of features, avoiding the need to apply computationally intensive alignment methods to all features and thereby reducing computational resource consumption.
2Measurement precision
If traditional chromatographic alignment methods are used to achieve accurate retention time alignment, then alignment quality is improved, but processing time increases
Solution Approach 1:
The patent divides chromatographic features into two subsets and applies different processing approaches: traditional alignment methods on the first subset and ML-based prediction on the second subset. This segmentation reduces the total processing time while maintaining alignment quality through the use of the trained ML model for the larger portion of features.
Solution Approach 2:
The patent performs preliminary training of the machine learning model on a training set of chromatographic features. Once trained, the model can quickly predict retention time offsets for new features without requiring repeated computationally intensive alignment calculations, thereby significantly reducing processing time for large numbers of samples.
3Measurement precision
If traditional chromatographic alignment methods are used to ensure accurate retention time alignment, then alignment accuracy is improved, but scalability decreases
Solution Approach 1:
The patent segments the workload into model training on a first subset of features and prediction on a second subset of features. This segmentation enables the system to scale to large numbers of samples because the ML model, once trained, can rapidly process additional features without proportionally increasing computational resource consumption, thus improving scalability while maintaining alignment accuracy.
Solution Approach 2:
The patent performs preliminary training of the machine learning model on a representative training set. This pre-computed model captures the essential alignment patterns and can be applied to numerous additional samples with minimal additional computational cost, enabling the system to scale efficiently to large numbers of samples while maintaining alignment accuracy.
4Productivity
If machine learning models are used to predict retention time offset values for all chromatographic features, then processing speed is improved, but measurement precision may deteriorate
Solution Approach 1:
The patent segments chromatographic features into two subsets: a first subset processed with traditional alignment methods to obtain accurate reference values, and a second subset processed with ML prediction. This segmentation allows the system to use the accurate traditional methods on a training set to teach the ML model, ensuring prediction accuracy while achieving speed improvements on the larger second subset.
Solution Approach 2:
The patent performs preliminary action by training the machine learning model on accurately aligned data from the first subset using traditional methods. This ensures the ML model learns from high-quality reference data, maintaining prediction accuracy while enabling fast processing of the second subset through the pre-trained model.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An exemplary chromatographic alignment system accesses a target file including data representative of a plurality of chromatographic features detected from a first sample and a reference file including data representative of a plurality of chromatographic features detected from a second sample. The system identifies, based on the target and reference files, a distinct retention time offset value for each chromatographic feature included in a first subset of the plurality of chromatographic features detected from the first sample. The system determines, based on the identified distinct retention time offset values for the chromatographic features included in the first subset and on a machine learning model, a distinct predicted retention time offset value for each chromatographic feature included in a second subset of the plurality of chromatographic features detected from the first sample. The system assigns the distinct predicted retention time offset value for each chromatographic feature included in the second subset.