Tandem Mass Spectrometry Clustering for Unknown Compound Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for discovering novel unnamed compounds in complex mixtures are time-consuming, prone to false positives, and lack automation, leading to a stagnation in reporting new compounds.
Innovation Solution
A fully automated method utilizing MS scans, which discards known features and focuses on forming component clusters based on precursor ion mass-to-charge ratio and retention index, to identify candidate unknown compounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a graphical interactive tool is used to assist in discovering unnamed compounds, then the scientist receives indispensable assistance, but the process becomes very time-consuming and discovers many false positives
Solution Approach 1:
The patent replaces the manual graphical interactive tool with an automated computational system that uses machine learning algorithms and data processing pipelines to analyze mass spectrometry data, eliminating the need for manual scientist intervention while improving both speed and accuracy
Solution Approach 2:
The system performs self-service by automatically filtering false positives, clustering compound features, and identifying unnamed compounds without requiring continuous scientist input, allowing the discovery process to run autonomously and efficiently
2Quantity of substance
If existing chromatography/mass spectrometry methods are used to detect unnamed compounds, then compounds are detectable, but they are not reported because they do not match any existing known compound in a database
Solution Approach 1:
The patent extracts and isolates features from mass spectrometry data that do not match known compounds in databases, separating them from the standard identification workflow and creating a dedicated pathway for novel compound discovery and reporting
Solution Approach 2:
The system performs preliminary clustering and grouping of detected features before database matching, organizing data in advance to identify patterns and relationships that indicate novel compounds, enabling proactive discovery rather than reactive identification
3Measurement precision
If manual processing and characterization of features is performed to enrich for true novel unnamed compounds, then accuracy may improve, but the process becomes time consuming and error-prone
Solution Approach 1:
The patent replaces manual processing and characterization steps with automated computational algorithms that perform feature clustering, filtering, and validation, eliminating human error while maintaining high precision and enabling scalable processing of large datasets
Solution Approach 2:
The system implements feedback loops where computational results are automatically evaluated and refined, with algorithms learning from previous analyses to improve precision in identifying novel compounds while maintaining high throughput through iterative optimization
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This method significantly reduces analysis time, increases accuracy, and efficiently filters out noise, enabling the analysis of large sample sets in a fraction of the time required by existing software.
Implementation Method 1
data for each sample being obtained from a component separation and tandem mass spectrometer system
Implementation Method 2
a separation portion, including a liquid chromatograph, a gas chromatograph, a supercritical fluid chromatograph, or a capillary electrophoresis analyzer
Data Source
AI summary
A data analysis method for a component separation/tandem mass spectrometer system, including first and second MS steps which data therefrom includes respective sample components (MS1-SC, MS2-SC), includes analyzing per sample a data set for MS2-SC to determine mass-to-charge ratio and retention index (m/z-RI) for each MS2-SC. m/z-RI for each MS2-SC is compared to a known compound library and matching MS2-SC removed from the data set, the remaining MS2-SC being candidate MS2-SC. Clusters are formed across the candidate MS2-SC, each having m/z-RI within respective ranges per cluster. For each sample within each cluster, MS1-SC within the cluster ranges are retrieved. For each cluster, at most one consensus MS1-SC represents each sample, with corresponding consensus MS2-SC and m/z-RI, and designated as a molecular ion or derivative thereof. Clusters are grouped by consensus RI and candidate clusters from each group are selected and correlated by consensus parameters with an unknown compound.


