Metabolite Identification Algorithm Using Isotopologue Probability Caching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods in ultra-high resolution mass spectrometry struggle to identify and interpret metabolites due to over 50% of detected metabolites remaining unidentified, with low confidence in identifications without standard comparisons, especially in stable isotope-resolved metabolomics experiments where comparing to millions of isotopically-different chemical standards is not feasible.

Innovation Solution

A novel algorithm and system that detects metabolites by calculating natural abundance probabilities, building and sorting large caches of molecular fragments, characterizing peaks, and statistically filtering results to identify specific elemental molecular formulas from cliques of related isotopologue peaks, enabling the identification of metabolites at the level of structural isomers and overcoming the limitations of existing methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current database-matching methods are used to identify metabolites, then identification speed is improved, but identification accuracy deteriorates because over 50% of detected metabolites remain unidentified

Engineering Contradiction:
Improveidentification speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the identification parameters from simple mass-to-charge ratio matching to a comprehensive multi-parameter approach including high-resolution mass accuracy (±2 ppm), isotopic pattern matching, fragmentation pattern analysis, and retention time correlation. This enables accurate identification of metabolites that cannot be matched in conventional databases by expanding the parameter space used for identification.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the identification process into multiple independent analytical components: (1) full-scan high-resolution mass spectrometry for accurate mass measurement, (2) targeted MS/MS fragmentation analysis for structural elucidation, (3) isotopic pattern deconvolution for molecular formula determination, and (4) database searching with expanded parameters. This segmentation allows each component to contribute independently to the overall identification accuracy.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If comparison with millions of isotopically-different chemical standards is performed, then identification accuracy is improved, but device complexity and time consumption worsen

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-processing the mass spectrometry data through high-resolution full-scan acquisition before targeted analysis. The accurate mass measurements and isotopic pattern deconvolution are performed upfront to generate a filtered list of candidate metabolites, which then reduces the scope of subsequent database searching. This preliminary filtering action dramatically reduces the effective search space from millions of standards to a manageable subset.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts and isolates the most discriminating features from the complex mass spectrometry data, specifically the high-accuracy mass-to-charge ratios and isotopic abundance patterns, separating these from the full spectral data. By extracting only these critical parameters for database comparison, the system achieves accurate identification without requiring comparison against all isotopically-different standards.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If low-confidence metabolite identifications are accepted, then productivity is improved, but reliability deteriorates

Engineering Contradiction:
Improvenumber of identificationsVSAvoidconfidence in identifications
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies partial action by implementing a tiered identification confidence system. Rather than requiring complete certainty for all metabolites, the system provides high-confidence identifications for those meeting strict criteria (multiple matching parameters) and lower-confidence annotations for others, still delivering productive results across the entire dataset while maintaining reliability for the high-confidence subset.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements feedback through a confidence scoring system that evaluates each metabolite identification based on multiple parameters including mass accuracy deviation, isotopic pattern correlation, fragmentation match quality, and database hit rate. This quantitative feedback mechanism allows the system to objectively assess identification reliability and guide further analysis of lower-confidence candidates.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10607723B2Method and system for identification of metabolites using mass spectra
Publication Date: 2020.03.31 UNIVERSITY OF KENTUCKY RESEARCH FOUNDATION
  • US10607723B2 patent drawing
  • US10607723B2 patent drawing
  • US10607723B2 patent drawing

AI summary

A method and system is provided for mass spectrometry for identification of a specific elemental formula for an unknown compound which includes but is not limited to a metabolite. The method includes calculating a natural abundance probability (NAP) of a given isotopologue for isotopes of non-labelling elements of an unknown compound. Molecular fragments for a subset of isotopes identified using the NAP are created and sorted into a requisite cache data structure to be subsequently searched. Peaks from raw spectrum data from mass spectrometry for an unknown compound. Sample-specific peaks of the unknown compound from various spectral artifacts in ultra-high resolution Fourier transform mass spectra are separated. A set of possible isotope-resolved molecular formula (IMF) are created by iteratively searching the molecular fragment caches and combining with additional isotopes and then statistically filtering the results based on NAP and mass-to-charge (m/2) matching probabilities. An unknown compound is identified and its corresponding elemental molecular formula (EMF) from statistically-significant caches of isotopologues with compatible IMFs.