Mass Spectra Tokenization for Chemical Structure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mass spectrometry techniques face challenges in determining the chemical structure of molecules from mass spectra due to high chemical structure diversity and noisy data, making it difficult to confidently identify molecules, especially in complex samples.

Innovation Solution

A computational metabolomics platform utilizing trained bidirectional transformer-based machine-learning models to predict and generate chemical structures and properties from known mass spectrometry data, including mass-to-charge (m/z) values and precursor mass.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If mass spectrometry is used to determine chemical structure, then molecular identification is achieved, but accuracy decreases due to high chemical structure diversity and noisy data

Engineering Contradiction:
Improvemolecular identification accuracyVSAvoididentification confidence
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces an intermediate computational processing layer between mass spectrometry data acquisition and molecular identification. This intermediary system uses machine learning models to process raw MS data, filter noise, and generate predicted chemical structures, thereby improving both measurement precision and identification confidence without requiring direct interpretation of noisy spectral data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional manual or rule-based chemical structure determination methods with automated machine learning-based prediction systems. This substitution enables more accurate and reliable molecular identification by using computational algorithms to interpret mass spectrometry data, overcoming the limitations of human analysis and traditional spectral interpretation methods

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If traditional mass spectrometry analysis is used, then molecular data is obtained, but time consumption increases due to the need to isolate and analyze each molecule individually

Engineering Contradiction:
Improvemolecular characterization throughputVSAvoidtime to identify molecules
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent creates a universal computational platform that can simultaneously analyze and predict structures for multiple different molecules from mass spectrometry data. This multi-functional system processes complex mixtures in parallel, eliminating the need for sequential isolation and analysis of individual compounds, thereby dramatically increasing productivity while reducing time loss

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary computational processing of mass spectrometry data to predict chemical structures before detailed analysis is required. By pre-processing the data and generating structure predictions upfront, the system enables faster subsequent analysis and reduces the overall time needed for molecular characterization of complex samples

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250189496A1Predicting chemical structure and properties based on mass spectra
Publication Date: 2025.06.12 ENVEDA THERAPEUTICS INC
  • US20250189496A1 patent drawing
  • US20250189496A1 patent drawing
  • US20250189496A1 patent drawing

AI summary

Methods for identifying a chemical structure of a compound based on mass spectrometry (MS) data using one or more computing devices are disclosed. The methods include receiving mass spectrometry (MS) data that includes a plurality of mass-to-charge values associated with fragments obtained from mass spectrometry performed on the compound, inputting the plurality of mass-to-charge values into a tokenizer trained to generate a plurality of tokens based on the plurality of mass-to-charge values, and determining one or more chemical structures of the compound based at least in part on the plurality of tokens.