Cut Vertex Graphs for Large-Molecule Metabolite Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current metabolite identification (MetID) approaches are inadequate for large molecules like therapeutic proteins and peptides, as they face computational complexity due to atom-based algorithms, inability to deconvolute monoisotopic peaks, and lack of consideration for distinct metabolic processes, leading to inefficient identification and visualization of metabolites.

Innovation Solution

A cut vertex method is employed to represent large molecules as minimum cleavable unit graphs, splitting them into components, generating line graphs, and using graph traversal algorithms to identify and display substructures, enabling efficient identification and visualization of metabolites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If atom-based algorithms are used for metabolite identification, then identification accuracy is improved, but computational complexity increases significantly for large biomolecules

Engineering Contradiction:
Improvemetabolite identification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the large biomolecule into smaller substructures or fragments. Instead of analyzing the entire molecule at once using atom-based algorithms, the system divides it into manageable pieces that can be processed independently, reducing computational complexity while maintaining identification accuracy through systematic reconstruction of metabolites from these segments

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If conventional small molecule MetID software is used for large biomolecules, then software availability is improved, but deconvolution accuracy of monoisotopic peaks deteriorates

Engineering Contradiction:
Improvesoftware availabilityVSAvoidmonoisotopic peak deconvolution accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the fundamental parameters and algorithms used for peak deconvolution, adapting them specifically for large biomolecules. Instead of using parameters optimized for small molecules, the system implements modified deconvolution algorithms that account for the unique mass spectral characteristics of large biomolecules, including their broader isotopic distributions and lower signal intensities

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If all substructures of a large molecule are analyzed, then comprehensive metabolite identification is improved, but memory requirements exceed available computer RAM

Engineering Contradiction:
Improvecomprehensive metabolite identificationVSAvoidmemory requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the comprehensive analysis into smaller analytical units based on molecular substructures. By dividing the molecule into fragments and analyzing each segment separately, the system reduces the memory footprint of each analysis step while maintaining comprehensive coverage through systematic aggregation of results from all segments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial analysis by focusing on biologically relevant substructures and metabolic pathways rather than exhaustively analyzing all possible substructures. This selective approach analyzes only the portions of the molecule most likely to produce metabolites, reducing memory requirements while maintaining practical comprehensiveness for drug metabolism studies

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3794599B1Cut vertex method for identifying complex molecule substructures
Publication Date: 2025.08.06 MERCK SHARP & DOHME LLC
  • EP3794599B1 patent drawingFigure 1
  • EP3794599B1 patent drawingFigure 2
  • EP3794599B1 patent drawingFigure 3

AI summary

Embodiments of the present invention avoid the processing problems associated with using conventional computer systems for identifying and characterizing all of the substructures (e.g., metabolites) of large complex molecules by using a defined minimum cleavable unit (MCU) and an MCU graph for a chosen molecule, as well as a "cut vertex" in the MCU graph for the chosen molecule. The system splits the MCU graph of the chosen molecule at the specified cut vertex to produce two separate MCU graph components (i.e., a first MCU subgraph and a second MCU subgraph) of the chosen molecule, and generates and traverses a first line graph component and a second line graph component, respectively, for the two MCU subgraph components with a graph traversing algorithm to generate and store in memory a first database of substructures and molecular weights for the first component, and a second database of substructures and molecular weights for the second line graph component. Subsequently, embodiments of the present invention can perform binary searches on the two databases (or the two subsections of a single database) to identify and produce graphic representations of all of the substructures of the chosen molecule that have molecular weights that match the query molecular weight (or range of query molecular weights), including the substructures of the chosen molecule that straddle (i.e., include) the cut vertex.