Peptide Sequencing from MS/MS Spectra with Candidate Graph Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing peptide sequencing methods struggle with efficiently sequencing data-encoded peptides using mass spectrometry, particularly due to the lack of suitable training sequences and the use of unnatural amino acids, which complicates the application of existing algorithms.

Innovation Solution

The proposed method involves a computer-implemented two-stage sequencing approach and a highest-intensity-tag based sequencing method, which utilize graph theory models and error-correction codes to infer partial sequences and refine peptide sequences from experimental spectra.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing peptide sequencing methods are used, then sequencing can be performed, but the speed and accuracy are insufficient due to the lack of suitable training sequences and use of unnatural amino acids

Engineering Contradiction:
Improvesequencing accuracyVSAvoidsequencing method complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the sequencing problem into two distinct stages: (1) generating candidate sequences from spectral data, and (2) verifying and refining these candidates using graph theory models and error-correction codes. This segmentation allows each stage to be optimized independently, improving overall accuracy without proportionally increasing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-processing spectral data to generate a reduced set of candidate sequences before full verification. Error-correction codes are applied in advance to encode validation information, enabling faster and more accurate sequencing by eliminating unlikely candidates early in the process.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If existing algorithms are applied to data-encoded peptides, then sequencing is possible, but the number of candidate sequences remains large reducing efficiency

Engineering Contradiction:
Improvesequencing speedVSAvoidcandidate sequence filtering
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent employs feedback mechanisms where graph theory models continuously evaluate candidate sequences against spectral data, providing feedback that guides the refinement process. Error-correction codes provide additional feedback layers, enabling the system to quickly identify and eliminate incorrect candidates, thereby increasing sequencing speed without losing valid sequences.

Inventive Principle:
Principle #23Feedback

3Reliability

If traditional mass spectrometry sequencing is used, then peptide sequences can be obtained, but reliability is reduced due to inability to effectively reject unlikely candidates

Engineering Contradiction:
Improvesequencing reliabilityVSAvoidverification process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-encoding error-detection and correction capabilities into the sequencing workflow. Graph theory models are prepared in advance with validation rules, and error-correction codes are applied to candidate sequences before full verification, enabling reliable rejection of unlikely candidates without excessive complexity during the main sequencing process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4261826B1Sequencing data-encoded peptides from tandem mass spectra
Publication Date: 2025.03.05 THE HONG KONG POLYTECHNIC UNIV
  • EP4261826B1 patent drawingFigure 1
  • EP4261826B1 patent drawingFigure 2
  • EP4261826B1 patent drawingFigure 2

AI summary

Peptide sequencing is important in decoding data stored in a data-encoded peptide. Tandem mass spectrometry (MS/MS) is particularly useful for peptide sequencing. In a computer-implemented method for sequencing the data-encoded peptide from an experimental spectrum, raw data of the experimental spectrum are first preprocessed to remove uninterpretable peaks to yield preprocessed data. A first set of one or more candidate sequences contending for a peptide sequence of the peptide is identified from a spectrum graph. The spectrum graph is formed according to the preprocessed data rather than the raw data for generating a fewer number of candidate sequences to thereby reduce a time cost in sequencing. The first candidate-sequence set is then processed to estimate the peptide sequence to thereby obtain a set of peptide-sequence estimate(s). Each estimate is verified whether it is invalid. The set of peptide-sequence estimate(s) is purged to remove any invalid estimate.