Peptide Sequencing from MS/MS Spectra with Candidate Graph Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing peptide sequencing methods struggle with efficiently sequencing data-encoded peptides using mass spectrometry, particularly due to the lack of suitable training sequences and the use of unnatural amino acids, which complicates the application of existing algorithms.
Innovation Solution
The proposed method involves a computer-implemented two-stage sequencing approach and a highest-intensity-tag based sequencing method, which utilize graph theory models and error-correction codes to infer partial sequences and refine peptide sequences from experimental spectra.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing peptide sequencing methods are used, then sequencing can be performed, but the speed and accuracy are insufficient due to the lack of suitable training sequences and use of unnatural amino acids
Solution Approach 1:
The patent applies segmentation by dividing the sequencing problem into two distinct stages: (1) generating candidate sequences from spectral data, and (2) verifying and refining these candidates using graph theory models and error-correction codes. This segmentation allows each stage to be optimized independently, improving overall accuracy without proportionally increasing complexity.
Solution Approach 2:
The patent implements preliminary action by pre-processing spectral data to generate a reduced set of candidate sequences before full verification. Error-correction codes are applied in advance to encode validation information, enabling faster and more accurate sequencing by eliminating unlikely candidates early in the process.
2Productivity
If existing algorithms are applied to data-encoded peptides, then sequencing is possible, but the number of candidate sequences remains large reducing efficiency
Solution Approach 1:
The patent employs feedback mechanisms where graph theory models continuously evaluate candidate sequences against spectral data, providing feedback that guides the refinement process. Error-correction codes provide additional feedback layers, enabling the system to quickly identify and eliminate incorrect candidates, thereby increasing sequencing speed without losing valid sequences.
3Reliability
If traditional mass spectrometry sequencing is used, then peptide sequences can be obtained, but reliability is reduced due to inability to effectively reject unlikely candidates
Solution Approach 1:
The patent applies preliminary action by pre-encoding error-detection and correction capabilities into the sequencing workflow. Graph theory models are prepared in advance with validation rules, and error-correction codes are applied to candidate sequences before full verification, enabling reliable rejection of unlikely candidates without excessive complexity during the main sequencing process.
Data Source
Figure 1
Figure 2
Figure 2
AI summary
Peptide sequencing is important in decoding data stored in a data-encoded peptide. Tandem mass spectrometry (MS/MS) is particularly useful for peptide sequencing. In a computer-implemented method for sequencing the data-encoded peptide from an experimental spectrum, raw data of the experimental spectrum are first preprocessed to remove uninterpretable peaks to yield preprocessed data. A first set of one or more candidate sequences contending for a peptide sequence of the peptide is identified from a spectrum graph. The spectrum graph is formed according to the preprocessed data rather than the raw data for generating a fewer number of candidate sequences to thereby reduce a time cost in sequencing. The first candidate-sequence set is then processed to estimate the peptide sequence to thereby obtain a set of peptide-sequence estimate(s). Each estimate is verified whether it is invalid. The set of peptide-sequence estimate(s) is purged to remove any invalid estimate.