Peptide Scoring via Graph Theory and Sequence Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mass spectrometry methods for protein and peptide identification face challenges in accurately identifying sequences due to variations from recognized protein states, such as post-translational modifications and single nucleotide polymorphisms, which can lead to false or no matches, decreasing confidence in protein identification.

Innovation Solution

The method employs a computer-based system that uses graph theory and sequence tagging to identify and score peptide sequences, considering modifications and variations, by generating hypotheses and scoring them based on experimental data, thereby accounting for differences in precursor masses and fragmentation patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional mass spectrometry methods are used for peptide identification, then the process is simple and fast, but accuracy decreases when post-translational modifications or single nucleotide polymorphisms are present

Engineering Contradiction:
Improvepeptide identification accuracyVSAvoidsearch space complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the peptide search space by dividing the protein sequence into multiple peptide hypotheses based on different modification states. Each segment represents a potential peptide variant, allowing systematic exploration of modification possibilities without overwhelming the entire search space at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by evaluating and scoring individual peptide hypotheses based on their match to experimental mass spectrometry data. Peptide regions with higher scores (better matches) are prioritized, allowing the system to focus computational resources on the most promising candidates while accounting for local variations caused by modifications.

Inventive Principle:
Principle #3Local quality

2Reliability

If all possible peptide variations are considered to account for modifications, then identification accuracy improves, but computational time and resources increase significantly

Engineering Contradiction:
Improveprotein identification confidenceVSAvoidcomputational time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-calculating and ranking peptide hypotheses based on theoretical properties before comparing them against experimental data. This preliminary sorting allows the system to evaluate only the most promising peptide variants, reducing the computational burden while maintaining high reliability in identification.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by considering only the most likely peptide variations rather than exhaustively evaluating all possibilities. By using scoring mechanisms to prioritize hypotheses, the system achieves sufficient reliability without the computational cost of complete exhaustiveness.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If peptide hypotheses are strictly matched to experimental spectra, then match precision is high, but false matches occur when modifications alter precursor mass and fragmentation

Engineering Contradiction:
Improvespectra matching precisionVSAvoidpeptide identification reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies parameter changes by allowing flexible matching criteria that can adapt to mass shifts caused by post-translational modifications. Instead of strict fixed-tolerance matching, the system adjusts mass tolerance parameters and fragmentation expectations based on the suspected modifications, maintaining both precision and reliability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses feedback mechanisms where the scoring of peptide hypotheses is continuously refined based on the degree of match between theoretical and experimental spectra. This feedback loop allows the system to identify and correct mismatches caused by modifications, improving overall identification reliability while maintaining precision.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8712695B2Method, system, and computer program product for scoring theoretical peptides
Publication Date: 2014.04.29 MDS INC
  • US8712695B2 patent drawing
  • US8712695B2 patent drawing
  • US8712695B2 patent drawing

AI summary

The present teachings provide for identification of peptides using small sequence tags to focus computational resources on searching regions of a protein database that are the most likely to yield correct identifications. They allow for the incorporation of modifications and in doing so focuses the search to peptides with a precursor mass match. Additionally, probability or relevance factors can be used to determine peptide hypotheses. Various embodiments are presented that search for peptides when a single precursor is selected or when multiple precursors are simultaneously fragmented.