Peptide Identification via Dynamic Reconstruction Counting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for peptide identification in tandem mass spectrometry are hindered by slow processing times and inaccuracies, particularly when dealing with large datasets and unknown proteomes, due to the lack of a concrete theoretical probability model and reliance on decoy databases, which are time-consuming and impractical.

Innovation Solution

The development of a mass-spectrometry-generating function (MS-GF) algorithm that computes the generating function of a spectrum to determine statistically significant peptide reconstructions without relying on decoy databases, allowing for efficient identification of peptides by scoring matches and optimizing the number of candidate peptides considered for analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If database search methods (SEQUEST, MASCOT) are used to match spectra against protein databases, then peptide identification can be performed, but processing time becomes excessively long for large datasets

Engineering Contradiction:
Improvepeptide identification accuracyVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the peptide identification process into distinct phases: (1) generating candidate peptide reconstructions from spectrum data, (2) scoring each candidate using a probabilistic model, and (3) selecting top candidates for database matching. This segmentation allows the system to quickly filter out unlikely candidates before performing computationally intensive database searches, thereby improving processing speed while maintaining identification accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by generating and scoring candidate peptide reconstructions before actual database searching. The probabilistic scoring model pre-evaluates candidates based on spectrum matching quality, preparing a refined list of promising candidates that are then subjected to database verification. This preliminary filtering reduces the burden on subsequent database search operations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the number of peptide reconstructions is increased to improve identification accuracy, then the likelihood of finding the correct peptide increases, but computational time and resources increase significantly

Engineering Contradiction:
Improvepeptide identification accuracyVSAvoidcomputational time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by generating a limited but sufficient number of peptide reconstructions rather than exhaustively searching all possible sequences. The system generates candidate reconstructions up to a configurable limit and uses probabilistic scoring to identify the most promising ones, avoiding the need to evaluate every possible peptide sequence while still maintaining high identification accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes parameters dynamically by adjusting the number of reconstructions generated based on spectrum quality metrics. For high-quality spectra with clear fragment patterns, the system generates fewer candidates since the correct peptide is more likely to be among top scorers. For lower quality spectra, it generates more candidates to ensure the correct peptide is not missed, optimizing computational resources based on actual data quality.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If de novo peptide sequencing is performed without database reference, then analysis can proceed for unknown proteomes, but accuracy decreases compared to database search methods

Engineering Contradiction:
Improvecapability to analyze unknown proteomesVSAvoidpeptide sequencing accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a probabilistic scoring model as an intermediary between raw spectrum data and final peptide identification. This scoring model evaluates candidate reconstructions based on multiple factors including fragment ion matching, mass accuracy, and spectral quality metrics. The intermediary scorer provides a quantitative assessment that improves de novo sequencing accuracy by objectively ranking candidates even in the absence of database references.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements dynamic adjustment of sequencing strategies based on spectrum characteristics. The system adapts the number of candidate reconstructions generated, the scoring thresholds applied, and the verification steps performed based on real-time assessment of spectrum quality, ion types present, and data completeness. This dynamic approach optimizes accuracy for each individual spectrum while maintaining versatility across different proteome types.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8639447B2Method for identifying peptides using tandem mass spectra by dynamically determining the number of peptide reconstructions required
Publication Date: 2014.01.28 RGT UNIV OF CALIFORNIA
  • US8639447B2 patent drawing
  • US8639447B2 patent drawing
  • US8639447B2 patent drawing

AI summary

A method for identifying peptides using tandem mass spectrometry takes the spectrum for a peptide to be analyzed and uses a scoring function to score a match between the spectrum and each candidate peptide in a peptide database. The scoring function has a value corresponding to a number of fragment peaks in the spectrum that match fragment peaks in a spectrum of the candidate peptide. Using the match scores, a generating function of the spectrum is computed to determine the number of peptide reconstructions at each value of the scoring function. The generating function is then used to determine the number of candidate peptides for each match score and the probability of a peptide having a given match score to the spectrum. A spectral probability can be determined by calculating the total probability of all peptides with scores equal to or larger than the given match score.