Protein Confidence Recalculation via Iterative Peptide Assignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Protein confidence values in proteomic analysis may be inaccurate due to the initial peptide confidence values not considering sample-specific information, leading to high false positive results, especially for singleton peptides.

Innovation Solution

A computer-implemented system that recalculates protein confidence values by updating peptide confidence values based on sample-specific information and hidden correlations between peptides, using a processor connected to a mass spectrometer and protein database, iteratively assigning peptides to proteins and recalculating confidence values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If initial peptide confidence values are used without considering sample-specific information, then the calculation process is simple and fast, but the protein confidence values become inaccurate with high false positive results

Engineering Contradiction:
Improveprotein confidence value accuracyVSAvoidconfidence calculation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-calculating and storing peptide confidence values in a database before actual protein identification. These pre-computed values consider various factors and can be quickly retrieved during analysis, avoiding the need to recalculate everything from scratch while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary confidence value calculation system that mediates between raw peptide data and final protein identification. This intermediary layer computes peptide confidence values using multiple factors (peptide mass accuracy, sequence coverage, database search scores) and uses these as inputs for protein confidence calculation, thereby improving accuracy without directly complicating the final protein identification step.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If peptide confidence values are updated based on sample-specific information and hidden correlations, then the accuracy of protein confidence values improves, but the computational time and processing complexity increase

Engineering Contradiction:
Improvepeptide confidence value accuracyVSAvoidrecalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary calculations of peptide confidence values and stores them in a database before actual protein identification experiments. When new samples are analyzed, these pre-computed values can be retrieved and adjusted with minimal recalculation, significantly reducing processing time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by updating peptide confidence values selectively based on sample-specific information. Instead of recalculating all peptide confidence values from scratch, the system identifies which peptides need updating based on their relevance to the current sample and updates only those, reducing overall computational burden.

Inventive Principle:
Principle #3Local quality

3Reliability

If iterative peptide assignment and confidence recalculation is performed, then false positive results are reduced, but the computational complexity and processing steps increase

Engineering Contradiction:
Improveprotein identification reliabilityVSAvoidanalysis process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements a dynamic iterative process where peptide assignments to proteins and confidence value calculations are performed in alternating steps. In each iteration, peptides are assigned to proteins based on current confidence values, then confidence values are recalculated based on new assignments, and this process repeats until convergence or a maximum number of iterations is reached. This dynamic approach improves reliability by progressively refining results.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms where the results of each iteration are used to inform the next iteration. Specifically, updated protein confidence values from one iteration feed into the peptide assignment process of the next iteration, and updated peptide assignments feed back into confidence value recalculations. This feedback loop allows the system to progressively eliminate false positives and converge on accurate results.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11875880B2Systems and methods for calculating protein confidence values
Publication Date: 2024.01.16 DH TECH DEVMENT PTE
  • US11875880B2 patent drawing
  • US11875880B2 patent drawing
  • US11875880B2 patent drawing

AI summary

A method is disclosed for calculating and recalculating protein confidence values by recalculating peptide confidence values in proteomic analysis. A plurality of scans are performed of a sample that is proteolytically digested into surrogate peptide analytes. A plurality of spectra are obtained, and a plurality of peptides are identified. A protein database is searched for proteins matching peptides from the plurality of peptides. Peptide confidence values are determined for the set of peptides. A protein confidence value is calculated for each protein in the set of proteins. A protein from the set of proteins with a largest protein confidence value is selected, the largest protein confidence value for the protein is saved, the protein from the set of proteins is removed, and peptides corresponding to the protein is removed from the set of peptides. The protein confidence value is recalculated for each protein in the set of proteins.