Peptide Source Assignment via FDR-Based Database Prioritization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In bioinformatics, assigning a putative source to de novo peptide sequences without experimental confirmation is challenging due to the reliability and timeliness of data sources, and existing methods struggle to accurately identify the origin of peptides, especially those that may be cis- or trans-spliced.

Innovation Solution

A method that orders peptide source search steps based on increasing random hit rates, using databases like the expanded human proteome and non-endogenous proteome, and applies simulated random queries to determine false discovery rates, facilitating a query support data structure for improved peptide source assignment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple databases are searched to identify peptide sources, then the comprehensiveness of source identification is improved, but the complexity of the search process increases

Engineering Contradiction:
Improvecompleteness of source identificationVSAvoidsearch process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The search process is segmented into multiple independent database searches (human proteome, non-endogenous proteome, etc.), each handled separately with its own random hit rate calculation. This segmentation allows comprehensive coverage while managing complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Random peptide sequences are generated in advance to calculate random hit rates for each database before actual peptide identification. This preliminary action enables systematic evaluation of each database's reliability without requiring experimental confirmation, establishing a framework for confident source assignment.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If search results are evaluated without considering false discovery rates, then the speed of analysis is improved, but the reliability of source assignment decreases

Engineering Contradiction:
Improveanalysis speedVSAvoidsource assignment reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

Random hit rates are calculated for each database using simulated random peptide sequences, providing feedback on the reliability of each search source. This feedback mechanism enables systematic evaluation of false discovery rates, allowing confident source assignment while maintaining efficient analysis through automated FDR-based filtering.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If experimental confirmation is required for source assignment, then the accuracy of peptide origin identification is improved, but the time required for analysis increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Random peptide sequences are generated in advance to calculate random hit rates for each database before actual peptide identification. This preliminary action establishes reliable FDR benchmarks that enable accurate source assignment without requiring time-consuming experimental confirmation, as the framework is pre-configured with reliability metrics.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The mechanical process of experimental confirmation is replaced with computational FDR-based evaluation. By substituting physical experimentation with algorithmic assessment of random hit rates, the system achieves comparable or superior identification accuracy while dramatically reducing analysis time through in silico evaluation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240153587A1Workflow to assign putative source to de novo peptide sequence
Publication Date: 2024.05.09 REGENERON PHARMACEUTICALS INC
  • US20240153587A1 patent drawing
  • US20240153587A1 patent drawing
  • US20240153587A1 patent drawing

AI summary

Methods and systems are described for optimizing search results through querying a plurality of databases according to false discovery, random hit rates are presented herein. Methods and systems adapted to assigning a putative source to a de novo peptide sequence and/or creating a workflow for performing said assignment are presented herein.