Deep Learning De Novo Peptide Sequencing from DIA Mass Spectrometry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

De novo peptide sequencing from mass spectrometry data acquired by data-independent acquisition is challenging due to noise and ambiguity, limiting the accuracy and practical use of mass spectrometry data, especially in interpreting highly multiplexed spectra where links between precursor and fragment ions are unknown.

Innovation Solution

A deep learning framework utilizing neural networks is applied to learn the 3D shapes of fragment ions along m/z and retention time dimensions, correlate precursor and fragment ions, and identify peptide sequence patterns, combining recurrent and beam-search mechanisms to improve sequencing accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data-independent acquisition is used to acquire mass spectrometry data, then productivity is improved by acquiring multiple precursor ions simultaneously, but measurement precision deteriorates due to noise and ambiguity in highly multiplexed spectra

Engineering Contradiction:
Improvethroughput of peptide sequencingVSAvoidaccuracy of peptide sequencing
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies deep learning neural networks that process mass spectrometry data in multiple dimensions simultaneously - using m/z ratio, retention time, and intensity as separate dimensional features. This multi-dimensional approach allows the system to resolve ambiguous fragment ion assignments in DIA data by considering temporal and intensity patterns alongside mass information, thereby maintaining high throughput while improving sequencing accuracy

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces deep learning neural networks as an intermediary computational layer between raw mass spectrometry data and peptide sequence identification. The neural network learns complex patterns and relationships in the data, serving as a mediator that translates noisy, multiplexed DIA spectra into reliable peptide sequence information, effectively resolving the contradiction between high-throughput acquisition and accurate interpretation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If deep learning framework is applied to learn fragment ion patterns and correlate precursor-fragment ions, then measurement precision is improved by reducing noise and ambiguity, but device complexity increases due to computational requirements

Engineering Contradiction:
Improveaccuracy of peptide sequencingVSAvoidcomputational complexity of sequencing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs deep learning neural networks that are trained in advance on large datasets of mass spectrometry spectra. This preliminary training phase allows the network to learn complex patterns of fragment ion relationships and noise characteristics beforehand. During actual peptide sequencing, the pre-trained network can rapidly process new data with high accuracy without requiring complex real-time computational resources, thus improving precision while managing system complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11694769B2Systems and methods for de novo peptide sequencing from data-independent acquisition using deep learning
Publication Date: 2023.07.04 BIOINFORMATICS SOLUTIONS
  • US11694769B2 patent drawing
  • US11694769B2 patent drawing
  • US11694769B2 patent drawing

AI summary

The present systems and methods introduce deep learning to de novo peptide sequencing from tandem mass spectrometry data, and in particular mass spectrometry data obtained by data-independent acquisition. The systems and methods achieve improvements in sequencing accuracy over existing systems and methods and enables complete assembly of novel protein sequences without assisting databases. To sequence peptides from mass spectrometry data obtained by data-independent acquisition, precursor profiles representing intensities of one or more precursor ion signals associated with a precursor retention time and fragment ion spectra representing signals from fragment ions and fragment retention times are fed into a neural network.