Reinforcement Learning for DIA Spectral Library Expansion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning prediction methods for extracting information from data-independent acquisition (DIA) mass spectrometry data lead to high false-negative rates and increased computational time due to large, crowded mass spaces and extensive libraries, which complicates the identification of compounds not present in spectral libraries.

Innovation Solution

A reinforcement learning algorithm is employed to search for related compounds and apply deep learning prediction algorithms to predict additional product ion spectra, iteratively refining the library and improving prediction models, thereby reducing false-negative rates and computational time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If deep learning prediction methods are used to extract information from DIA data, then proteome coverage is improved, but false-negative rate increases and computational time increases

Engineering Contradiction:
Improveproteome coverageVSAvoidfalse-negative rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent performs preliminary actions by first identifying compounds using traditional spectral library matching before applying deep learning prediction methods. This staged approach allows the system to benefit from the expanded proteome coverage of deep learning while maintaining the reliability of traditional methods for confirmed identifications, thereby reducing false-negative rates without sacrificing coverage

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary verification step where predicted compound identifications are cross-checked against the original DIA data and spectral libraries. This intermediary layer acts as a filter to validate deep learning predictions, reducing false-negative rates while preserving the expanded proteome coverage capability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If deep learning prediction methods are used to extract information from DIA data, then proteome coverage is improved, but computational time increases

Engineering Contradiction:
Improveproteome coverageVSAvoidcomputational time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies partial action by using deep learning prediction methods selectively rather than universally. It focuses computational resources on predicting spectra for compounds that are most likely to be present based on preliminary filtering, thereby achieving improved proteome coverage without the excessive computational time required for exhaustive prediction of all possible compounds

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent segments the compound identification process into distinct phases: traditional spectral library matching, deep learning prediction for candidate compounds, and verification. This segmentation allows computational time to be distributed efficiently across different methods, achieving high proteome coverage without the prohibitive computational cost of applying deep learning to all compounds simultaneously

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If spectral libraries are expanded to include more compounds, then compound identification capability is improved, but false discovery rate increases

Engineering Contradiction:
Improvecompound identification capabilityVSAvoidfalse discovery rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the results from deep learning predictions are fed back into the identification process for verification and refinement. This feedback loop allows the system to adjust and validate predictions against the expanded spectral library, ensuring that increased compound identification capability does not compromise the false discovery rate through unverified matches

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The approach effectively enhances the identification of compounds by iteratively expanding the spectral library with related proteins or compounds, reducing false discovery rates and improving proteome coverage without significantly increasing computational time.

Implementation Method 1

an ion source device that ionizes one or more compounds of a sample, producing an ion beam

Methodology Applied
Scientific EffectIonization: Ionisation

Implementation Method 2

for each window of the n windows, fragments precursor ions of each window and mass analyzes resulting product ions from the fragmentation

Methodology Applied
Scientific EffectCollision-induced dissociation:

Data Source

PatentUS20240428893A1Methods for enhancing complete data extraction of DIA data
Publication Date: 2024.12.26 DH TECH DEVMENT PTE
  • US20240428893A1 patent drawing
  • US20240428893A1 patent drawing
  • US20240428893A1 patent drawing

AI summary

The n spectra of a DIA method are compared to a library of product ion spectra to identify an initial i compounds corresponding to l spectra. A reinforcement learning algorithm (RLA) is performed. (a) An agent of the RLA performs an action At that includes searching one or more compound databases for compounds related to the i compounds, producing j related compounds, and applying one or more deep learning prediction algorithms to predict k spectra for the i+j compounds. (b) An environment of the RLA compares the k spectra to the n spectra, producing a state, St, in which i+j compounds produce m matching compounds and a reward, Rt, for the agent if m>i. (c) If the Rt is produced, the i compounds are set to the m compounds and the l spectra are set to the k spectra, and steps (a)-(c) are repeated.