Deconvolution of Mixed Molecular Information in Complex Samples

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying organisms in biological samples are limited by the need for specific detection reagents and cannot provide a comprehensive assay, struggling with mixtures and unknown organisms due to shared peptides and the increasing density of sequenced genomes.

Innovation Solution

A computer-implemented method that generates a data set from peptide or nucleic acid sequences, compares it to a database, and uses a signature function to quantify and identify organisms based on phylogenetic distance, allowing for the identification of any organism present in a sample without the need for isolation or cultivation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If mass spectrometry is used to analyze protein content and match peptides to organisms, then identification capability is improved, but the method fails to comprehensively identify all organisms in mixtures due to shared peptides

Engineering Contradiction:
Improveidentification capabilityVSAvoidcomprehensive identification
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the identification process into multiple computational stages: (1) peptide spectrum matching to generate candidate organisms, (2) construction of a contribution matrix decomposing total spectra into organism-specific contributions, (3) iterative deconvolution to separate shared peptides from unique peptides, and (4) quantification based on deconvolved contributions. This segmentation allows the method to handle mixture complexity systematically.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimensional approach by constructing a contribution matrix that adds a quantitative dimension to traditional binary matching. Instead of simply determining presence/absence, the method quantifies the contribution of each organism to the total spectral signal, enabling discrimination of mixtures through mathematical deconvolution of overlapping signals across multiple dimensions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the number of sequenced genomes increases to improve database coverage, then more organisms can be identified, but the number of shared peptides increases making mixture analysis more difficult

Engineering Contradiction:
Improvedatabase coverageVSAvoidmixture analysis complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements feedback through iterative deconvolution where the contribution matrix is continuously refined. Initial estimates of organism contributions are made, then the model predicts expected spectral contributions, compares with actual observed spectra, and adjusts contributions to minimize residuals. This feedback loop progressively separates shared peptides from unique signatures even as database size increases.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The method dynamically adjusts parameters during analysis including contribution weights, phylogenetic distance thresholds, and deconvolution iteration counts. By changing these parameters adaptively based on database size and mixture complexity, the method maintains analytical effectiveness despite increasing genomic data density.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If specific detection reagents are used to identify target organisms, then detection specificity is improved, but the method cannot provide comprehensive assay of all organisms present

Engineering Contradiction:
Improvedetection specificityVSAvoidcomprehensive assay capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal identification system that simultaneously performs multiple functions: (1) detects known organisms through database matching, (2) identifies unknown organisms through de novo sequencing, (3) quantifies mixture components through deconvolution, and (4) provides phylogenetic classification. This multi-functional approach replaces the need for multiple specific reagents with a single comprehensive mass spectrometry-based platform.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Instead of using physical specific reagents for each organism, the patent creates virtual copies of organism-specific peptide signatures through computational matching against virtual peptide sequences derived from genomic databases. This digital replication allows comprehensive coverage without requiring physical reagents for every possible target.

Inventive Principle:
Principle #26Copying

4Productivity

If clustering methods are used to assign peptides to organisms, then identification speed is improved, but accuracy decreases due to numerous shared peptides between organisms

Engineering Contradiction:
Improveidentification speedVSAvoidassignment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces a contribution matrix as an intermediary between raw peptide spectra and organism identification. This matrix serves as a mediator that decomposes the complex many-to-many mapping between peptides and organisms into a quantitative system where each peptide's contribution to each organism can be calculated and separated, improving assignment accuracy without sacrificing speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3030997B1Method of deconvolution of mixed molecular information in a complex sample to identify organism(s)
Publication Date: 2023.03.29 BERTIN TECHNOLOGIES
  • EP3030997B1 patent drawingFigure 1
  • EP3030997B1 patent drawingFigure 2
  • EP3030997B1 patent drawingFigure 3

AI summary

The present invention relates to methods to determine the identity of one or more organisms present in a sample (if these are already reported in a taxonomic database) or the identity of the closest related organism reported in a taxonomic database. The present invention does this by comparing a data set acquired by analysing at least one component of the biological sample to a database, so as to match each component of the analysed content of the sample to one or more taxon(s) and then collating the phylogenetic distance between each taxa and the taxon with the highest number of matches in the data set. A deconvolution function is then generated for the taxon with the highest number of matches, based on a correlation curve between the number of matches per taxon (Y axis) and said phylogenetic distance (X axis), the outcome of this function providing the identity of the organism or the closest known organism to it.