Species-Level Microbiome Classification via Machine Learning and MED

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for taxonomic classification of the human microbiome in the aerodigestive tract are limited to genus-level resolution, failing to distinguish medically important pathogens from harmless bacteria, and require high-resolution databases and algorithms for species-level identification.

Innovation Solution

The use of a machine learning classifier coupled with Minimum Entropy Decomposition (MED) and the expanded Human Oral Microbiome Database (eHOMD) for species-level taxonomic classification, leveraging the V1-V3 region of the 16S rRNA gene for enhanced resolution and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If genus-level classification methods are used, then the analysis is simpler and faster, but the resolution is insufficient to distinguish pathogens from harmless bacteria

Engineering Contradiction:
Improvetaxonomic resolutionVSAvoidclassification system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the taxonomic classification process into multiple hierarchical levels (phylum, class, order, family, genus, species) and applies different computational methods at each level. Species-level classification uses machine learning classifiers trained on specific genomic markers, while higher levels use simpler compositional methods, resolving the contradiction by dividing the complex task into manageable segments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters used for classification by transitioning from generic genus-level markers to specific species-level genomic markers and machine learning models. This parameter change enables species-level resolution while maintaining computational efficiency through optimized training sets and feature selection

Inventive Principle:
Principle #35Parameter changes

2Reliability

If high-resolution species-level databases are used, then classification accuracy improves, but data processing complexity and computational requirements increase

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and trains machine learning classifiers on specific, high-value genomic markers and features rather than processing entire genomes. This extraction approach maintains high classification accuracy by focusing on discriminative features while reducing data processing complexity by excluding redundant information

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary action by pre-training machine learning classifiers on curated training sets before actual classification. This preliminary training incorporates complex pattern recognition and feature weighting, allowing rapid and accurate species-level classification during execution without requiring complex real-time processing

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If machine learning classifiers with training sets are used, then species-level identification accuracy increases to over 90%, but the system complexity and training data requirements increase

Engineering Contradiction:
Improvespecies-level identification accuracyVSAvoidmachine learning system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent optimizes machine learning parameters including feature selection, class weighting, and threshold settings to achieve over 90% accuracy. By carefully tuning these parameters and using curated training sets, the system achieves high precision while managing complexity through parameter optimization rather than increased model capacity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent incorporates feedback mechanisms where classification results are validated against reference databases and used to refine future training sets. This feedback loop continuously improves accuracy while stabilizing system complexity by learning from past performance and adjusting parameters accordingly

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20220122696A1System and method for achieving high gene data resolution using training sets
Publication Date: 2022.04.21 ADA FORSYTH INSTITUTE INC
  • US20220122696A1 patent drawing
  • US20220122696A1 patent drawing
  • US20220122696A1 patent drawing

AI summary

Systems, methods, and computer program products for generating an enhanced set of sequences for taxonomical classification are disclosed. In various embodiments, a plurality of reference sequences are received. Each of the plurality of reference sequences corresponds to a taxonomical classification. A label corresponding to at least one of the reference sequences is assigned to each of a plurality of supplemental sequences. Each of the plurality of supplemental sequences and each of the plurality of reference sequences are truncated to a region of interest to thereby generate a truncated set of sequences. Similarity is measured between pairs of truncated sequences in the truncated set of sequences to determine whether the similarity is above a predetermined threshold. An intermediate taxonomical label is assigned to the pair of truncated sequences in the truncated set of sequences when the similarity is above the predetermined threshold to thereby generate an enhanced set of sequences.