Self-consistent k-mer database for accurate organism identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Standard taxonomic classification of microbes in official databases contains errors, leading to inaccurate organism identification and classification, which can result in incorrect treatment of patients and wrongful attribution of responsibility in outbreaks, as well as invalid conclusions in scientific and industrial research.

Innovation Solution

A self-consistent k-mer database is created by classifying k-mers using a taxonomy based on genetic distances, assigning unique IDs, and calculating weights/probabilities to ensure accurate taxonomic profiling of sequenced nucleic acids, thereby isolating errors in metadata and providing robust classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If standard taxonomic classification from official databases is used, then organism classification can be performed using existing databases, but classification accuracy deteriorates due to metadata errors and incorrect genome identification

Engineering Contradiction:
Improveease of database constructionVSAvoidclassification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by calculating genetic distances between all genomes in the database before constructing the taxonomy. This pre-computed distance information is stored and used to build a self-consistent taxonomy that is independent of erroneous metadata, thereby resolving the contradiction between ease of construction and classification accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces genetic distance as an intermediary metric between genome sequences and taxonomic classification. Instead of directly using potentially erroneous metadata for classification, the system uses genetic distance calculations as a mediator to determine accurate taxonomic relationships, improving classification accuracy while maintaining database construction feasibility

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If k-mer based classification with standard taxonomy is used, then classification speed is improved, but identification accuracy deteriorates due to errors propagating through the taxonomic tree

Engineering Contradiction:
Improveclassification speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses genetic distance as an intermediary to build a self-consistent taxonomy that serves as an accurate reference for fast k-mer based classification. The pre-computed genetic distances eliminate metadata errors from the reference, allowing rapid classification to proceed with high reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary genetic distance calculations and self-consistent taxonomy construction before the actual classification task. This pre-processing creates a reliable reference database that enables both fast and accurate k-mer based classification without error propagation

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If standard taxonomy is used despite metadata errors, then database construction can proceed with available data, but classification reliability deteriorates leading to incorrect treatment and wrongful attribution

Engineering Contradiction:
Improvedatabase usabilityVSAvoidclassification reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces genetic distance as a mediator that enables the database to remain usable with available data while simultaneously ensuring classification reliability. The genetic distance calculations provide an objective, error-resistant basis for taxonomy that maintains database versatility while preventing incorrect classifications

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the fundamental parameter used for taxonomy construction from metadata-based classification to genetic distance-based classification. This parameter change allows the database to remain usable with existing genome data while dramatically improving classification reliability by eliminating dependence on erroneous metadata

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11830580B2K-mer database for organism identification
Publication Date: 2023.11.28 MARS INC
  • US11830580B2 patent drawing
  • US11830580B2 patent drawing
  • US11830580B2 patent drawing

AI summary

A large collection of sample genomes containing misclassified k-mers and metadata errors from a reference taxonomy was converted to a self-consistent k-mer database comprising a self-consistent taxonomy. The self-consistent taxonomy was based on genetic distances calculated using the MinHash method or the Meier-Koltoff method. An agglomerative clustering algorithm was used to calculate the self-consistent taxonomy. Each k-mer of the sample genomes was assigned to only one node of the self-consistent taxonomy. In another step, each node of the self-consistent taxonomy was mapped to the reference taxonomy, thereby preserving in the self-consistent taxonomy links to the reference taxonomy while correcting for the misclassification errors therein. The self-consistent k-mer database can be used to taxonomically profile sequenced nucleic acids with greater specificity compared to systems relying on the reference taxonomy.