Self-consistent k-mer database for accurate organism identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Standard taxonomic classification of microbes in official databases contains errors, leading to inaccurate organism identification and classification, which can result in incorrect treatment of patients and wrongful attribution of responsibility in outbreaks, as well as invalid conclusions in scientific and industrial research.
Innovation Solution
A self-consistent k-mer database is created by classifying k-mers using a taxonomy based on genetic distances, assigning unique IDs, and calculating weights/probabilities to ensure accurate taxonomic profiling of sequenced nucleic acids, thereby isolating errors in metadata and providing robust classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If standard taxonomic classification from official databases is used, then organism classification can be performed using existing databases, but classification accuracy deteriorates due to metadata errors and incorrect genome identification
Solution Approach 1:
The patent performs preliminary actions by calculating genetic distances between all genomes in the database before constructing the taxonomy. This pre-computed distance information is stored and used to build a self-consistent taxonomy that is independent of erroneous metadata, thereby resolving the contradiction between ease of construction and classification accuracy
Solution Approach 2:
The patent introduces genetic distance as an intermediary metric between genome sequences and taxonomic classification. Instead of directly using potentially erroneous metadata for classification, the system uses genetic distance calculations as a mediator to determine accurate taxonomic relationships, improving classification accuracy while maintaining database construction feasibility
2Productivity
If k-mer based classification with standard taxonomy is used, then classification speed is improved, but identification accuracy deteriorates due to errors propagating through the taxonomic tree
Solution Approach 1:
The patent uses genetic distance as an intermediary to build a self-consistent taxonomy that serves as an accurate reference for fast k-mer based classification. The pre-computed genetic distances eliminate metadata errors from the reference, allowing rapid classification to proceed with high reliability
Solution Approach 2:
The system performs preliminary genetic distance calculations and self-consistent taxonomy construction before the actual classification task. This pre-processing creates a reliable reference database that enables both fast and accurate k-mer based classification without error propagation
3Adaptability or versatility
If standard taxonomy is used despite metadata errors, then database construction can proceed with available data, but classification reliability deteriorates leading to incorrect treatment and wrongful attribution
Solution Approach 1:
The patent introduces genetic distance as a mediator that enables the database to remain usable with available data while simultaneously ensuring classification reliability. The genetic distance calculations provide an objective, error-resistant basis for taxonomy that maintains database versatility while preventing incorrect classifications
Solution Approach 2:
The patent changes the fundamental parameter used for taxonomy construction from metadata-based classification to genetic distance-based classification. This parameter change allows the database to remain usable with existing genome data while dramatically improving classification reliability by eliminating dependence on erroneous metadata
Data Source
AI summary
A large collection of sample genomes containing misclassified k-mers and metadata errors from a reference taxonomy was converted to a self-consistent k-mer database comprising a self-consistent taxonomy. The self-consistent taxonomy was based on genetic distances calculated using the MinHash method or the Meier-Koltoff method. An agglomerative clustering algorithm was used to calculate the self-consistent taxonomy. Each k-mer of the sample genomes was assigned to only one node of the self-consistent taxonomy. In another step, each node of the self-consistent taxonomy was mapped to the reference taxonomy, thereby preserving in the self-consistent taxonomy links to the reference taxonomy while correcting for the misclassification errors therein. The self-consistent k-mer database can be used to taxonomically profile sequenced nucleic acids with greater specificity compared to systems relying on the reference taxonomy.


