Self-Consistent Genome Taxonomy for k-mer Database Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing genome databases face inaccuracies and inefficiencies due to errors in metadata and the non-self-consistent taxonomic classifications, leading to incorrect organism identification and classification, which can result in inappropriate treatments and wrongful penalties.

Innovation Solution

A method is developed to create a self-consistent taxonomy for k-mer databases by calculating genetic distances, removing clusters with insufficient data, recalculating taxonomy, and correcting metadata, resulting in a more accurate and efficient classification system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If standard taxonomy from NCBI is used for genome classification, then existing genome databases can be quickly accessed and used, but the taxonomy contains metadata errors and is not self-consistent, leading to incorrect organism identification

Engineering Contradiction:
Improveaccuracy of organism classificationVSAvoidcomplexity of taxonomy construction
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-calculating genetic distances between all genome pairs and pre-cluster genomes into taxonomic groups before the actual classification query. This preprocessing creates a self-consistent taxonomy structure that eliminates metadata errors, allowing accurate classification without manual curation during runtime.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of the taxonomy structure that is self-consistent and error-free, separate from the original NCBI taxonomy containing metadata errors. This copied taxonomy is built from actual genetic distance calculations and cluster assignments, providing a reliable alternative for classification that can be used without modifying the original database.

Inventive Principle:
Principle #26Copying

2Reliability

If manual curation is performed to correct taxonomy errors, then classification accuracy improves, but the process becomes time-consuming and cannot keep pace with hundreds of thousands of genomes

Engineering Contradiction:
Improveaccuracy of taxonomic labelingVSAvoidspeed of database curation
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements self-service by enabling the taxonomy to correct its own metadata errors automatically through algorithmic processes. The system uses genetic distance calculations and cluster analysis to identify and correct mislabeled genomes without human intervention, making the curation process autonomous and scalable to handle hundreds of thousands of genomes efficiently.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the fundamental parameter of taxonomy construction from metadata-based classification to genetic-distance-based classification. By using calculated genetic distances as the primary sorting criterion rather than relying on potentially erroneous metadata fields, the system automatically produces accurate taxonomic labels at scale without manual curation.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If k-mer database is built with inaccurate taxonomy, then database construction is fast, but the database loses ability to accurately identify specific taxa and produces wrong identifications

Engineering Contradiction:
Improvespecificity of taxonomic identificationVSAvoidease of database construction
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates a corrected copy of the taxonomy structure specifically for k-mer database construction, eliminating metadata errors while maintaining the computational efficiency needed for fast database building. This self-consistent taxonomy copy ensures that the k-mer database achieves both high identification specificity and ease of construction through automated processes.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11347810B2Methods of automatically and self-consistently correcting genome databases
Publication Date: 2022.05.31 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11347810B2 patent drawing
  • US11347810B2 patent drawing
  • US11347810B2 patent drawing

AI summary

A method is described for automatically correcting metadata errors in a k-mer database. A k-mer database having a self-consistent taxonomy based on genome-genome distance was constructed from a set of sample and reference genomes whose metadata included taxonomic labeling from a reference taxonomy (the standard NCBI taxonomy), which is not based on genetic distance. As a result, genomes of a given taxonomic ID of the self-consistent taxonomy could be separated into clusters based on the differences in the metadata. Genomes of the clusters less than a minimum cluster size Cmin were removed and profiled against the remaining genomes, correcting metadata automatically for those genomes that could be mapped back. The resulting k-mer database showed improved specificity for genetic profiling. Another method is described for identifying and handling chimeric genomes using the self-consistent taxonomy. Another method is described for correcting a classification database.