OTU Classification Using Segmented Marker Gene Databases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for taxonomic classification of metagenomic samples using short read amplicon sequences are inaccurate due to reliance on short regions of phylogenetic marker genes, leading to variable OTU clustering results and sub-optimal identification of operational taxonomic units (OTUs).

Innovation Solution

A system and method that create a customized OTU database (OTUX) using predefined segments of nucleotide sequences, calculating propensity scores for OTUs, and building a mapping matrix to enhance the accuracy of OTU classification, allowing for improved classification of short read amplicon sequences into appropriate OTUs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short read amplicon sequences are used for OTU classification, then cost effectiveness and throughput are improved, but classification accuracy deteriorates

Engineering Contradiction:
ImprovethroughputVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the full-length marker gene into multiple predefined segments or regions. Each segment is independently clustered to create segment-specific OTU databases. This segmentation allows short reads to be classified using a database optimized for their specific region, improving accuracy while maintaining the cost-effectiveness of short-read sequencing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary clustering of sequences from each predefined segment to create customized OTU databases before actual classification. This preliminary action prepares segment-specific reference databases that capture the variability within each region, enabling more accurate classification of short reads without requiring full-length sequencing.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If full-length marker genes are used for OTU database construction, then classification accuracy is improved, but sequencing cost and complexity increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidsequencing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Instead of requiring full-length marker gene sequencing, the patent divides the marker gene into multiple predefined segments. Each segment is clustered separately to build specialized OTU databases. This approach achieves high classification accuracy using only short-read sequencing of individual segments, avoiding the complexity and cost of full-length sequencing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses a partial approach by sequencing only specific predefined segments of the marker gene rather than the entire length. This partial action is sufficient for accurate classification when combined with segment-specific OTU databases, reducing sequencing complexity while maintaining or improving accuracy compared to full-length approaches.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If conventional OTU databases are used for classification, then compatibility with existing methods is maintained, but classification accuracy for short reads deteriorates

Engineering Contradiction:
Improvemethod compatibilityVSAvoidclassification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates customized OTU databases segmented by predefined regions of the marker gene. Each segment database is built from clustered sequences specific to that region, allowing short reads to be classified with high accuracy against a database optimized for their length and origin, while maintaining compatibility with existing classification workflows.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the organizational parameter of OTU databases from whole-gene based to segment-based. By clustering and storing sequences according to predefined segments rather than full-length genes, the database structure adapts to the characteristics of short-read sequencing, improving classification accuracy while maintaining interface compatibility with standard tools.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11996170B2Method and system for identification and classification of operational taxonomic units in a metagenomic sample
Publication Date: 2024.05.28 TATA CONSULTANCY SERVICES LTD
  • US11996170B2 patent drawing
  • US11996170B2 patent drawing
  • US11996170B2 patent drawing

AI summary

A system and method for identification and classification of operational taxonomic units (OTUs) in a metagenomic sample using short read amplicon sequences has been described. The disclosure enables accurate identification of OTU in a metagenomic sample and provides a framework for easy cross comparison of microbiome community structures sampled across different disconnected metagenomic studies. Instead of using a reference database consisting of full-length marker genes directly for taxonomic classification or OTU-picking, the present disclosure creates customized OTU databases for different hyper-variable regions of a marker gene. These databases consist of reference OTUs obtained through independent clustering of sequences pertaining to different selected hyper-variable regions of the marker gene. In another embodiment, mapping back is also provided facilitating cross comparison between results obtained from different studies that may have utilized different hyper-variable regions. The system results in enhanced accuracy of classification of operational taxonomic units (OTUs) in the metagenomic sample.