OTU Classification Using Segmented Marker Gene Databases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for taxonomic classification of metagenomic samples using short read amplicon sequences are inaccurate due to reliance on short regions of phylogenetic marker genes, leading to variable OTU clustering results and sub-optimal identification of operational taxonomic units (OTUs).
Innovation Solution
A system and method that create a customized OTU database (OTUX) using predefined segments of nucleotide sequences, calculating propensity scores for OTUs, and building a mapping matrix to enhance the accuracy of OTU classification, allowing for improved classification of short read amplicon sequences into appropriate OTUs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short read amplicon sequences are used for OTU classification, then cost effectiveness and throughput are improved, but classification accuracy deteriorates
Solution Approach 1:
The patent segments the full-length marker gene into multiple predefined segments or regions. Each segment is independently clustered to create segment-specific OTU databases. This segmentation allows short reads to be classified using a database optimized for their specific region, improving accuracy while maintaining the cost-effectiveness of short-read sequencing.
Solution Approach 2:
The patent performs preliminary clustering of sequences from each predefined segment to create customized OTU databases before actual classification. This preliminary action prepares segment-specific reference databases that capture the variability within each region, enabling more accurate classification of short reads without requiring full-length sequencing.
2Measurement precision
If full-length marker genes are used for OTU database construction, then classification accuracy is improved, but sequencing cost and complexity increase
Solution Approach 1:
Instead of requiring full-length marker gene sequencing, the patent divides the marker gene into multiple predefined segments. Each segment is clustered separately to build specialized OTU databases. This approach achieves high classification accuracy using only short-read sequencing of individual segments, avoiding the complexity and cost of full-length sequencing.
Solution Approach 2:
The patent uses a partial approach by sequencing only specific predefined segments of the marker gene rather than the entire length. This partial action is sufficient for accurate classification when combined with segment-specific OTU databases, reducing sequencing complexity while maintaining or improving accuracy compared to full-length approaches.
3Adaptability or versatility
If conventional OTU databases are used for classification, then compatibility with existing methods is maintained, but classification accuracy for short reads deteriorates
Solution Approach 1:
The patent creates customized OTU databases segmented by predefined regions of the marker gene. Each segment database is built from clustered sequences specific to that region, allowing short reads to be classified with high accuracy against a database optimized for their length and origin, while maintaining compatibility with existing classification workflows.
Solution Approach 2:
The patent changes the organizational parameter of OTU databases from whole-gene based to segment-based. By clustering and storing sequences according to predefined segments rather than full-length genes, the database structure adapts to the characteristics of short-read sequencing, improving classification accuracy while maintaining interface compatibility with standard tools.
Data Source
AI summary
A system and method for identification and classification of operational taxonomic units (OTUs) in a metagenomic sample using short read amplicon sequences has been described. The disclosure enables accurate identification of OTU in a metagenomic sample and provides a framework for easy cross comparison of microbiome community structures sampled across different disconnected metagenomic studies. Instead of using a reference database consisting of full-length marker genes directly for taxonomic classification or OTU-picking, the present disclosure creates customized OTU databases for different hyper-variable regions of a marker gene. These databases consist of reference OTUs obtained through independent clustering of sequences pertaining to different selected hyper-variable regions of the marker gene. In another embodiment, mapping back is also provided facilitating cross comparison between results obtained from different studies that may have utilized different hyper-variable regions. The system results in enhanced accuracy of classification of operational taxonomic units (OTUs) in the metagenomic sample.


