Microbial Sequencing Identification Using Hierarchical Reference Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current microbial species identification technologies, particularly using next-generation sequencing (NGS), face challenges in distinguishing between highly similar long sequences like 16S rRNA genes and are limited to short sequence tests, leading to high costs, low sensitivity, and unreliable results due to non-microbial host nucleic acids and sequence similarity issues.
Innovation Solution
A method involving targeted enrichment and amplification of microbial characteristic nucleic acid sequences, followed by NGS sequencing, and a hierarchical clustering of reference sequences in a database to accurately identify microbial species, using a comparative analysis that iteratively filters and aligns sequencing data with stringent metrics to achieve species-level resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If metagenomic NGS sequences all nucleic acids in samples indiscriminately, then comprehensive microbial detection is achieved, but sequencing data waste increases and test sensitivity decreases due to overwhelming host nucleic acid background
Solution Approach 1:
The patent extracts and enriches only the relevant microbial characteristic sequences from the complex sample nucleic acid mixture using targeted capture methods. This extraction principle separates the useful microbial signals from the overwhelming host background, achieving high sensitivity detection while minimizing waste of sequencing resources on irrelevant host sequences.
Solution Approach 2:
The patent changes the enrichment parameters by using specific probes and capture conditions that selectively bind to microbial characteristic sequences. This parameter optimization allows differential enrichment of microbial sequences versus host sequences, resolving the contradiction between comprehensive detection and data efficiency.
2Ease of operation
If short fragments are used for NGS sequencing to accommodate platform limits, then sequencing feasibility is improved, but assembly accuracy decreases and chimeric sequences are generated due to high sequence similarity among species
Solution Approach 1:
The patent performs preliminary enrichment and capture of full-length or long-fragment microbial characteristic sequences before sequencing. This preliminary action preserves longer sequence lengths that maintain species-specific discriminatory power, allowing accurate assembly and identification while still being compatible with NGS platforms through appropriate library preparation.
Solution Approach 2:
The patent transitions from analyzing short individual fragments to analyzing enriched long-fragment sequences in a different dimensional space. By enriching for longer fragments that span multiple variable regions, the patent achieves both NGS compatibility and sufficient sequence length for accurate species discrimination without generating chimeric assemblies.
3Productivity
If variable regions only are sequenced to reduce complexity, then sequencing cost and time are reduced, but identification resolution decreases and cannot achieve species-level distinction
Solution Approach 1:
The patent designs a universal enrichment approach that captures multiple variable regions simultaneously within longer sequence fragments. This multi-functional capture strategy maintains high sequencing efficiency while obtaining sufficient phylogenetic information across multiple regions, enabling species-level identification without requiring separate sequencing of each variable region.
Solution Approach 2:
The patent merges the sequencing of multiple variable regions into a single enriched library preparation step. By combining the capture of multiple phylogenetically informative regions into one workflow, the patent achieves both productivity (efficient single-step enrichment) and precision (sufficient sequence length and content for species-level resolution).
Data Source
AI summary
This invention relates to the area of microorganism identification, specifically involving a method of obtaining microorganism identities and related information by sequencing. The method includes: i) obtaining sequencing data, said sequencing data are obtained by amplification of microbial characteristic sequences using primers followed by sequencing the amplification products using next-generation sequencing technology; ii) comparing said sequencing data with characteristic sequence database to identify microbial composition in said samples tested; wherein perform clustering on said characteristic sequence database in advance based on the sequence similarity among reference sequences containing said characteristic sequences, obtain one or more tiers of clusters, there is at least one child seed in each cluster, and there are several seeds as reference sequences in the bottom tier cluster.


