Microbial Sequencing Identification Using Hierarchical Reference Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current microbial species identification technologies, particularly using next-generation sequencing (NGS), face challenges in distinguishing between highly similar long sequences like 16S rRNA genes and are limited to short sequence tests, leading to high costs, low sensitivity, and unreliable results due to non-microbial host nucleic acids and sequence similarity issues.

Innovation Solution

A method involving targeted enrichment and amplification of microbial characteristic nucleic acid sequences, followed by NGS sequencing, and a hierarchical clustering of reference sequences in a database to accurately identify microbial species, using a comparative analysis that iteratively filters and aligns sequencing data with stringent metrics to achieve species-level resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If metagenomic NGS sequences all nucleic acids in samples indiscriminately, then comprehensive microbial detection is achieved, but sequencing data waste increases and test sensitivity decreases due to overwhelming host nucleic acid background

Engineering Contradiction:
Improvemicrobial identification sensitivityVSAvoidsequencing data waste
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent extracts and enriches only the relevant microbial characteristic sequences from the complex sample nucleic acid mixture using targeted capture methods. This extraction principle separates the useful microbial signals from the overwhelming host background, achieving high sensitivity detection while minimizing waste of sequencing resources on irrelevant host sequences.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the enrichment parameters by using specific probes and capture conditions that selectively bind to microbial characteristic sequences. This parameter optimization allows differential enrichment of microbial sequences versus host sequences, resolving the contradiction between comprehensive detection and data efficiency.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If short fragments are used for NGS sequencing to accommodate platform limits, then sequencing feasibility is improved, but assembly accuracy decreases and chimeric sequences are generated due to high sequence similarity among species

Engineering Contradiction:
Improvesequencing compatibilityVSAvoidsequence assembly accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent performs preliminary enrichment and capture of full-length or long-fragment microbial characteristic sequences before sequencing. This preliminary action preserves longer sequence lengths that maintain species-specific discriminatory power, allowing accurate assembly and identification while still being compatible with NGS platforms through appropriate library preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transitions from analyzing short individual fragments to analyzing enriched long-fragment sequences in a different dimensional space. By enriching for longer fragments that span multiple variable regions, the patent achieves both NGS compatibility and sufficient sequence length for accurate species discrimination without generating chimeric assemblies.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If variable regions only are sequenced to reduce complexity, then sequencing cost and time are reduced, but identification resolution decreases and cannot achieve species-level distinction

Engineering Contradiction:
Improvesequencing efficiencyVSAvoidspecies identification resolution
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent designs a universal enrichment approach that captures multiple variable regions simultaneously within longer sequence fragments. This multi-functional capture strategy maintains high sequencing efficiency while obtaining sufficient phylogenetic information across multiple regions, enabling species-level identification without requiring separate sequencing of each variable region.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the sequencing of multiple variable regions into a single enriched library preparation step. By combining the capture of multiple phylogenetically informative regions into one workflow, the patent achieves both productivity (efficient single-step enrichment) and precision (sufficient sequence length and content for species-level resolution).

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250372206A1Methods, Devices, Computer Readable Storage Media, and Electronic Devices for Obtaining Microbial Species Identity and Related Information by Sequencing
Publication Date: 2025.12.04 INQUIRE LIFE DIAGNOSTICS INC
  • US20250372206A1 patent drawing
  • US20250372206A1 patent drawing
  • US20250372206A1 patent drawing

AI summary

This invention relates to the area of microorganism identification, specifically involving a method of obtaining microorganism identities and related information by sequencing. The method includes: i) obtaining sequencing data, said sequencing data are obtained by amplification of microbial characteristic sequences using primers followed by sequencing the amplification products using next-generation sequencing technology; ii) comparing said sequencing data with characteristic sequence database to identify microbial composition in said samples tested; wherein perform clustering on said characteristic sequence database in advance based on the sequence similarity among reference sequences containing said characteristic sequences, obtain one or more tiers of clusters, there is at least one child seed in each cluster, and there are several seeds as reference sequences in the bottom tier cluster.