Phylogenetic Database for Bacterial Species Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current metagenomic sequencing methods, such as 16S profiling and whole genome sequencing, face limitations in distinguishing between closely related species and determining relative abundances of organisms, especially when dealing with unculturable bacteria and the complexity of the human microbiome, which hinders the identification and treatment of dysbiosis and diseases.
Innovation Solution
A method utilizing a database that stores reference genomes and their phylogenetic relationships, allowing for the analysis of sequence reads to determine unique mappings and normalize their counts based on the uniqueness of each genome, thereby providing relative abundance indications within a sample.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If 16S profiling is used to analyze microbial composition, then the method is highly efficient and standardized, but the resolution is limited to Family or Genus level and cannot distinguish between closely related species
Solution Approach 1:
The patent segments the genomic analysis by using specific marker genes (such as rpoB, gyrB, recA) that provide higher phylogenetic resolution than the 16S gene. This segmentation allows the method to maintain efficiency while achieving species-level or strain-level differentiation by focusing on informative genomic regions rather than analyzing entire genomes.
Solution Approach 2:
The patent changes the parameter of genetic marker selection from the highly conserved 16S rRNA gene to more variable marker genes (rpoB, gyrB, recA, etc.) that exhibit greater sequence divergence among closely related species. This parameter change enables the method to achieve higher phylogenetic resolution while maintaining the efficiency of targeted gene analysis.
2Adaptability or versatility
If whole genome de-novo assembly is used to overcome reference genome limitations, then culturing issues are addressed, but the method is extremely computationally intensive and requires substantial sequence coverage
Solution Approach 1:
The patent extracts and analyzes specific marker genes (rpoB, gyrB, recA, etc.) from the whole genome rather than performing complete de-novo assembly. This extraction approach maintains the advantage of not requiring reference genomes while dramatically reducing computational complexity and sequence coverage requirements by focusing only on informative regions.
Solution Approach 2:
The patent applies partial action by analyzing only specific marker genes rather than assembling entire genomes. This partial approach provides sufficient phylogenetic resolution for species and strain differentiation while avoiding the excessive computational burden of complete genome assembly, achieving an optimal balance between accuracy and efficiency.
3Speed
If lowest common ancestor approach is used for fast classification, then classification speed is improved, but the method provides information about read mapping rather than true relative abundances
Solution Approach 1:
The patent uses marker gene sequences as intermediaries to bridge the gap between fast classification and accurate abundance measurement. By mapping reads to specific marker genes (rpoB, gyrB, recA) rather than using lowest common ancestor approaches, the method maintains classification speed while providing more accurate relative abundance information through targeted gene analysis.
4Loss of information
If shotgun metagenomic sequencing is used to analyze complete DNA, then comprehensive genomic information is obtained, but the method cannot define complete genomic units and is limited to sequenced regions
Solution Approach 1:
The patent applies preliminary action by pre-selecting and analyzing specific marker genes (rpoB, gyrB, recA, etc.) that are known to provide complete phylogenetic information. This preliminary targeting approach ensures complete genomic unit definition for phylogenetic analysis while maintaining the comprehensive information gathering benefits of shotgun sequencing by focusing computational resources on informative regions.
Data Source
Figure 1~2(b)
Figure 3~4
Figure 5
AI summary
Methods are provided of using a database that stores a plurality of reference genomes and phylogenetic information which relates the stored reference genomes to each other in a phylogenetic structure. These methods are useful in analysing the bacteria and/or bacterial lineages present in a sample and to identify a bacterium for use in therapy.