Microflora identification method, system and device

By constructing a lightweight ribosomal hidden Markov model database, identifying and extracting ribosomal RNA sequences, the problems of large amounts of metagenomic/transcriptomical classification data and their inability to be reused were solved, enabling efficient and flexible identification of microbial communities.

CN121237228AActive Publication Date: 2025-12-30TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511767022.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2025-12-30
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

Existing metagenomic/transcriptomics classification data based on whole sequence alignment is large in volume and the output data cannot be reused, resulting in high computational resource requirements and limited analysis speed, making it impossible to seamlessly integrate with traditional amplicon analysis workflows.

Method used

A lightweight ribosomal hidden Markov model database is used for global search. By constructing various types of ribosomal RNA sequence models, ribosomal RNA sequences are identified and extracted. Combined with hypervariable region sequences and quality information feedback, a hypervariable region sequence file containing complete quality information is generated.

Benefits of technology

It reduces storage space and computing resource requirements, improves computing efficiency, generates reusable data that supports both traditional and in-depth analysis, and ensures identification accuracy and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237228A_ABST
    Figure CN121237228A_ABST
Patent Text Reader

Abstract

The invention discloses a microbial community identification method and system, and particularly relates to the technical field of microbiomics. A hidden Markov model database constructed on the basis of multiple types of ribosome RNA sequences is adopted for global search, a huge whole genome database in the prior art is replaced, and the size of the database is reduced; moreover, the hypervariable region sequence is extracted from the identified ribosome RNA sequence in a targeted manner, and the extracted sequence data and the quality information in the original sequencing data are pasted back, so that the hypervariable region sequence file containing the complete quality information is generated. The sequence data contains sequence information for identification, and a ribosome RNA marker gene sequence can be accurately identified and extracted from metagenome or metatranscriptome sequencing data; the original sequencing data can be used for a traditional amplicon analysis process, and can also support downstream deep analysis, so that seamless connection with the traditional amplicon analysis process is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microbiome technology, and in particular to a method and system for identifying microbial communities. Background Technology

[0002] With the development of high-throughput sequencing technology, metagenomics and metatranscriptomics have become mainstream methods for studying complex microbial communities. Metagenomics reveals the species composition and genetic blueprint of microbial communities by directly extracting total DNA from environmental samples for sequencing; metatranscriptomics, on the other hand, captures the real-time metabolic activities and functional responses of microbial communities by extracting total RNA (mainly messenger RNA) for sequencing.

[0003] Currently, the observation of microbial community structure mainly relies on two technical approaches: one is the traditional method based on amplicon sequencing, such as the QIIME2 procedure for bacterial 16S rRNA genes or the MOTHUR procedure for fungal ITS regions. This method uses specific primers to amplify target marker genes via PCR, achieving efficient and low-cost identification of microbial communities. However, when studies require both functional gene information and community structure information, metagenomic sequencing and amplicon sequencing must be performed separately, leading to a significant increase in experimental costs and increased complexity in data analysis. The other approach is metagenomic / transcriptomics classification tools based on whole-sequence alignment, such as Kraken2 and Bracken. These tools compare the short reads generated by sequencing with a vast database containing numerous microbial whole-genome sequences, directly identifying the species origin of the sequences and estimating species abundance.

[0004] While metagenomic / transcriptome classification based on whole-sequence alignment avoids PCR amplification bias and can simultaneously acquire community structure and functional gene information, its reference databases are typically enormous (often hundreds of gigabytes), leading to cumbersome software deployment, high computational resource requirements, large computational overhead, and limited analysis speed. Furthermore, the output of metagenomic / transcriptome classification is usually a taxonomic table, and the output data does not include complete 16S or 18S rRNA gene sequences with original sequencing quality information. As a result, the output data cannot be used for downstream in-depth analysis, nor can it be seamlessly integrated with traditional amplicon analysis workflows, limiting the potential for data reuse. Summary of the Invention

[0005] The main objective of this invention is to address the technical problems of large data volume and non-reusability of output data in existing metagenomic / transcriptomics classification methods based on whole sequence alignment. It proposes a lightweight, efficient, and reusable method that can accurately identify and extract ribosomal RNA marker gene sequences from metagenomic / transcriptomics data.

[0006] To achieve the above objectives, this invention proposes a method for identifying microbial communities, comprising the following steps: S100, acquiring raw sequencing data of metagenomics or metatranscriptomics from environmental samples; S200, converting the raw sequencing data containing quality information into sequence data without quality information; S300, searching the sequence data without quality information based on a model database constructed from multiple types of ribosomal RNA sequences, and globally identifying ribosomal RNA sequences; S400, screening 16S and / or 18S ribosomal RNA sequences from the ribosomal RNA sequences, and extracting hypervariable sequences from the 16S and / or 18S ribosomal RNA sequences, and constructing a hypervariable region identification model based on the hypervariable region sequences; S500, pasting the hypervariable region sequences back with the quality information in the raw sequencing data to generate a hypervariable region sequence file containing quality information; S600, analyzing the hypervariable region sequence file using analysis software to obtain microbial community and structural information.

[0007] Preferably, in step S200: the original sequencing data is a FASTQ format file and the sequence data is a FASTA format file; the FASTQ format file is converted into a FASTA format file using a sequence format conversion tool.

[0008] Preferably, step S300 specifically includes the following steps: S301, obtaining a full-length reference ribosomal RNA sequence from a standard ribosomal RNA sequence database; S302, performing multiple sequence alignment on the obtained sequence data using a multiple sequence alignment tool; and using a hidden Markov model construction tool to construct a ribosomal hidden Markov model as a model database based on the aligned sequence, wherein the fs mode is used for construction to identify incomplete rRNA gene fragments.

[0009] Preferably, the size of the model database is optimized to the MB level.

[0010] Preferably, in step S400, based on the gene type identifier in the sequence annotation information, gene sequences belonging to prokaryotes (16S rRNA) and eukaryotes (18S rRNA) are specifically screened from the ribosomal RNA sequences as hypervariable sequences.

[0011] Preferably, in step S400, constructing a high-variability region identification model based on the high-variability region sequence specifically includes the following steps: S401, obtaining representative sequences of high-variability regions from the NCBI database; S402, performing multi-sequence alignment on the obtained representative sequences using a multi-sequence alignment tool; S403, using a hidden Markov model construction tool to construct a high-variability region hidden Markov model for identifying specific high-variability regions based on the aligned sequences.

[0012] Preferably, in step S500, the sequence identifier of the hypervariable region sequence in the sequence data is used to match and reattach it with the original sequencing quality value.

[0013] This invention also provides a system for identifying microbial communities, comprising: a data acquisition and conversion unit, used to acquire raw sequencing data of metagenomics or metatranscriptomics from environmental samples, and convert raw sequencing data files in FASTQ format into sequence data files in FASTA format, the FASTA format sequence data files including multiple sequence identifiers and nucleic acid sequences; a global identification and extraction unit, including a pre-constructed ribosomal hidden Markov model database, used to perform sequence searches on the FASTA format files to identify and extract gene sequences belonging to ribosomal RNA, the ribosomal hidden Markov model database including multiple ribosomal hidden Markov models constructed based on 5S, 16S, 18S, 23S, and 28S ribosomal RNA sequences; and a sequence screening unit, used to screen from the identified gene sequences belonging to ribosomal RNA to identify 16S rRNA belonging to prokaryotes and 18S rRNA belonging to eukaryotes. The rRNA gene sequence; the targeted extraction section, which is used to search for hypervariable region sequences of the selected 16S rRNA and 18S rRNA gene sequences from a pre-constructed hypervariable region Hidden Markov Model database, so as to extract at least one specific hypervariable region sequence; the quality information re-attachment section, which is used to associate the targeted extracted specific hypervariable region sequence with the sequencing quality information in the original sequencing data to generate a sequence file containing complete sequencing quality information.

[0014] Compared with existing technologies, the advantages of this invention are as follows: By employing a Hidden Markov Model database constructed based on multiple types of ribosomal RNA sequences for global search, it replaces the massive whole-genome database in existing technologies, reducing database size and thus lowering the demand for storage space and computing resources, resulting in a significant improvement in deployment convenience and computational efficiency. Furthermore, this invention generates a hypervariable region sequence file containing complete quality information by targeting and extracting hypervariable region sequences from the identified ribosomal RNA sequences and then back-linking the extracted sequence data with the quality information from the original sequencing data. Moreover, by using a lightweight ribosomal Hidden Markov Database for initial screening and combining the two steps of targeted extraction of hypervariable region sequences and quality information back-linking, synergy is achieved, ensuring identification accuracy while greatly reducing computational and storage overhead, and generating reusable data that supports both traditional and in-depth analysis. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0016] Figure 1 This is a schematic diagram of the process for extracting ribosomal RNA marker genes from metagenomic or metatranscriptomic data in one embodiment; Figure 2 As shown in one embodiment, the analysis results of the microbial community structure of nine groups of samples (s1 to s9) were obtained by the method of the present invention. Figure 3 This is a diagram showing the analysis results of the microbial community structure of nine samples (s1 to s9) using Kraken2 software in one embodiment. The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0017] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments.

[0018] Implementation 1 In this embodiment, as Figure 1 The method for identifying a microbial community includes steps S100 to S600.

[0019] In this embodiment, S100, raw sequencing data of metagenomics or metatranscriptomics are obtained from environmental samples; specifically, sequencing data of total metagenomic DNA or total RNA from environmental samples are obtained; when processing total RNA sequencing data, this method aims to identify and extract ribosomal RNA (rRNA) sequences from it.

[0020] In this embodiment, the raw sequencing data is a FASTQ format file from the Illumina platform (a sequencing platform) that has undergone quality control. Quality control of the raw sequencing data can be achieved using fastp (a FASTQ preprocessing tool) to remove low-quality sequences, adapters, and trim low-quality bases.

[0021] S200. Convert the raw sequencing data containing quality information into sequence data without quality information; specifically, use SeqKit (a cross-platform sequence processing tool) to convert FASTQ format files containing quality values ​​into FASTA format files without quality information for subsequent model search.

[0022] In other possible embodiments, sequence processing tools such as bioawk, seqtk, or self-written Python scripts (using the Biopython library) can be used to convert FASTQ format files containing quality values ​​into FASTA format files without quality information.

[0023] S300, a model database built based on multiple types of ribosomal RNA sequences, searches sequence data that does not contain quality information and globally identifies ribosomal RNA sequences.

[0024] In this embodiment, step S300 specifically includes the following steps: S301. Obtain the full-length reference ribosomal RNA sequence from a standard ribosomal RNA sequence database; S302. Perform multiple sequence alignment on the obtained sequence data using a multiple sequence alignment tool, and construct a ribosomal hidden Markov model as a model database based on the aligned sequence using a hidden Markov model construction tool, wherein the fs mode is used for construction to identify incomplete rRNA gene fragments.

[0025] Specifically, in step S301, full-length 5S, 16S, 18S, 23S, and 28S ribosomal RNA sequences are obtained from a standard database.

[0026] Specifically, in step S302, the ribosomal RNA sequence obtained in step S301 is used for multiple sequence alignment to construct a ribosomal hidden Markov model (HMM) for recognizing ribosomal RNA.

[0027] In this embodiment, the model database of the ribosomal hidden Markov model contains only the core rRNA (ribosomal RNA) sequence information, thereby carefully optimizing and reducing the size of the database to be much smaller than that of the whole genome database.

[0028] In this embodiment, when constructing the ribosomal hidden Markov model, the default parameters of the MAFFT (Multiple Alignment using Fast Fourier Transform) software are first used for multiple sequence alignment. Then, the fs (full sequence) mode of hmmbuild (a tool in the HMMER (Standard Tools) package used to construct HMM profile models from multiple sequence alignment files) is used for construction in order to better identify incomplete rRNA gene fragments.

[0029] In this embodiment, the standard database selected is the 5S ribosome database, PR2, or Sliva.

[0030] S400. Screening 16S and / or 18S ribosomal RNA sequences from ribosomal RNA sequences, extracting hypervariable regions from 16S and / or 18S ribosomal RNA sequences, and constructing a hypervariable region identification model based on the hypervariable region sequences.

[0031] Specifically, screening for 16S and / or 18S ribosomal RNA sequences from ribosomal RNA sequences involves using the ribosomal hidden Markov model database constructed in step S300 to search the FASTA format sequence data obtained in step 200 using the hmmsearch command (a tool in the HMMER3 software package that uses an HMM model to search the sequence database and outputs matching regions and significance scores). This comprehensively identifies and extracts gene sequences belonging to 5S, 16S, 18S, 23S, and 28S ribosomal RNAs from the massive amount of raw sequencing data. This step serves as a preliminary screening, rapidly locating all possible ribosomal RNA sequences from the vast amount of raw sequencing data, significantly reducing the data scale for subsequent analyses.

[0032] In this embodiment, hypervariable regions are extracted from 16S and / or 18S ribosomal RNA sequences. A hypervariable region identification model constructed based on these sequences includes the following steps: S401. Obtain representative sequences of the hypervariable region from the NCBI database. Specifically, using the gene type identifier in the sequence annotation information, specifically screen gene sequences belonging to prokaryotic 16S rRNA and eukaryotic 18S rRNA from all ribosomal RNA gene sequences as the target sequence set for subsequent hypervariable region analysis. Then, obtain representative sequences of V4 (Variable region 4 of 16S rRNA) and V9 (Variable region 9 of 16S rRNA) from standard databases [e.g., NCBI (National Center for Biotechnology Information)]. Construct dedicated hypervariable region Hidden Markov Models for V4 and V9 respectively. Using the hypervariable region Hidden Markov Models, directly identify and extract specific sequences of regions V4 and V9 from the original FASTQ data.

[0033] S402. Use a multiple sequence alignment tool to perform multiple sequence alignment on the obtained representative sequences; In this embodiment, the multiple sequence alignment tool is Clustal Omega or MUSCLE (Multiple Sequence Comparison by Log-Expectation).

[0034] S403. Using a Hidden Markov Model (HMM) construction tool, construct a high-variability region HMM based on the aligned sequences to identify specific high-variability regions. Specifically, target and extract the V4 and V9 high-variability regions of the 16S and 18S rRNA sequences obtained in step S401.

[0035] In this embodiment, the construction process and parameters of the hypervariable region hidden Markov model are similar to those of the ribosomal hidden Markov model. Both first perform multiple sequence alignment using the default parameters of the MAFFT software, and then use the fs mode of hmmbuild for construction.

[0036] S500: The hypervariable region sequences are back-attached to the quality information in the original sequencing data, generating a hypervariable region sequence file containing quality information. Specifically, for specific sequences of regions V4 and V9 directly identified and extracted from the original FASTQ data, their sequence identifiers in the FASTQ file are used to accurately associate and back-attach them with the quality values ​​of the original sequencing data. In steps S200, S300, and S400, all processing tools (such as SeqKit and hmmsearch) must use parameter settings that retain the original sequence identifiers (Sequence IDs) to ensure that the sequence IDs output in each step are completely consistent with the IDs in the original FASTQ file, providing an accurate mapping basis for subsequent quality information back-attachment. This results in two sets of V4 / V9 region sequence data: one in FASTA format without quality values, and the other in FASTQ format containing complete sequencing quality information.

[0037] S600. Using the two sets of sequence data obtained from step S500, the extracted hypervariable region sequence files can be analyzed using analysis software for downstream bioinformatics analysis.

[0038] In this embodiment, the original sequencing quality information is preserved and reattached while extracting specific hypervariable regions. This allows for flexible selection during downstream bioinformatics analysis, enabling the use of either FASTA data without quality values ​​for standard analysis, or FASTQ data with quality values ​​for more precise error correction and ASV (Amplicon Sequence Variant) analysis, depending on the requirements. This approach is more flexible than existing technologies and meets the needs of different analytical depths and precisions.

[0039] In this embodiment, bioinformatics analysis refers to the analysis of the structure and diversity of microbial communities, thereby enabling the analysis and structural identification of microbial communities.

[0040] In this embodiment, the extracted rRNA sequences can also be checked for chimerism using analysis software, and sequences identified as chimeras can be removed to ensure stable classification of the microbial community structure.

[0041] In this embodiment, the analysis software used to analyze the hypervariable region sequences is DADA2 (Denoising Algorithm for Amplicon Data using Divisive Amplicon Denoising Algorithm 2), QIIME2 (Quantitative Insights Into Microbial Ecology 2), USEARCH (Ultra-fast Search), or VSEARCH (Open-source replacement for USEARCH). In this embodiment, DADA2, QIIME2, USEARCH, and VSEARCH are all commercially available software, and their specific algorithms are not detailed here.

[0042] In this embodiment, the model database of the ribosomal hidden Markov model contains only the core rRNA (ribosomal RNA) sequence information, and the size of the ribosomal hidden Markov model database is optimized to the MB level after the original sequencing data is processed by fastp and fastuniq.

[0043] In one possible embodiment, the ribosomal hidden Markov model database is constructed based on the SSU (small subunit) and LSU (large subunit) rRNA sequences of bacteria and archaea in the SILVA database (version 138) and the SSU (small subunit) and LSU (large subunit) database in the PR2 database (version 5.1.1), resulting in a database file size of approximately 51.1 MB.

[0044] Common model databases used for sequencing (e.g., the Kraken2 standard database) typically weigh approximately 59GB to 69GB after decompression, depending on the type of reference library chosen, such as bacteria, viruses, or fungi. Therefore, the model database of this invention is significantly smaller than existing databases, offering substantial advantages in terms of lightweight design, ease of downloading, and deployment.

[0045] Example 2 This embodiment provides a comparative example based on Embodiment 1.

[0046] In this embodiment, as Figure 2 and Figure 3As shown, Figure 2 This is a microbial community structure distribution map obtained after performing biological analysis on nine groups of samples (s1 to s9) using the method described in Example 1. Figure 3 This is a microbial community structure distribution map obtained after biological analysis of nine groups of samples (s1-s9) using Kraken2 software. (Comparison) Figure 2 and Figure 3 It is evident that the results obtained by analyzing the top ten most abundant microbial communities in the nine sample tests, using the method described in Example 1 and the results obtained by analyzing with Kraken2 software, show extremely small errors.

[0047] Example 3 This embodiment provides a system for identifying microbial communities, based on any of the above embodiments.

[0048] This embodiment of a system for identifying microbial communities includes: a data acquisition and conversion unit, used to acquire raw sequencing data of metagenomics or metatranscriptomics from environmental samples, and convert raw sequencing data files in FASTQ format into sequence data files in FASTA format, the FASTA format sequence data files including multiple sequence identifiers and nucleic acid sequences; a global identification and extraction unit, including a pre-constructed ribosomal hidden Markov model database, used to perform sequence searches on FASTA format files to identify and extract gene sequences belonging to ribosomal RNA, the ribosomal hidden Markov model database including multiple ribosomal hidden Markov models constructed based on 5S, 16S, 18S, 23S, and 28S ribosomal RNA sequences; and a sequence screening unit, used to screen from the identified gene sequences belonging to ribosomal RNA to identify 16S rRNA belonging to prokaryotes and 18S rRNA belonging to eukaryotes. The rRNA gene sequence; the targeted extraction section, which is used to search for hypervariable region sequences of the selected 16S rRNA and 18S rRNA gene sequences from a pre-constructed hypervariable region Hidden Markov Model database, so as to extract at least one specific hypervariable region sequence; the quality information re-attachment section, which is used to associate the targeted extracted specific hypervariable region sequence with the sequencing quality information in the original sequencing data to generate a sequence file containing complete sequencing quality information.

[0049] Example 4 This embodiment provides a microbial community identification device based on any of the above embodiments.

[0050] This embodiment of a microbial community identification device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned method for identifying microbial communities.

[0051] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, several equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or purpose, should be considered within the scope of protection of the present invention.

Claims

1. A method for identifying a microbial community, characterized by, The method comprises the following steps: S100, obtaining raw sequencing data of metagenome or metatranscriptome in an environmental sample; S200, converting the raw sequencing data containing quality information into sequence data not containing quality information; S300, searching the sequence data not containing quality information based on a model database constructed based on multiple types of ribosomal RNA sequences, and globally identifying ribosomal RNA sequences; S400, screening 16S and / or 18S ribosomal RNA sequences from the ribosomal RNA sequences, extracting high variable region sequences from the 16S and / or 18S ribosomal RNA sequences, and constructing a high variable region identification model based on the high variable region sequences; S500, back mapping the high variable region sequences to the quality information in the raw sequencing data to generate a high variable region sequence file containing quality information; S600, analyzing the high variable region sequence file by an analysis software to obtain microbial community and structure information.

2. The method of identifying a microbial community according to claim 1, wherein In step S200: The raw sequencing data is a file in FASTQ format, and the sequence data is a file in FASTA format; The FASTQ format file is converted into a FASTA format file by a sequence format conversion tool.

3. The method of identifying a microbial community according to claim 1, wherein Step S300 specifically comprises the following steps: S301, obtaining full-length reference ribosomal RNA sequences from a standard ribosomal RNA sequence database; S302, performing multiple sequence alignment on the obtained sequence data using a multiple sequence alignment tool; and constructing the ribosomal hidden Markov model as the model database based on the aligned sequences using a hidden Markov model construction tool.

4. The method of identifying a microbial community according to claim 3, wherein The volume of the model database is optimized to the MB level.

5. The method of identifying a microbial community according to claim 1, wherein In step S400, based on the gene type identification in the sequence annotation information, the 16S rRNA gene sequence belonging to prokaryotes and the 18S rRNA gene sequence belonging to eukaryotes are specifically screened from the ribosomal RNA sequences as high variable region sequences.

6. The method of identifying a microbial community according to claim 5, wherein In step S400, constructing a high variable region identification model based on high variable region sequences specifically comprises the following steps: S401, obtaining representative sequences of high variable regions from an NCBI database; S402, performing multiple sequence alignment on the obtained representative sequences using a multiple sequence alignment tool; S403, constructing a high variable region hidden Markov model for identifying specific high variable regions based on the aligned sequences using a hidden Markov model construction tool.

7. The method of identifying a microbial community according to claim 1, wherein In step S500, the high variable region sequence is matched and back mapped to the original sequencing quality value through the sequence identifier of the high variable region sequence in the sequence data.

8. A system for the identification of a microbial community, characterized in that It comprises: A data acquisition and conversion unit, which is used for obtaining raw sequencing data of metagenome or metatranscriptome derived from an environmental sample, and converting a raw sequencing data file in FASTQ format into a sequence data file in FASTA format, the sequence data file in FASTA format comprising multiple sequence identifiers and nucleic acid sequences; a global identification and extraction unit comprising a pre-constructed ribosomal hidden Markov model database for sequence search of the FASTA format file to identify and extract gene sequences belonging to ribosomal RNA, the ribosomal hidden Markov model database comprising a plurality of ribosomal hidden Markov models constructed based on 5S, 16S, 18S, 23S and 28S ribosomal RNA sequences; a sequence screening unit for screening from the identified gene sequences belonging to ribosomal RNA, gene sequences belonging to 16S rRNA of prokaryotes and 18S rRNA of eukaryotes; a targeted extraction unit for hypervariable region sequence search of the screened 16S rRNA and 18S rRNA gene sequences from a pre-constructed hypervariable region hidden Markov model database to target extraction of at least one specific hypervariable region sequence; a quality information backsticking unit for associating the targeted extracted specific hypervariable region sequence with sequencing quality information in the original sequencing data to generate a sequence file containing complete sequencing quality information.

9. A microbial community identification device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the microbial community identification method of any one of claims 1-7. a global identification and extraction unit comprising a pre-constructed ribosomal hidden Markov model database for sequence search of the FASTA format file to identify and extract gene sequences belonging to ribosomal RNA, the ribosomal hidden Markov model database comprising a plurality of ribosomal hidden Markov models constructed based on 5S, 16S, 18S, 23S and 28S ribosomal RNA sequences; a sequence screening unit for screening from the identified gene sequences belonging to ribosomal RNA, gene sequences belonging to 16S rRNA of prokaryotes and 18S rRNA of eukaryotes; a targeted extraction unit for hypervariable region sequence search of the screened 16S rRNA and 18S rRNA gene sequences from a pre-constructed hypervariable region hidden Markov model database to target extraction of at least one specific hypervariable region sequence; a quality information backsticking unit for associating the targeted extracted specific hypervariable region sequence with sequencing quality information in the original sequencing data to generate a sequence file containing complete sequencing quality information. The processor executes the computer program to implement the microbial community identification method of any one of claims 1-7.

Citation Information

Patent Citations

  • Method for constructing environmental microbial genome draft

    CN105420375A

  • Method for accurately identifying unknown microbial community in water body based on metagenomics analysis

    CN112786102A

  • Metagenomic library and natural product discovery platform

    US20210257056A1

  • Systems and methods for identifying microbial species and improving health

    US20250054578A1