Construction method of dehalogenation microorganism functional genome database and analysis method of dehalogenation microorganism
By constructing a hierarchical classification method based on multiple databases, the problems of low coverage and low analysis efficiency of dehalogenated microbial functional genome database in the prior art are solved, and efficient and accurate construction and analysis of dehalogenated microbial functional genome database are achieved.
Patent Information
- Application Number
- CN202510829729.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
When constructing a functional genome database for dehalogenated microorganisms, the existing technology has problems such as difficult to cultivate microorganisms, high limitations in molecular biological methods, low database coverage and inconsistent functional annotation, resulting in low analysis efficiency.
By constructing a hierarchical classification method based on Uniprot, arCOG, KOG, COG, eggNOG and KEGG databases, and combining with the NCBI RefSeq database, a functional genome database of dehalogenated microbials is constructed, and a variety of database comparison and deredundancy technologies are used to improve the comprehensiveness and accuracy of the database.
An efficient and accurate functional genome database of dehalogenated microbial organisms is realized, which can quickly identify and analyze functional communities of dehalogenated microbial organisms in the environment, and improve analysis efficiency and accuracy.
Smart Images

Figure CN120340631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method for constructing a functional genomic database of dehalogenating microorganisms and an analysis method for dehalogenating microorganisms. Background Art
[0002] Organohalides are organic compounds containing halogens such as fluorine, chlorine, bromine, and iodine. They are mainly produced by industrial and agricultural processes, such as the production of pesticides, drugs, and personal care products, etc. They are widely present in nature and have toxicity, persistence, and bioaccumulation. These organohalides often enter the environment through various channels, causing pollution and posing long-term hazards to human health and the ecosystem. Microorganisms in the environment usually exist in the form of communities. Through their interactions with the environment and each other, they drive the global geochemical cycle and significantly change the living environment. These complex microorganisms can participate in the dehalogenation process of organohalides to achieve the effective degradation of such substances. Microbial dehalogenation reactions are catalyzed by dehalogenases that cleave carbon-halogen bonds and are the key initial steps in the mineralization and detoxification of organohalides.
[0003] The dehalogenase-catalyzed dehalogenation reaction has a similar catalytic mechanism, and halogen substitution occurs through nucleophilic substitution. Currently, there are seven main types of dehalogenation reactions, namely reductive dehalogenation, oxidative dehalogenation, hydrolytic dehalogenation, thiolytic dehalogenation, dehydrohalogenation, methyl transfer dehalogenation, and intramolecular substitution reaction. These reactions occur under the action of reductases, oxidases (monooxygenases or dioxygenases), and haloacid dehalogenases, etc., with the participation of superoxide radicals, hydrogen electrons, and hydroxyl groups in the bond-breaking process. The degradation of such structurally complex organohalides by microorganisms involves the action of multiple functional genes and enzymes. For example, the dehalogenation reaction of microbial haloalkanes is catalyzed by dehalogenases encoded by dhmA, dhaAF, and linB, and the reductive dehalogenation of chloroform involves the enzyme encoded by the cfrA gene. With the continuous development of high-throughput sequencing technology, current technicians can clearly understand the classification, function, and phylogenetic diversity of various organohalide biodegradation genes through 16S rRNA and metagenomic sequencing methods, and clearly understand the microbial degradation process of various organohalides in complex environments.
[0004] There are currently methods for constructing a functional genome database of dehalogenating microorganisms, such as purification culture, molecular biology, and sequencing. The purification culture method involves screening some microorganisms in environmental samples to obtain dehalogenating microorganisms, followed by isolation and culture, and further study of their degradation ability, degradation rate, and adaptability. Usually, the dilution plate method or isolation culture method is used for isolation to gradually obtain single dehalogenating microorganisms. The molecular biology method directly extracts DNA from environmental samples and uses PCR technology to amplify genes related to dehalogenation reactions, such as the rdhA gene (reductive dehalogenase gene), pceA gene (tetrachloroethylene reductase gene), tceA gene (trichloroethylene reductase gene), etc. By detecting these genes, dehalogenating microorganisms can be quickly screened. The sequencing method can analyze the microbial community through high-throughput sequencing of the 16S rRNA gene to understand the types and community structures of dehalogenating microorganisms in the environment. However, the vast majority of dehalogenating microorganisms in the environment are difficult to culture under laboratory conditions, and it is also very difficult to isolate these microorganisms. Therefore, both the bacterial culture method and the PCR amplification method have significant limitations and are time-consuming and laborious. Moreover, these methods highly depend on the quality of the available sequence database. Although common public sequence databases such as eggNOG, KEGG, and Uniprot contain rich sequence information, there are problems such as inconsistent functional annotations, low coverage of target gene families, difficulty in distinguishing homologous genes, and long database search times in metagenomic analysis. Summary of the Invention
[0005] To solve the problems existing in the prior art, the present invention provides a method for constructing a functional genome database of dehalogenating microorganisms and an analysis method for dehalogenating microorganisms.
[0006] In the first aspect, the present invention provides a method for constructing a functional genome database of dehalogenating microorganisms, including: (1) Obtaining genes related to dehalogenation reactions; (2) Based on the Uniprot protein database, obtaining the seed sequences and homologous sequences of the families to which the genes related to dehalogenation reactions belong to obtain a core database; (3) Comparing the core database with a homologous protein database, and merging the sequences of the dehalogenating functional gene families obtained by the comparison with the core database to obtain a preliminary database; (4) Comparing the preliminary database with the Refseq database of microorganisms in NCBI, and obtaining a complete database through classification, merging homologous sequences, and removing redundancy.
[0007] In recent years, the amount of data generated by sequencing technology has been increasing. However, existing databases cannot meet the needs of analyzing the species and abundances of dehalogenating microorganisms and related functional genes. The present invention provides a construction method. The Uniprot protein database is the most comprehensive protein database currently. All protein entries in its Swiss-Prot sub-library have been manually annotated by experts and are supported by literature, with high data reliability. It is the source of protein sequences for constructing seed sequences in the present invention. The construction of the core database is to compare the seed sequences with the TrEMBL database in the UniProt protein database. The TrEMBL sub-library contains a large number of protein sequences that have not been manually annotated and can be used to supplement the seed sequences. Protein sequences in the UniProt database that have not been manually annotated but have dehalogenation functions are collected into the core database. Subsequently, by comparing the core database with homologous protein databases such as arCOG, KOG, COG, eggCOG, and the KEGG database and extracting sequences, the comprehensiveness of the database is improved, and homologous gene families are identified and incorporated into the complete database to reduce false positives that may occur in database retrieval. Among the five databases selected in the present invention, the arCOG, KOG, and COG databases are all based on expert manual annotation and have high reliability. The KEGG and eggNOG databases have the advantages of full species coverage, support for large-scale data analysis and systems biology analysis, etc. Moreover, all five databases adopt hierarchical functional classification, which is convenient for annotation and cross-study comparison. Finally, information on bacterial, archaeal, and fungal species covered by the database is obtained through the RefSeq database of microorganisms in NCBI. The sequences in the NCBI RefSeq database are of high quality and non-redundant, cover a wide range of species and are well-annotated, with clear classification hierarchies, making it convenient to obtain species information of protein sequences. A functional genomic database of dehalogenating microorganisms with high accuracy is constructed through the screening and utilization of each database, which has high analysis efficiency and accuracy when applied to the analysis of dehalogenating microorganisms.
[0008] Further, in step (1), the obtaining of genes related to dehalogenation reactions includes: Referring to the microbial dehalogenation reaction process in the KEGG database, genes related to dehalogenation reactions are obtained through literature research; the dehalogenation reactions include one or more of reductive dehalogenation, oxidative dehalogenation, hydrolytic dehalogenation, thiolytic dehalogenation, dehydrohalogenation, methyl transfer dehalogenation, or intramolecular substitution reaction dehalogenation.
[0009] Further, the genes related to dehalogenation reactions include: Genes related to reductive dehalogenation reactions, genes related to oxidative dehalogenation reactions, genes related to intramolecular substitution reactions, genes related to hydrolytic dehalogenation reactions, genes related to dehydrohalogenation reactions, genes related to methyl transfer dehalogenation reactions, and genes related to thiolytic dehalogenation reactions; The genes related to reductive dehalogenation include: cfrA, cprA, cprA5, dcaA, dcpA, dcrA, DebcprA, linD, pceA, and vcrA; The genes related to oxidative dehalogenation reactions include: bnz, bphA, bphB, bphC, bphD, cbdA, cbdB, cbdC, clcA, cpo, GLG, hadA, LIP, LPOA, mnp, pcpB, tcbC, tecA, and tfdC; The genes related to intramolecular substitution reactions include: hheA, hheB, and hheC; The genes related to hydrolytic dehalogenation reactions include: atzA, caa43, dcmA, dehH, dhaA, dhaAF, dhlA, dhlB, dhlC, dehI, dhmA, fac-dex, fcbB, hadL, HADs, linB, and RPA1163; The gene related to dehydrohalogenation reactions includes: linA; The genes related to methyltransferase dehalogenation reactions include: cmuA; The genes related to thiolytic dehalogenation reactions include: bbdE, bbdI, gstA, gstB, GST, pcpC, and TCHQD.
[0010] Based on personal experience and through literature research, the present invention obtained 59 genes related to dehalogenation reactions. When these genes related to dehalogenation reactions are combined with the aforementioned construction method, the constructed functional genomic database of dehalogenating microorganisms has relatively better comprehensiveness and accuracy.
[0011] Furthermore, in step (2), the core database is constructed by the following method: Obtain the seed sequences of the families to which the genes related to dehalogenation reactions belong based on the Swiss-Prot database in the Uniprot protein database; align the seed sequences with the TrEMBL database, and merge the homologous sequences obtained by the alignment with the seed sequences to obtain the core database; Preferably, search for the seed sequences by gene name in the Swiss-Prot database. The seed sequences include at least 150 amino acids, and ClustalW is used for sequence alignment and generation of a phylogenetic tree file during the search process; More preferably, in the TrEMBL database, search for homologous sequences by Usearch to obtain the ids of the homologous sequences, with a homology cut-off rate > 30%; based on the ids of the homologous sequences, extract the sequences by Seqkit.
[0012] Further, in step (3), the homologous protein database includes one or more of the arCOG, KOG, COG, eggNOG, and KEGG homologous protein databases; Preferably, the blastp module of Diamond is used to align the homologous protein database to obtain the sequence ids of the dehalogenation functional gene family. The alignment parameters include: E-value < 10-5, percentage of identical matches > 80, bitscore > 40; based on the sequence ids, sequences are extracted through Seqkit.
[0013] Further, in step (4), the Refseq database of microorganisms in NCBI includes the Refseq databases of archaea, bacteria, and fungi; Preferably, the alignment parameters of the preliminary database include: E-value < 10-5, percentage of identical matches > 80; More preferably, the complete taxonomic level information of each sequence is determined through TaxonKit; Even more preferably, the CD-hit software is used to set the parameter id = 1.0 for redundancy removal.
[0014] The present invention further optimizes each process and selects software that is stable and efficient for realizing the function, which helps to improve the comprehensiveness and accuracy of the dehalogenating microbial functional genome database constructed by the present invention.
[0015] In a second aspect, the present invention provides a dehalogenating microbial functional genome database, which is constructed by the foregoing construction method.
[0016] In a third aspect, the present invention provides the application of the foregoing dehalogenating microbial functional genome database in the identification or analysis of dehalogenating microbial functional communities.
[0017] In a fourth aspect, the present invention provides an analysis method for dehalogenating microorganisms, including: (1) Extracting the total DNA of the sample to be tested, and performing sequencing and quality control to obtain clean sequences; (2) Assembling the clean sequences to obtain contig sequences; (3) Performing ORFs prediction and clustering on the contig sequences to obtain a non-redundant gene set; (4) Based on the contig sequences, non-redundant high-quality MAGs are obtained through binning, redundancy removal, and quality control, and ORFs prediction is performed on the non-redundant high-quality MAGs to obtain MAGs protein sequences; (5) Align at least one of the cleaning sequence, the non-redundant gene set, or the MAGs protein sequence to the aforementioned dehalogenating microbial functional genomic database to obtain one or more of the following information in the test sample: i) Composition of microbial dehalogenation functional genes, ii) Abundance of microbial dehalogenation functional genes, iii) Composition of dehalogenating microorganisms, iv) Abundance of dehalogenating microorganisms.
[0018] Furthermore, the quality control in step (1) includes: Use FastQC to check the sequencing quality, and use KneadData to remove adapter sequences, low-quality sequences, and host sequences; and / or, The assembly in step (2) includes: Use Megahit to assemble the cleaning sequence and filter the overlapping sequences with a length less than 500 bp; and / or, In step (3), use Prodigal for ORFs prediction and use CD-Hit for clustering; and / or, In step (4), use metaWRAP encapsulating three binning methods of Maxbin2, metaBAT, and CONCOCT for binning, use dRep for redundancy removal, and retain non-redundant high-quality MAGs with a completeness ≥ 70% and a contamination rate ≤ 5%; and / or, In step (5), for the cleaning sequence, use Diamond blastx for alignment and calculation of the abundance of microbial dehalogenation functional genes, and use Kraken2 for species annotation; For the non-redundant gene set, use Diamond blastx for alignment, use Salmon for calculation of the abundance of microbial dehalogenation functional genes, and use Kraken2 and GTDB database for species annotation; For the MAGs protein sequence, use Diamond blastp for alignment to obtain MAGs containing dehalogenation functional genes, and use GTDB-tk to align the MAGs containing dehalogenation functional genes to the GTDBs database to obtain species information.
[0019] Even further, in step (1), use Trimmomactic in KneadData to remove adapter sequences and low-quality sequences, and use Bowtie2 to remove host sequences.
[0020] Furthermore, for the cleaning sequence, the obtained sequence depth statistics use the stats module of Seqkit software to normalize the relative abundance of microbial dehalogenation functional genes to RPKM.
[0021] Furthermore, for the non-redundant gene set, use coverM to delete sequences with an alignment sequence length < 75% and a nucleic acid identity < 95%, then use tpmean to calculate the average read depth, and normalize it according to the number of base pairs; the abundance information is combined with the species annotation information to obtain the final microbial classification.
[0022] The present invention has the following beneficial effects: Based on the analysis and search results of multiple databases, the present invention constructs an accurate dehalogenation functional gene database. And based on this dehalogenation functional gene database, combined with methods such as functional gene analysis and host information recognition, a bioinformatics data analysis process for a complete in-situ dehalogenating microbial functional community is established, which can accurately and effectively analyze the dehalogenating microbial functional community in the environment. Brief Description of the Drawings
[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic diagram of the method for constructing a dehalogenating microbial functional genome database and further analyzing dehalogenating microorganisms provided in Embodiment 1 of the present invention.
[0025] Figure 2 It is the comparison result of the accuracy rate, sensitivity, false negative rate, and false positive rate of the DehaloDB, KEGG, COG, KOG, eggNOG, and arCOG databases provided in Embodiment 2 of the present invention. Detailed Embodiments
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0027] The experimental methods involved in the following examples are all conventional methods in the art unless otherwise specified. For example, they can be referred to in the experimental manuals in the art or carried out according to the conditions recommended in the manufacturer's instructions.
[0028] The experimental materials and reagents involved in the following examples can be obtained from commercial sources unless otherwise specified.
[0029] Example 1 The present invention provides a method for constructing a functional genome database of dehalogenating microorganisms based on metagenomic samples from industrial wastewater treatment plants and further analyzing dehalogenating microorganisms (as Figure 1 shown), including the following processes: 1. Construction of the functional genome database of dehalogenating microorganisms (1) According to literature research, referring to the microbial dehalogenation reaction processes provided in the KEGG database, all genes involved in seven dehalogenation reactions were summarized.
[0030] These dehalogenation reaction processes include: reductive dehalogenation, oxidative dehalogenation, hydrolytic dehalogenation, thiolytic dehalogenation, dehydrohalogenation, methyl transfer dehalogenation, and intramolecular substitution reaction dehalogenation.
[0031] All genes of the seven dehalogenation reactions obtained include: reductive dehalogenation reaction: cfrA, cprA, cprA5, dcaA, dcpA, dcrA, DebcprA, linD, pceA, vcrA; oxidative dehalogenation reaction: bnz, bphA, bphB, bphC, bphD, cbdA, cbdB, cbdC, clcA, cpo, GLG, hadA, LIP, LPOA, mnp, pcpB, tcbC, tecA, tfdC; intramolecular substitution reaction: hheA, hheB, hheC; hydrolytic dehalogenation reaction: atzA, caa43, dcmA, dehH, dhaA, dhaAF, dhlA, dhlB, dhlC, dehI, dhmA, fac-dex, fcbB, hadL, HADs, linB, RPA1163; dehydrohalogenation reaction: linA; methyl transfer dehalogenation reaction: cmuA; thiolytic dehalogenation reaction: bbdE, bbdI, gstA, gstB, GST, pcpC, TCHQD.
[0032] (2) Construction of seed sequences: Obtain the seed sequences of the gene families corresponding to all genes of the aforementioned seven dehalogenation reactions from the Swiss-Prot database in the Uniprot protein database, search by gene name, and the selected sequences include at least 150 amino acids. During the search process, ClustalW is used for sequence alignment and a phylogenetic tree file is generated; The seed sequence was aligned with the TrEMBL database in the Uniprot protein database. The Usearch software was used to search for homologous sequences of the seed sequence in TrEMBL. The homologous cutoff rate was set at id = 30%, and a file containing homologous sequence ids was obtained. The homologous sequences were extracted according to the ids using the Seqkit software, and the obtained homologous sequences were downloaded and merged with the seed sequence to obtain the Core database.
[0033] (3)The Core database was aligned with the existing homologous protein databases arCOG, KOG, COG, eggNOG, and KEGG. The blastp module of Diamond was used for sequence alignment. The alignment parameters were set to meet E-value < 10 -5 , percentage of identical matches > 80, and bitscore > 40 to screen the sequences to obtain the sequence ids belonging to the dehalogenation functional gene family. Subsequently, Seqkit was used again for sequence extraction and merged with the Core database to form the pre-full database.
[0034] (4)The pre-full database was aligned to the Refseq database of microorganisms in NCBI (E-value < 10 -5 , percentage of identical matches > 80), the taxonomic generalization of the pre-full database was classified to determine the coverage of dehalogenation functional species, and the TaxonKit was used to determine the complete taxonomic level information of each sequence. Homologous sequences were identified, extracted, and merged to generate the Full database.
[0035] (5)The obtained sequences were de-redundant using the CD-hit software with the parameter id = 1.0 to form the DehaloDB.
[0036] 2. Sample collection and pretreatment Different types of industrial wastewater treatment plants in a certain industrial park were selected as the research objects. Samples of influent water, the mud-water mixture in the biological unit, and the effluent from the secondary sedimentation tank were collected. A 0.22 μm mixed cellulose ester filter membrane (Millipore, Germany) was used to filter and intercept the DNA in the water sample. After obtaining 3 - 4 filter membranes at each sampling point, they were stored in a -80 °C refrigerator for DNA extraction and sequencing.
[0037] 3. DNA extraction and metagenomic sequencing The 0.22-μm filter membrane was cut into squares with a size of 2-3 mm and transferred into the DNA extraction tube of the DNA extraction kit. The PowerSoil DNA Isolation Kit (Mo Bio, USA) was used for DNA extraction, and the specific operation was referred to the operation manual of the kit. The qualified DNA samples were used for library construction. For the qualified libraries, the Illumina HiSeq 2500 high-throughput sequencing platform was used for PE150 sequencing to obtain the raw sequencing sequences (Raw reads).
[0038] 4. Metagenomic data analysis 4.1 Quality control, assembly and binning of metagenomic sequences (1) First, the FastQC (v0.11.9) software was used to check the overall quality of the raw reads, and the KneadData process was used for sequence quality control, including using the Trimmomactic (v0.39) software to remove the paired-end adapter sequences and low-quality sequences of the sequences (parameter settings: SLIDINGWINDOW:4:20 MINLEN:50), and the Bowtie2 (v2.3.6) software to remove human host contamination.
[0039] (2) The software MEGAHIT (v1.2.9) was used to perform de novo assembly on the clean reads (default parameter settings) to obtain the contig sequences, and the contigs with a splicing length of more than 500 bp were screened for subsequent analysis.
[0040] (3) The software Prodigal (v2.6.3) was used to predict the ORFs of the contigs (parameter settings: -p meta), and the software CD-Hit (v4.8.1) was used to cluster the obtained ORFs (parameter settings: -aS 0.9 -c 0.95) to obtain a non-redundant gene set; the software Salmon (v1.4.0) was used to calculate the reads count of each gene and further standardize it to obtain the TPM value, which was used to characterize the abundance information of the gene in each sample.
[0041] (4)Based on the quality-controlled clean reads and assembled contigs, the software metaWRAP (v1.3.2) (which encapsulates three binning methods: Maxbin2, metaBAT, and CONCOCT) was used to extract and purify the draft microbial genomes (MAGs) in each sample, and the gene sequences of MAGs were predicted. The screening threshold for MAGs was set as contamination ≤ 5% and completeness ≥ 70%. The software dRep (v3.5.0) was used to dereplicate the obtained MAGs (parameter settings: -sa 0.95, -nc 0.30).
[0042] 4.2 Annotation of metagenomic sequencing data by the DehaloDB database: (1)Short sequence annotation and abundance calculation: The paired-end sequences were aligned to the constructed DehaloDB database using the blastx algorithm of the Diamond software. The specific parameter settings were E-value < 10 -10 , percentage of identical matches > 80. The sequence depth statistics of the obtained sequences were calculated using the stats module of the Seqkit software, and the relative abundance of dehalogenation functional genes was normalized to RPKM (reads per kilobase per million mapped reads). The obtained sequences were annotated for species using Kraken2 (either the full Kraken2 database or the GTDB database was selected).
[0043] (2)Assembly sequence annotation and abundance calculation: The protein sequences of the non-redundant gene set were aligned with the DehaloDB database using the blastp algorithm of the Diamond software. The selected parameters were the same as those described in (1). Species annotation was performed using Kraken2; sequences with an alignment sequence length < 75% and nucleic acid identity < 95% were removed using coverM, and then the average read depth was calculated using tpmean in it and normalized according to the number of base pairs. The abundance information was combined with the species annotation information to obtain the final classification.
[0044] (3)Genome sequence annotation and abundance calculation: The obtained MAGs were aligned with the DehaloDB database using the blastp algorithm of the Diamond software. The selected parameters were the same as those in (1) above. The species information was obtained by aligning it with the GTDBs database using GTDB-tk. The relative abundance and coverage of MAGs were calculated using coverM.
[0045] Example 2 The present invention constructs an artificial protein sequence dataset. 400 protein sequences are collected from the NCBI database, including 150 microbial dehalogenation functional protein sequences, 50 non-target homologous gene protein sequences, and 200 protein sequences that are completely unrelated to the dehalogenation reaction. The artificial dataset is used to calculate the false-positive rate and false-negative rate. A sequence that is misannotated as a certain functional gene is regarded as a false positive, and a sequence that is not annotated but actually has the target functional gene is regarded as a false negative. The artificial dataset is searched in COG, arCOG, KOG, and DehaloDB using Diamond, with an E-value ≤ 10 -5 , sequence identity > 40%, and searched in the eggNOG database using eggNOG-mapper, with an E-value ≤ 10 -5 , and the artificial dataset is searched in the KEGG database using the functional annotation tool KofamKOALA provided by the KEGG database, with an E-value ≤ 10 -5 , to compare the annotation accuracies of these databases.
[0046] The results are as Figure 2 shown: The present invention evaluates the DehaloDB, KEGG, COG, KOG, eggNOG, and arCOG databases using the artificial dataset. The sensitivity evaluation results show that the DehaloDB database correctly annotates all sequences belonging to the dehalogenation function into the dehalogenation functional gene family; the accuracy results show that almost all dehalogenation functional sequences are annotated into the correct dehalogenation functional gene family, and other non-dehalogenation functional genes and non-target homologous gene sequences are not annotated (accuracy is 98.3%).
[0047] The public orthologous databases are not as good as DehaloDB in annotating dehalogenation functional genes. Among them, the annotation accuracy of the eggNOG database is the highest among the public orthologous databases, but it does not exceed 50%, only 47.23%, and it has a false-negative rate of 50.67%; the sensitivity of the COG database is very high (75.33%), but many homologous sequences are misidentified as dehalogenation functional genes because its false-positive rate is very high. The KOG and arCOG sequences of fungi and archaea have low sensitivity and accuracy for the dehalogenation function. Therefore, DehaloDB provided in Example 1 of the present invention has higher sensitivity and accuracy than other orthologous databases in the analysis of microbial dehalogenation functional genes.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a functional genomic database of dehalogenating microorganisms, characterized in that, Including: (1) Obtaining genes related to dehalogenation reactions; (2) Based on the Uniprot protein database, obtaining the seed sequences and homologous sequences of the family to which the genes related to dehalogenation reactions belong to obtain a core database; (3) Comparing the core database with a homologous protein database, and merging the sequences of the dehalogenation function gene family obtained by the comparison with the core database to obtain a preliminary database; (4) Comparing the preliminary database to the Refseq database of microorganisms in NCBI, and obtaining a complete database through classification, merging homologous sequences, and removing redundancy.
2. The construction method according to claim 1, wherein In step (1), the obtaining of genes related to dehalogenation reactions includes: Referring to the microbial dehalogenation reaction process in the KEGG database, obtaining genes related to dehalogenation reactions through literature research; the dehalogenation reactions include one or more of reductive dehalogenation, oxidative dehalogenation, hydrolytic dehalogenation, sulfuryl dehalogenation, dehydrohalogenation, methyl transfer dehalogenation, or intramolecular substitution reaction dehalogenation.
3. The construction method according to claim 2, wherein, The genes related to dehalogenation reactions include: Genes related to reductive dehalogenation reactions, genes related to oxidative dehalogenation reactions, genes related to intramolecular substitution reactions, genes related to hydrolytic dehalogenation reactions, genes related to dehydrohalogenation reactions, genes related to methyl transfer dehalogenation reactions, and genes related to sulfuryl dehalogenation reactions; The genes related to reductive dehalogenation reactions include: cfrA, cprA, cprA5, dcaA, dcpA, dcrA, DebcprA, linD, pceA, and vcrA; The genes related to oxidative dehalogenation reactions include: bnz, bphA, bphB, bphC, bphD, cbdA, cbdB, cbdC, clcA, cpo, GLG, hadA, LIP, LPOA, mnp, pcpB, tcbC, tecA, and tfdC; The genes related to intramolecular substitution reactions include: hheA, hheB, and hheC; The genes related to hydrolytic dehalogenation reactions include: atzA, caa43, dcmA, dehH, dhaA, dhaAF, dhlA, dhlB, dhlC, dehI, dhmA, fac-dex, fcbB, hadL, HADs, linB, and RPA1163; The genes related to dehydrohalogenation reactions include: linA; The genes related to methyl transfer dehalogenation reactions include: cmuA; The genes related to sulfuryl dehalogenation reactions include: bbdE, bbdI, gstA, gstB, GST, pcpC, and TCHQD.
4. The construction method according to any one of claims 1-3, characterized in that, In step (2), the core database is constructed by the following method: Based on the Swiss-Prot database in the Uniprot protein database, obtaining the seed sequences of the family to which the genes related to dehalogenation reactions belong; comparing the seed sequences with the TrEMBL database, and merging the homologous sequences obtained by the comparison with the seed sequences to obtain a core database; Search for the seed sequence in the Swiss-Prot database by gene name. The seed sequence includes at least 150 amino acids. During the search, use ClustalW to align the sequences and generate a phylogenetic tree file. In the TrEMBL database, search for homologous sequences by Usearch to obtain the ids of homologous sequences, with a homology cut-off rate > 30%. Based on the ids of the homologous sequences, extract the sequences by Seqkit.
5. The construction method according to any one of claims 1-3, characterized in that In step (3), the homologous protein database includes one or more of the following: arCOG, KOG, COG, eggNOG, KEGG homologous protein databases. The sequence IDs of the dehalogenation functional gene family were obtained by aligning with the homologous protein database using the blastp module of Diamond. The alignment parameters included: E-value < 10 -5 , percentage of identical matches > 80, bitscore > 40; Based on the sequence IDs, sequences were extracted using Seqkit.
6. The construction method according to any one of claims 1-3, characterized in that In step (4), the Refseq database of microorganisms in NCBI includes the Refseq databases of archaea, bacteria, and fungi. The comparison parameters of the preliminary database include: E-value < 10 -5 , percentage of identical matches > 80; Determine the complete taxonomic level information of each sequence by TaxonKit. Use the CD-hit software to set the parameter id = 1.0 for redundancy removal.
7. A dehalogenating microorganism functional genomics database, characterized in that, The dehalogenating microorganism functional genome database is constructed by the construction method according to any one of claims 1-6.
8. Use of the dehalogenating microorganism functional genome database according to claim 7 in the identification or analysis of dehalogenating microorganism functional communities.
9. A method for analyzing a dehalogenating microorganism, characterized in that, Including: (1) Extract the total DNA of the test sample, perform sequencing and quality control to obtain clean sequences. (2) Assemble the clean sequences to obtain contig sequences. (3) Perform ORFs prediction and clustering on the contig sequences to obtain a non-redundant gene set. (4) Based on the contig sequences, obtain non-redundant high-quality MAGs through binning, redundancy removal, and quality control. Perform ORFs prediction on the non-redundant high-quality MAGs to obtain MAGs protein sequences. (5) Align at least one of the clean sequences, non-redundant gene set, or MAGs protein sequences to the dehalogenating microorganism functional genome database according to claim 7 to obtain one or more of the following information in the test sample: i) Composition of microbial dehalogenation functional genes, ii) Abundance of microbial dehalogenation functional genes, iii) Composition of dehalogenating microorganisms, iv) Abundance of dehalogenating microorganisms.
10. The method according to claim 9, wherein In step (1), the quality control includes: Use FastQC to check the sequencing quality, and use Kneaddata to remove adapter sequences, low-quality sequences, and host sequences; and / or, In step (2), the assembly includes: Use Megahit to assemble the clean sequences and filter out overlapping sequences with a length less than 500 bp; and / or, In step (3), use Prodigal for ORFs prediction and CD-Hit for clustering; and / or, In step (4), use metaWRAP, which encapsulates three binning methods, Maxbin2, metaBAT, and CONCOCT, for binning, use dRep for redundancy removal, and retain non-redundant high-quality MAGs with a completeness ≥ 70% and a contamination rate ≤ 5%; and / or, In step (5), for the said cleaning sequence, Diamond blastx is used for alignment and calculation of the abundance of microbial dehalogenation functional genes, and Kraken2 is used for species annotation; For the non-redundant gene set, Diamond blastx is used for alignment, Salmon is used for calculation of the abundance of microbial dehalogenation functional genes, and Kraken2 and GTDB database are used for species annotation; For MAGs protein sequences, Diamond blastp is used for alignment to obtain MAGs containing dehalogenation functional genes, and GTDB-tk is used to align the MAGs containing dehalogenation functional genes with the GTDBs database to obtain species information.
Citation Information
Patent Citations
Method for analyzing microbial population function by using metagenome data
CN108804875A
Method for constructing dissimilatory arsenate reductase protein database
CN110310708A
Microbial species and functional composition analysis method for metagenome sequencing data
CN112599198A
Microbial metagenome sequencing analysis method and system based on Hi-C
CN118782149A
A browsable database for biological use
WO2004053769A2