A method for predicting the abundance of quorum-sensing genes in microbial populations
By constructing the QSG database and hidden Markov model, the shortcomings of the existing database in the prediction of induction gene abundance of microbial populations are solved, accurate prediction and functional evaluation of gene abundance are achieved, and annotation efficiency is improved.
Patent Information
- Application Number
- CN202210484869.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-06
AI Technical Summary
The existing population sensing-related databases have problems such as limited gene count, many missing entries in the database, and large analysis errors when predicting the abundance of microbial population sensing genes, which is difficult to meet the needs of high-throughput sequencing technology.
A QSG database was constructed, and a population sensing related database with complete gene species and rich sequences was established through literature classification, cluster analysis, hidden Markov model and protein conserved site comparison, and a population sensing related database was established with the hidden Markov model for gene abundance prediction, and protein sequence annotation and alignment were combined with Diamond and Hmmer software.
Accurate prediction of the abundance of microbial population sensing genes is achieved, the efficiency of annotation results is improved, the microbial population sensing function can be more comprehensively evaluated, the processing process is simplified, and the annotation results are enriched by 7 times.
Smart Images

Figure CN114842910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microbiology, and more particularly to a method for predicting the abundance of microbial quorum sensing genes. Background Art
[0002] Bacteria were once generally regarded as simple cellular life forms with only single physiological behaviors. However, more and more studies have confirmed that bacteria have multicellular group cooperative behaviors. These cooperative behaviors include processes such as bioluminescence, production of virulence factors, production of secondary metabolites, gene upregulation, and formation of biofilms. These processes often require individual cells to coordinate within the group. To better coordinate group behaviors, the process of cell-to-cell communication utilized by bacteria is called quorum sensing. The quorum sensing signals that have been identified mainly include acyl-homoserine lactones (HSLs), other autoinducer-2 (AI-2), diffusible signal factor (DSF), α-hydroxy ketones (AHKs), cyclic diguanylate monophosphate (c-di-GMP), Pseudomonas quinolone signal (PQS), and autoinducing peptide (AIP), etc. Quorum sensing is mainly regulated by the following processes: production, release, accumulation, and population-wide detection of autoinducing signals. Specifically, in a specified bacterial population, individual cells express specific genes to produce signal molecules, and the signal molecules are continuously accumulated. Once a certain concentration is reached, the signal molecules can be sensed and regulate the expression of corresponding target genes.
[0003] With the development of high-throughput sequencing technology, it has become a challenge to predict the production, reception, and quenching capabilities of signaling molecules in a community by detecting the related genes of signaling molecules in the community. Conventional methods for predicting functional genes are based on metagenomic sequencing or specific functional gene probe detection, and these methods require a large number of quorum sensing gene sequences. However, there are very limited reports on quorum sensing-related databases. The first SigMol database (http: / / bioinfo.imtech.res.in / manojk / sigmol) for quorum sensing signals was published in 2015. Starting from signaling molecules, it classified 2,900 articles, including 182 unique signaling molecules and their related genes from 215 organisms. In 2020, Du et al. published the QSDB database (https: / / github.com / qhmu / QSDB), which was modified according to the quorum sensing-related classification groups in the KEGG (Kyoto Encyclopedia of Genes and Genomes) database and included 39,981 related sequences, covering seven quorum sensing systems. However, in actual use, many gene entries were missing in this database, and the number of genes was limited, making it difficult to predict the abundance of quorum sensing genes in microorganisms. In 2021, the quorum sensing database QSDB proposed by Klein et al. showed the sensing and quenching relationships between microorganisms related to the human microbiome and signaling molecules and provided a graphical description of the sensing mechanism, but it was not applicable to the processing of metagenomes and could not perform gene abundance prediction. Chinese Patent Application No.: CN2020111179792, titled: A method for analyzing the quorum sensing effect of microorganisms based on metagenomic data, only obtained 891 HdtS sequences from the non-redundant protein database, and the number of obtained sequences was limited, resulting in a large error in analyzing the abundance of quorum sensing-related genes in the community. For other articles using metagenomic methods for analysis, the work was complex and repetitive and the basic sequences were not publicly available. Therefore, the present invention optimizes the method for analyzing quorum sensing-related genes in high-throughput sequencing technology, establishes a quorum sensing-related database with sufficient sequences, and predicts the abundance of quorum sensing-related genes in microorganisms through it, providing a powerful tool for analyzing the abundance of quorum sensing-related genes in the community. Summary of the Invention
[0004] The present invention provides a method for predicting the abundance of quorum sensing-related genes in microorganisms using a database to solve the technical problems of aligning and annotating quorum sensing target genes and providing an assessment of the capabilities related to the quorum sensing function of microorganisms.
[0005] To achieve the above object, the present invention provides a method for predicting the abundance of quorum sensing genes in microorganisms, including the following steps:
[0006] S101. Obtain the literature related to quorum sensing in the first database;
[0007] S102. Classify the literature related to quorum sensing to form a quorum sensing literature database with signals and corresponding genes, and obtain the sequences related to quorum sensing genes;
[0008] S103. Obtain highly certain sequences in the second database;
[0009] S104. Merge the sequences related to quorum sensing genes and the highly certain sequences, remove duplicate sequences, number them according to "signal - major gene category - minor gene category - serial number", then perform clustering analysis and remove redundant sequences to form seed sequences;
[0010] S105. Align the protein conserved sites of the seed sequences and construct a hidden Markov model of the protein conserved sites;
[0011] S106. Align the hidden Markov model of the conserved sites in the third database. For the sequences that meet the conserved regions, screen the found potential functional gene sequences using NR annotation information and credibility, and retain the sequences with E - value ≥ 1e -5 and the protein sequences obtained by searching for specific signal molecules, gene names, or functional descriptions in the annotation;
[0012] S107. Screen all the protein sequences obtained in S106, extract the proteins containing highly credible sequence keywords related to quorum sensing in the annotation to form a highly credible sequence database; according to the search results, use the Seqkit software to extract the sequences in the NR database, add the word "gene name %" to the download results for convenient subsequent processing; merge all the sequences to form a quorum sensing sequence database named QSG database;
[0013] S108. Statistically calculate the sequence recovery quantity and gene recovery rate corresponding to each gene to construct a standardized database; through the sequences in the QSG database, use the Hmmsearch program to adjust the T - value to 20 - 300 for searching respectively; calculate the sequence recovery quantity and gene recovery rate corresponding to each gene respectively, take the T - value with a recovery rate greater than 70% and the best recovery quantity as the optimal value, and write the optimal value into the hidden Markov model (HMM model).
[0014] The formula for the recovery rate is:
[0015] where N 目的基因 represents the number of highly credible sequences belonging to the gene in the search results, and N 搜索结果 represents the total number of search result sequences.
[0016] Write the optimal T value in all models into the HMM model, merge the texts using the cat command, and use the Hmmpress program to build a standardized database for the merged model. Subsequently, the Hmmscan program can be used for searching.
[0017] S109. Annotate and align the protein sequences.
[0018] Diamond alignment method: Use the Blastp program in the Diamond software to perform alignment analysis on the samples. The recommended parameters are "-e 0.00001 --id 50", that is, retain those with e-value < 1e -5 and identify ≥ 50%; The merge function in R language can be used to process and statistically analyze the results, and then the final protein sequence annotation results can be obtained.
[0019] Hmmer alignment method: Use the Hmmscan program in the Hmmer software to align the merged HMM data model described in S31 with the samples, using the parameter - cut_ga, that is, use the optimal parameters set in S31; The obtained results can be statistically classified using the grep function by adding the words "gene name %" according to S21, and then the annotation results of the protein sequences can be obtained.
[0020] Preferably, the first database is the PubMed database of the National Center for Biotechnology Information in the United States; the second database is the Uniport, KEGG, and QSDB databases of the National Center for Biotechnology Information in the United States; the third database is the non-redundant database (NR) of the National Center for Biotechnology Information (NCBI) in the United States (https: / / www.ncbi.nlm.nih.gov / ).
[0021] Preferably, the quorum sensing literature library of signals and corresponding genes described in step S102 includes 6 signal molecules and 24 related gene sequences related to them; the gene sequences include synthase, receptor enzyme, and degradation enzyme directly related to the signal molecule.
[0022] Preferably, the highly certain sequence described in step S103 is a sequence that contains related genes or functions and whose protein length is 80% - 140% of the average length of the sequences reported in the literature.
[0023] Preferably, the method for removing duplicate sequences described in step S104 includes retaining only one sequence with a similarity greater than 90%.
[0024] Preferably, the optimization method of the hidden Markov model of the protein conserved sites in step S105 includes using the Hmmsearch program to search for the conserved regions of the protein seed sequences and screening the score values; adjusting the T parameter to be between 20 and 300 respectively, with the adjustment step size of the T parameter being 5, and all sequences with scores exceeding the scoring T parameter will be retained.
[0025] Preferably, the protein sequence annotation comparison method in step S109 includes formatting the input protein sequence as fasta, performing Diamond similarity alignment using the seed sequence, and retaining the results with e-value < 1e -5 and identify ≥ 50%.
[0026] Preferably, the protein sequence annotation comparison method in step S109 includes using the Hmmscan function of the Hmmer software to perform annotation alignment on the protein sequence based on the constructed conserved site model, with the parameter being —cut_ga.
[0027] The beneficial effects of the present invention are as follows:
[0028] The present invention establishes a conserved site model for quorum sensing genes, constructs a quorum sensing-related gene database with complete gene types and rich sequence numbers based on this, and provides a method for predicting the abundance of quorum sensing-related genes in microorganisms using the hidden Markov model or database. Through the constructed database, more comprehensive quorum sensing-related genes in the microbial community can be annotated, which can be further applied to environmental samples to compare and annotate the quorum sensing target genes, provide an assessment of the quorum sensing function-related capabilities of microorganisms, and provide technical support for the treatment of actual samples.
[0029] The present invention collects the key genes related to the signal action process for six relatively well-studied signal molecules, constructs a hidden Markov model, establishes a systematic quorum sensing-related gene database, can evaluate the homology alignment and annotation of target proteins, and is further applied to actual environmental samples to perform homology alignment and annotation on target genes, and provide predictive analysis of the signal generation, sensing, and quenching capabilities of microbial quorum sensing.
[0030] The method provided by the present invention can obtain more novel relevant gene information through literature classification, and each gene can be traced in the literature database, which is convenient for personalized in-depth research. The hidden Markov model of quorum sensing genes constructed for the first time in the present invention simplifies the process of layer-by-layer processing in the prior art and finally obtains 159,893 sequences. Directly using the protein conserved sites for alignment and annotation enriches the annotation results, and the annotation efficiency is improved by 7 times compared with the prior art. Therefore, the method of the present invention can predict more proteins with potential functions but incomplete annotation information in the non-redundant protein data using the hidden Markov model. Brief Description of the Drawings
[0031] Figure 1 is a schematic diagram of the QSG database construction process of the present invention;
[0032] Figure 2 is a schematic diagram of the QSG literature index database of the present invention;
[0033] Figure 3 is a schematic diagram of the QSG database structure of the present invention;
[0034] Figure 4 is the QSG database sequence file of the present invention;
[0035] Figure 5 is the hidden Markov model of the present invention;
[0036] Figure 6 is the Hmmsearch search result of the present invention;
[0037] Figure 7 is the annotation comparison result of the Diamond software of the present invention;
[0038] Figure 8 is the Hmmscan search result of the present invention;
[0039] Figure 9 is the comparison of the results of different annotation methods of the present invention. Detailed Embodiments
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0041] A method for predicting the abundance of quorum sensing genes in microorganisms, comprising the following steps:
[0042] S101. Obtain quorum sensing-related literature in the first database;
[0043] S102. Classify the quorum sensing-related literature to form a quorum sensing literature library of signals and corresponding genes, and obtain sequences related to quorum sensing genes;
[0044] S103. Obtain highly certain sequences in the second database;
[0045] S104. Combine the sequences related to quorum sensing genes and the highly certain sequences, remove duplicate sequences, number them according to "signal - major gene category - minor gene category - serial number", and then perform clustering analysis to remove redundant sequences to form seed sequences;
[0046] S105. Align the protein conserved sites of the seed sequences, and construct a hidden Markov model of the protein conserved sites with the conserved sequences;
[0047] S106. Align the hidden Markov model of the conserved site in the third database, and according to E-value ≥ 1e -5 Retain the obtained protein sequences;
[0048] S107. Screen the obtained protein sequences in S106, and extract the proteins containing highly credible sequence keywords related to quorum sensing in the annotation to form a highly credible sequence database;
[0049] S108. Count the sequence recovery number and gene recovery rate corresponding to the gene, and construct a standardized database;
[0050] S109. Annotate and align the protein sequences.
[0051] The first database is the PubMed database of the National Center for Biotechnology Information of the United States; the second database is the Uniport, KEGG and QSDB databases of the National Center for Biotechnology Information of the United States; the third database is the non-redundant database of the National Center for Biotechnology Information of the United States.
[0052] Among them, the quorum sensing literature library of the signal and the corresponding gene in step S102 includes 6 signal molecules and 24 related gene sequences related to them; the gene sequences include synthase, receptor enzyme and degrading enzyme directly related to the signal molecule.
[0053] Among them, the highly deterministic sequence in step S103 is a sequence that contains the relevant gene or function and whose protein length is 80% - 140% of the average length of the literature-reported sequence.
[0054] Among them, the method for removing duplicate sequences in step S104 includes retaining only one sequence with a similarity greater than 90%.
[0055] Among them, the optimization method of the hidden Markov model of the protein conserved site in step S105 includes using the Hmmsearch program to search for the conserved region of the protein seed sequence and screening the score value; adjusting the search threshold to be between 20 and 300, modifying the step size to 5, and retaining the sequences exceeding the scoring search threshold.
[0056] Among them, the calculation formula for the gene recovery rate in step S108 is:
[0057]
[0058] Among them, N 目的基因 represents the number of highly credible sequences belonging to the gene in the search results, and N 搜索结果 represents the total number of search result sequences.
[0059] The protein sequence annotation comparison method in step S109 includes inputting the protein sequence in the format of fasta, using the seed sequence for Diamond similarity comparison, and retaining e-value < 1e -5 And identify ≥ 50% results.
[0060] The protein sequence annotation comparison method in step S109 includes using the Hmmscan function of the Hmmer software to perform annotation comparison on the protein sequence based on the constructed conservative site model using the parameter -cut_ga.
[0061] Example 1
[0062] like Figure 1 As shown, the present invention constructs a database process, and the specific steps are as follows:
[0063] S101. Search all literature related to quorum sensing in the PubMed database of the National Center for Biotechnology Information (NCBI).
[0064] The quorum sensing literature library involved in the present invention retrieved all documents related to the keyword "quorum" from the PubMed database in NCBI, and a total of 12,596 documents were collected as of October 30, 2021.
[0065] S102. Use Endnote software to manually classify the literature, and eventually form a quorum sensing literature library based on signals and corresponding genes, which includes 6 signal molecules and 24 related genes.
[0066] As attached Figure 2 As shown, a partial screenshot of the literature index database of the present invention, after removing 2167 irrelevant documents, a total of 4722 documents were screened according to the classification of signal molecules. The literature was classified mainly according to the abstract information and the full text content using Endnote software, including 6 signal molecules, namely: HSLs, AI-2, DSF, AHKs, c-di-GMP and PQS, including 36 "group sets", a total of about 580 "groups". Finally, the literature index database was constructed through manual inspection and screening.
[0067] S103. According to the classification of the literature in the quorum sensing literature library, keywords of signal-related enzymes were determined. The keywords can be viewed in the literature index database. Search for the keywords in databases such as the Universal Protein database (Uniport database), Nucleotide Sequence database (NR database), QSDB database, Kyoto Encyclopedia of Genes and Genomes database (KEGG database), etc., and download sequences with relatively high certainty according to whether the annotation information contains relevant genes or functions and the protein length is 80% - 140% of the average length of the reported sequences in the literature.
[0068] As shown in the appendix Figure 3 As shown, the quorum sensing sequence database involved in the present invention is mainly constructed around 6 signal molecules and contains 24 gene sequences related to the 6 signal molecules. The gene sequences contain synthase, receptor enzyme and degrading enzyme directly related to the signal molecules. The signal molecules are HSLs, AI-2, DSF, AHKs, c-di-GMP and PQS. Among them, no directly degrading enzyme was found for the AHKs signal, so only synthase and receptor enzyme are included.
[0069] As shown in the appendix Figure 4 As shown, the number of the 24 signal-related gene sequences is: LuxI (5399), LuxM (1051), HdtS (159893), LuxR (139017), LuxN (2864), Acylase (57116), Lactonases (13954), LuxS (20083), LsrB (7835), LuxP (626), dCACHE domain (18245), LsrG (274), RpfF (31506), RpfC (745), RpfB (95413), PqsH (252), PqsR (296), pqsD (23936), AqdC (69), CqsA (1846), CqsS (158), DGC (448252), PDE (150456) and Clp (35). A total of about 1,169,320 sequences are included, and each sequence contains "> gene name %".
[0070] S104. Merge all the downloaded sequences, use the Seqkit software to remove duplicate sequences, and use a self-written python program to number them according to "signal - major gene category - minor gene category - serial number". Use the CD-HIT software to cluster and remove redundancy of all the sequences to form seed sequences.
[0071] S105. Use the Muscle software to align the protein conserved sites of the seed sequences, and use the HMMER 3.0 software to construct a hidden Markov model of the protein conserved sites from the conserved sequences.
[0072] Optimization of the protein conserved site model: As shown in the appendix Figure 5 The protein conserved site model designed in the present invention is a hidden Markov model constructed using the Hmmbuild program based on the collected protein seed sequences. As shown in the appendix Figure 6 shown, use the Hmmsearch program to search for the conserved regions of the 24 signal-related protein seed sequences. Hmmsearch can screen the score values in the results. Therefore, adjust the "-T" parameter in Hmmsearch. Adjust the T parameter to be between 20 and 300, with a step size of 5 for the T parameter adjustment. The results will retain all sequences with a score exceeding the scoring T parameter. The software will score each search result and retain the sequences with a search result exceeding the T parameter.
[0073] S106. Use the Hmmsearch function to align the hidden Markov model in the NR database, and retain the searched protein sequences according to the E-value greater than or equal to 1e -5 Retain the searched protein sequences.
[0074] S107. Screen the protein sequences obtained by the search in S106, extract the proteins with highly credible sequence keywords related to quorum sensing in the annotation as highly credible sequences, and obtain the QSG database constructed in the present invention. The highly credible sequence keywords are shown in Table 1 below.
[0075] Statistically count the sequence recovery quantity and gene recovery rate corresponding to each gene respectively. Take the search threshold T value with a recovery rate greater than 80% and the optimal recovery quantity as the optimal value, and write the optimal value into the HMM model. The final optimal T value is shown in Table 2 below.
[0076] The formula for the recovery rate is:
[0077]
[0078] Among them, N target gene represents the number of highly credible sequences belonging to the gene in the search results, and N search results represents the total number of search result sequences.
[0079] Write the optimal T values in all models into the HMM model, merge the texts using the cat command, and use the Hmmpress program to construct a standardized database for the merged model. Subsequently, the Hmmscan program can be used for searching.
[0080] Protein sequence annotation and alignment
[0081] The quorum sensing sequence database involved in the present invention can annotate and align protein sequences. The input protein sequence format is in fasta format. Diamond similarity alignment is performed using seed sequences, and results with e-value < 1e-5 and identify ≥ 50% are retained.
[0082] In addition, the protein conserved site model involved in the present invention can also annotate and align protein sequences. The Hmmscan function of the Hmmer software is used to annotate and align protein sequences with the parameter of —cut_ga.
[0083]
[0084]
[0085] Table 1. Keywords for screening highly credible sequences
[0086]
[0087] Table 2. Optimization results of the HMM model
[0088] The protein sequences of the annotation results of two alternative methods for annotating the same sample are annotated and aligned. The input protein sequence format is in fasta format. Diamond similarity alignment is performed using seed sequences, and results with e-value < 1e-5 and identify ≥ 50% are retained. The annotation and alignment results are shown in Figure 7 . In addition, using the Hmmscan function of the Hmmer software, the protein sequences are annotated and aligned based on the constructed conserved site model with the parameter of —cut_ga, and the results are shown in Figure 8 . Both methods can predict a large number of quorum sensing-related proteins. Among them, Diamond alignment can obtain more results and is better.
[0089] As shown in the appendix Figure 9As shown, it is the result of comparing the QSDB database sequences reported in the prior art, the protein conserved site model described in the present invention, and the sequence database described in the present invention for the same sample. Twenty-four signal-related genes included in the database were annotated. Among them, the Diamond parameters were all selected as e-value < 1e-5 and identify ≥ 50%. In addition, the annotation results using the Hmmscan method were included, with the parameter being —cut_ga. The results show that compared with the results obtained by annotating the reported QSDB database (a total of 4,197 sequences), the annotation results that can be obtained by the method of the present patent for most genes are as follows: a total of 10,771 sequences were obtained by alignment with the Hmmer software, and a total of 33,550 sequences were obtained by alignment with the Diamond software. In terms of the total number, the sequences obtained by alignment with the Diamond software are 7.99 times the sequences predicted by the prior art QSDB database. The alignment results using the Hmmer and Diamond software show that the Diamond alignment method has more advantages for genes such as DGC, RpfF, dCACHE, LuxS, and CqsS, while the Hmmer software can annotate more results for the CqsA and LuxI genes.
[0090] Compared with the prior art, the present invention can predict more proteins, and more results are obtained by the Diamond method. This is because during the optimization process of the built-in parameters of the HMM model, only the genes with more complete annotation information are retained, and the proteins with incomplete annotation in the NR database are ignored. Users can adjust the parameters according to their needs.
[0091] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A method for predicting the abundance of microbial quorum sensing genes, characterized in that, It includes the following steps: S101. Obtain the literature related to quorum sensing in the first database; S102. Classify the literature related to quorum sensing, form a quorum sensing literature library with signals and corresponding genes, and obtain the sequences related to quorum sensing genes; The quorum sensing literature library with signals and corresponding genes includes 6 signal molecules and 24 related gene sequences related thereto; the gene sequences include synthases, receptor enzymes, and degrading enzymes directly related to the signal molecules; S103. Obtain highly certain sequences in the second database; S104. Combine the sequences related to quorum sensing genes and the highly certain sequences, remove the duplicate sequences, number them according to "signal - major gene category - minor gene category - serial number", and then perform clustering analysis to remove redundant sequences to form seed sequences; S105. Align the protein conserved sites of the seed sequences, and construct a hidden Markov model of the protein conserved sites; S106. Align the hidden Markov model of the conservative site in the third database. According to E-value ≥ 1e -5 Retain the obtained protein sequences; S107. Screen the protein sequences obtained by the search in S106, and extract the proteins with highly credible sequence keywords related to quorum sensing in the annotation to form a highly credible sequence database; The formation of the highly credible sequence database includes extracting sequences in the NR database using the Seqkit software according to the search results, and combining all the sequences to form a quorum sensing sequence database, named the QSG database; S108. Statistically analyze the sequence recovery quantity and gene recovery rate corresponding to the genes, and construct a standardized database; The construction of the standardized database includes searching through the sequences in the QSG database using the Hmmsearch program with the T value adjusted to 20 - 300 respectively; calculating the sequence recovery quantity and gene recovery rate corresponding to each gene respectively, taking the T value with a recovery rate greater than 70% and the best recovery quantity as the optimal value, writing the optimal value into the hidden Markov model, merging the obtained texts using the cat command, and constructing a standardized database for the merged model using the Hmmpress program; S109. Annotate and align the protein sequences; the protein sequence annotation and comparison method adopts the Diamond alignment method or the Hmmer alignment method.
2. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, wherein The first database is the PubMed database of the National Center for Biotechnology Information of the United States; the second database is the Uniport, KEGG, and QSDB databases of the National Center for Biotechnology Information of the United States; the third database is the non - redundant database of the National Center for Biotechnology Information of the United States.
3. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, wherein The highly certain sequences in step S103 are sequences that contain related genes or functions and whose protein length is 80% - 140% of the average length of the sequences reported in the literature.
4. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, characterized in that, The method for removing duplicate sequences in step S104 includes retaining only one sequence with a similarity greater than 90%.
5. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, wherein The optimization method for the hidden Markov model of the protein conserved sites in step S105 includes searching for the conserved regions of the protein seed sequences using the Hmmsearch program and screening the score values; adjusting the search threshold to be between 20 and 300, modifying the step size to 5, and retaining the sequences whose search results exceed the search threshold.
6. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, wherein The formula for calculating the gene recovery rate described in step S108 is as follows: Among them, N 目的基因 represents the number of highly reliable sequences belonging to the gene in the search results, and N 搜索结果 represents the total number of search result sequences.
7. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, characterized in that The Diamond alignment method in step S109 includes formatting the input protein sequence as fasta, performing Diamond similarity alignment using the seed sequence, and retaining the results with e-value < 1e-5 and identify ≥ 50%.
8. The method for predicting the abundance of microbial quorum sensing genes according to claim 1, wherein The Hmmer alignment method in step S109 includes using the Hmmscan function of the Hmmer software to perform annotation alignment on the protein sequence with the parameter —cut_ga based on the constructed conserved site model.
Citation Information
Patent Citations
Method for obtaining coding genes of gram-positive bacterium quorum sensing signal peptides
CN105219787A
Method for accurately identifying unknown microbial community in water body based on metagenomics analysis
CN112786102A