Screening Method and System for Cellulose-Degrading Bacteria

By determining gene abundance and calculating a cellulose degradation score, the method enables efficient and consistent large-scale screening of cellulose-degrading microorganisms, overcoming the complexity and precision issues of existing methods.

CN119964638BActive Publication Date: 2025-07-15THREE GORGES ENVIRONMENTAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429818.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-15
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The method of screening cellulose degradation bacteria in the prior art is complicated to operate and is not suitable for large-scale applications, resulting in low screening efficiency.

Method used

By determining the gene abundance data of all microorganisms to be screened in the sample to be detected, the target genes are screened based on the functional gene set, and the cellulose degradation score is calculated to achieve high-throughput screening of cellulose degradation bacteria.

Benefits of technology

Improve screening efficiency, ensure high repetition and accuracy of screening, and enable large-scale analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964638B_ABST
    Figure CN119964638B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and system for screening cellulose-degrading bacteria. The method includes: determining all genes of all microorganisms to be screened in a sample to be detected and first gene abundance data of all the genes; screening out a plurality of target genes from all the genes based on a functional gene set; determining second gene abundance data of the plurality of target genes according to the first gene abundance data; determining a cellulose degradation score for each microorganism to be screened according to the second gene abundance data; and screening out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation score of each microorganism to be screened. High-throughput screening can be performed, and a large number of microorganisms to be screened can be analyzed simultaneously, improving the screening efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of bioinformatics technology, and in particular, to a method and system for screening cellulose-degrading bacteria. Background Art

[0002] In the exploration of the aerobic fermentation process of organic solid wastes such as agricultural crop straws and sludge, microorganisms play a core driving role in fermentation. Screening microorganisms that can efficiently degrade cellulose plays a key role in improving fermentation efficiency. Currently, the methods for measuring and screening cellulose-degrading bacteria are mainly experimental methods, such as Congo red staining method and reducing sugar method.

[0003] The Congo red staining method utilizes the property that Congo red dye forms a red complex with cellulose in the culture medium. When cellulose is decomposed by cellulase, the Congo red-cellulose complex cannot be formed, so a transparent circle will appear around the cellulose-decomposing bacteria in the culture medium. The cellulose degradation ability of microorganisms is screened and measured by observing whether a transparent circle is produced and the size of the transparent circle. However, this method is relatively complex in operation and usually requires a large number of dilution processes to clearly observe the formation of the hydrolysis circle, so it is not suitable for large-scale screening.

[0004] The principle of the reducing sugar method for measuring cellulase activity is based on the hydrolysis of cellulose by cellulase. Cellulase decomposes cellulose into reducing sugars, such as glucose. In the experiment, the dinitrosalicylic acid colorimetric method is usually used to measure the content of reducing sugars. However, this method has high requirements for precise control of experimental conditions, including factors such as temperature and reaction time, making the experimental operation relatively complex and not suitable for large-scale screening.

[0005] Whether it is the Congo red staining method or the reducing sugar method, both have the problems of relatively complex operation and high requirements for precise control of experimental conditions, so that large-scale screening cannot be carried out and the screening efficiency is low. Summary of the Invention

[0006] The embodiments of the present application provide a method and system for screening cellulose-degrading bacteria, so as to achieve high-throughput screening, analyze a large number of microorganisms to be screened at the same time, and improve the screening efficiency.

[0007] In a first aspect, the embodiments of the present application provide a method for screening cellulose-degrading bacteria, including:

[0008] Determine all genes of all microorganisms to be screened in the sample to be detected and the first gene abundance data of the all genes, where the first gene abundance data includes the gene abundance value of each gene in each of the all genes in each microorganism to be screened;

[0009] Based on the functional gene set, a plurality of target genes are screened out from all the genes, wherein the functional gene set includes a plurality of functional genes for indicating the cellulose degradation function;

[0010] According to the first gene abundance data, the second gene abundance data of the plurality of target genes is determined, wherein the second gene abundance data includes the gene abundance values of each target gene in each of the microorganisms to be screened;

[0011] According to the second gene abundance data, the cellulose degradation score of each of the microorganisms to be screened is determined;

[0012] According to the cellulose degradation scores of each of the microorganisms to be screened, the cellulose-degrading bacteria are screened out from all the microorganisms to be screened.

[0013] In a possible implementation manner, the determining the cellulose degradation score of each of the microorganisms to be screened according to the second gene abundance data includes: according to the second gene abundance data and the first gene abundance data, determining the first preset number of control genes and the third gene abundance data of each target gene, wherein the third gene abundance data includes the average gene abundance of the first preset number of control genes of each target gene in each of the microorganisms to be screened; determining the gene attribute vectors of all the target genes, wherein the gene attribute vectors are used to indicate the activation attributes or inhibition attributes of all the target genes; and calculating the cellulose degradation scores of each of the microorganisms to be screened according to the gene attribute vectors, the second gene abundance data, and the third gene abundance data.

[0014] In a possible implementation manner, the determining the first preset number of control genes and the third gene abundance data of each target gene according to the second gene abundance data and the first gene abundance data includes: summing the first gene abundance data to obtain the total gene abundance value of each gene in all the genes in each of the microorganisms to be screened; sorting the total gene abundance values of each gene according to the magnitude to generate an abundance sequence; based on the abundance sequence, dividing all the genes into a second preset number of groups; determining the groups where all the target genes are located, and randomly selecting the first preset number of control genes from the groups to which each target gene belongs to generate multiple groups of the first preset number of control genes; and obtaining the average gene abundance of each group of the first preset number of control genes in each of the microorganisms to be screened according to the first gene abundance data to generate the third gene abundance data.

[0015] In a possible implementation, calculating the cellulose degradation score of each microorganism to be screened according to the gene attribute vector, the second gene abundance data, and the third gene abundance data includes: generating a first gene abundance matrix according to the second gene abundance data; generating a second gene abundance matrix according to the third gene abundance data; calculating a degradation score matrix according to the gene attribute vector, the first gene abundance matrix, and the second gene abundance matrix, where the degradation score matrix includes the cellulose degradation scores of each microorganism to be screened.

[0016] In a possible implementation, screening out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation scores of each microorganism to be screened includes: sorting each microorganism to be screened according to the cellulose degradation score of each microorganism to be screened to obtain a degradation score sequence; screening out cellulose-degrading bacteria from all the microorganisms to be screened according to the degradation score sequence.

[0017] In a possible implementation, determining all genes of all the microorganisms to be screened in the sample to be detected and the first gene abundance data of all the genes includes: performing metagenomic sequencing on the sample to be detected, and performing data preprocessing on the obtained sequencing data to obtain the transcripts per million (TPM) value corresponding to each contig of each microorganism to be screened in the sample to be detected; calculating the standard abundance value of each microorganism to be screened according to the TPM value corresponding to each contig of each microorganism to be screened; calculating the gene abundance value of each gene in all the genes in each microorganism to be screened according to the standard abundance value of each microorganism to be screened to generate the first gene abundance data.

[0018] In a possible implementation, before determining all genes of all the microorganisms to be screened and the first gene abundance data of all the genes, it further includes: obtaining pathway information related to cellulose degradation from a preset database; obtaining cellulases related to cellulose degradation and genes of the cellulases according to the pathway information; obtaining regulatory genes for regulating the cellulases from a preset database; generating the functional gene set according to the genes of the cellulases and the regulatory genes.

[0019] In a possible implementation, the regulatory genes include transcriptional activator genes and transcriptional repressor genes.

[0020] In a possible implementation, the formula for calculating the degradation score matrix is:

[0021]

[0022] Wherein, D is the degradation fraction matrix; X is the first gene abundance matrix; B is the second gene abundance matrix; vector(G) is the gene attribute vector.

[0023] In a possible implementation, the formula for calculating the standard abundance value of each microorganism to be screened is:

[0024]

[0025] Wherein, Bin k _abundance_value represents the standard abundance value of the microorganism k to be screened; T k,j represents the TPM value of the j-th contig in the microorganism k to be screened; L k,j represents the length of the j-th contig in the microorganism k to be screened;

[0026] The formula for calculating the gene abundance value of each gene in all the genes in each microorganism to be screened is:

[0027]

[0028] Wherein, Gene_abundance k,i represents the gene abundance value of gene i in the microorganism k to be screened; mapped_gene_read_length k,i represents the length of the sequencing data mapped to gene i; gene_length i represents the gene length of gene i; Bin k _abundance_value represents the standard abundance value of the microorganism k to be screened.

[0029] In a second aspect, an embodiment of the present application provides a screening system for cellulose-degrading bacteria, including:

[0030] A determination module, configured to determine all genes of all microorganisms to be screened in a sample to be detected and first gene abundance data of all the genes, where the first gene abundance data includes the gene abundance value of each gene in all the genes in each microorganism to be screened;

[0031] A screening module, configured to screen out a plurality of target genes from all the genes based on a functional gene set, where the functional gene set includes a plurality of functional genes for indicating cellulose degradation function;

[0032] The determination module is further configured to determine second gene abundance data of the plurality of target genes according to the first gene abundance data, where the second gene abundance data includes the gene abundance value of each target gene in all the genes in each microorganism to be screened;

[0033] The determining module is further configured to determine a cellulose degradation score for each of the microorganisms to be screened according to the second gene abundance data.

[0034] The screening module is further configured to screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation score of each of the microorganisms to be screened.

[0035] The method and system for screening cellulose-degrading bacteria provided by the embodiments of the present application determine all genes of all the microorganisms to be screened and the first gene abundance data of all the genes; screen out a plurality of target genes from all the genes based on a functional gene set; determine the second gene abundance data of the plurality of target genes according to the first gene abundance data; determine a cellulose degradation score for each of the microorganisms to be screened according to the second gene abundance data; and screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation score of each of the microorganisms to be screened. It overcomes the limitation of strict control of experimental conditions in experimental screening, eliminates the problem of inconsistent screening conditions, ensures high repeatability of screening, enables high-throughput screening, analyzes a large number of microorganisms to be screened simultaneously, and improves the screening efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0037] Figure 1 It is a schematic structural diagram of a computer device provided by an embodiment of the present application;

[0038] Figure 2 It is a schematic flowchart of a method for screening cellulose-degrading bacteria provided by an embodiment of the present application;

[0039] Figure 3 It is a schematic diagram of the experimental results of a cellulose Congo red staining experiment in a test example of the present application;

[0040] Figure 4 It is a schematic structural diagram of a system for screening cellulose-degrading bacteria provided by an embodiment of the present application.

[0041] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and the textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of systems and methods consistent with some aspects of the present application as detailed in the appended claims.

[0043] The screening method of cellulose-degrading bacteria provided by the present application determines the gene abundance of target genes related to cellulose degradation function among all microorganisms to be screened; further determines the cellulose degradation score of each microorganism to be screened according to the gene abundance of the target gene; and screens out cellulose-degrading bacteria by ranking the cellulose degradation scores of each microorganism to be screened, thereby solving the technical problems of the prior art that large-scale screening cannot be carried out and the screening efficiency is low.

[0044] Figure 1 It is a schematic structural diagram of a computer device provided by an embodiment of the present application, which is used to execute the screening method of cellulose-degrading bacteria provided by an embodiment of the present application. As Figure 1 shown, the computer device provided in this embodiment includes: a receiving device 101, a processor 102, and a display device 103.

[0045] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the item recognition method. In other feasible embodiments of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenario and will not be limited here. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0046] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can obtain the data to be screened.

[0047] The processor 102 can process the data to be screened to determine the screening result.

[0048] The display device 103 can be used to display the above screening results and the like.

[0049] The display device can also be a touch display screen, which is used to receive user instructions while displaying the above content to realize interaction with the user.

[0050] It should be understood that the above processor can be implemented by the processor reading instructions in the memory and executing the instructions, or can be implemented by a chip circuit.

[0051] In addition, the network architecture and service scenarios described in the embodiments of this application are to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art can know that with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0052] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0053] Figure 2 It is a schematic flowchart of a method for screening cellulose-degrading bacteria provided by an embodiment of this application. As Figure 2 shown, this method includes:

[0054] S201: Determine all genes of all microorganisms to be screened in the sample to be detected and the first gene abundance data of all genes, where the first gene abundance data includes the gene abundance value of each gene in each microorganism to be screened among all genes.

[0055] Specifically, S201 specifically includes S2011 to S2013:

[0056] S2011: Perform metagenomic sequencing on the sample to be detected, and perform data preprocessing on the obtained sequencing data to obtain the Transcripts Per Million (TPM) value corresponding to each contig of each microorganism to be screened in the sample to be detected.

[0057] Specifically, perform data quality control, sequence assembly, binning, gene prediction, and gene annotation on the obtained sequencing data to obtain the gene data of each microorganism to be screened in the sample to be detected, where the gene data includes contigs; map the sequencing data to each contig of each microorganism to be screened to obtain the TPM value corresponding to each contig.

[0058] Specifically, perform adapter sequence detection and quality control on the obtained sequencing data, remove adapters and low-quality sequences to obtain first-processed data; perform separate assembly on the first-processed data to obtain second-processed data; bin the second-processed data according to the nucleic acid composition information characteristics and microbial abundance information characteristics of different microorganisms to obtain third-processed data. Use the de novo prediction method to perform gene prediction and gene annotation on the third-processed data to obtain fourth-processed data; generate gene data for each microorganism to be screened based on the first-processed data, second-processed data, third-processed data, and fourth-processed data.

[0059] Among them, the second-processed data includes contigs of each microorganism to be screened; the third-processed data includes genomic sequences of each microorganism to be screened; the fourth-processed data includes gene sequences of each microorganism to be screened and corresponding translated amino acid sequences.

[0060] S2012: Calculate the standard abundance value of each microorganism to be screened according to the TPM value corresponding to each contig of each microorganism to be screened.

[0061] Among them, the formula for calculating the standard abundance value of each microorganism to be screened is:

[0062]

[0063] In the formula, Bin k _abundance_value represents the standard abundance value of the microorganism k to be screened; T k,j represents the TPM value of the j-th contig in the microorganism k to be screened; L k,j represents the length of the j-th contig in the microorganism k to be screened.

[0064] S2013: Calculate the gene abundance value of each gene in each microorganism to be screened for all genes to generate first gene abundance data.

[0065] Among them, the formula for calculating the gene abundance value of each gene in each microorganism to be screened for all genes is:

[0066]

[0067] In the formula, Gene_abundance k,i represents the gene abundance value of gene i in the microorganism k to be screened; mapped_gene_read_length k,i represents the length of the sequencing data mapped to gene i; gene_length i represents the gene length of gene i, in base pairs; Bin k_abundance_value represents the standard abundance value of the microorganism k to be screened.

[0068] S202: Based on the functional gene set, multiple target genes are screened out from all genes, where the functional gene set includes multiple functional genes for indicating the cellulose degradation function.

[0069] Specifically, according to the gene identifiers in the functional gene set, they are matched with all genes to obtain multiple target genes.

[0070] S203: According to the first gene abundance data, determine the second gene abundance data of multiple target genes, where the second gene abundance data includes the gene abundance values of each target gene in each microorganism to be screened.

[0071] Specifically, according to the representation of multiple target genes, they are matched with the first gene abundance data to obtain the second gene abundance data.

[0072] S204: According to the second gene abundance data, determine the cellulose degradation score of each microorganism to be screened.

[0073] Specifically, determine the gene attributes of each target gene; according to the second gene abundance data and each gene attribute, calculate the sum of the gene abundance values of each target gene in each microorganism to be screened to obtain the cellulose degradation score of each microorganism to be screened.

[0074] S205: Screen out the cellulose-degrading bacteria among all the microorganisms to be screened according to the cellulose degradation score of each microorganism to be screened.

[0075] Specifically, S205 specifically includes S2051~S2052:

[0076] S2051: Sort each microorganism to be screened according to the cellulose degradation score of each microorganism to be screened to obtain a degradation score sequence.

[0077] Specifically, sort each microorganism to be screened from largest to smallest or from smallest to largest according to the cellulose degradation score to obtain a degradation score sequence.

[0078] S2052: Screen out the cellulose-degrading bacteria among all the microorganisms to be screened according to the degradation score sequence.

[0079] Specifically, according to the degradation score sequence, select the microorganisms to be screened with a cellulose degradation score higher than a preset threshold or within a preset range of the cellulose degradation score ranking as cellulose-degrading bacteria.

[0080] The screening method of cellulose-degrading bacteria provided by the embodiments of the present application includes determining all genes of all microorganisms to be screened and the first gene abundance data of all genes; screening out a plurality of target genes from all genes based on a functional gene set; determining the second gene abundance data of the plurality of target genes according to the first gene abundance data; determining the cellulose degradation score of each microorganism to be screened according to the second gene abundance data; and screening out cellulose-degrading bacteria from all microorganisms to be screened according to the cellulose degradation score of each microorganism to be screened. It overcomes the limitations of strict control of experimental conditions in experimental screening, eliminates the problem of inconsistent screening conditions, ensures high repeatability of screening, enables high-throughput screening, analyzes a large number of microorganisms to be screened simultaneously, and improves the screening efficiency.

[0081] In one embodiment of the present application, on the basis of the above embodiment, in step S204, another implementation manner is further provided, which is described in detail as follows:

[0082] S2041: According to the second gene abundance data and the first gene abundance data, determine the first preset number of control genes for each target gene and the third gene abundance data, where the third gene abundance data includes the average gene abundance of the first preset number of control genes for each target gene in each microorganism to be screened.

[0083] Among them, the first preset number can be 100.

[0084] Specifically, S2041 specifically includes Sa~Se:

[0085] Sa: Sum the first gene abundance data to obtain the total gene abundance value of each gene in each microorganism to be screened among all genes.

[0086] Among them, the first gene abundance data can be a table with the column headers being the names of the microorganisms to be screened, the row headers being the gene names, and the data being the gene abundance values of the corresponding genes in the microorganisms to be screened.

[0087] Specifically, sum the gene abundance data in the first gene abundance data row by row to obtain the total gene abundance value of each gene in each microorganism to be screened among all genes.

[0088] Sb: Sort the total gene abundance values of each gene according to their magnitudes to generate an abundance sequence.

[0089] Sc: Based on the abundance sequence, divide all genes into a second preset number of groups.

[0090] Among them, the second preset number can be 50.

[0091] Specifically, an abundance ranking table is generated according to the abundance sequence; all genes are divided into a second preset number of groups according to the abundance ranking table.

[0092] Among them, the column headers of the abundance ranking table include: gene name, gene abundance, ranking, etc.

[0093] Table 1 is an abundance ranking table provided by an embodiment of the present application, as shown in Table 1:

[0094] Table 1

[0095]

[0096] Sd: Determine the groups where all target genes are located, and randomly select a first preset number of control genes from the groups to which each target gene belongs to generate multiple groups of first preset number of control genes.

[0097] Specifically, a gene grouping table is generated according to the group information of all genes; the groups where each target gene is located are obtained from the gene grouping table; a first preset number of control genes corresponding to each target gene are randomly selected in each group to generate multiple groups of first preset number of control genes.

[0098] Among them, the column headers of the gene grouping table include gene name, gene group, gene type, etc.

[0099] Among them, the gene type includes functional genes and others.

[0100] Table 2 is a gene grouping table provided by an embodiment of the present application, as shown in Table 2:

[0101] Table 2

[0102]

[0103] Se: According to the first gene abundance data, obtain the average gene abundance of each group of first preset number of control genes in each microorganism to be screened to generate third gene abundance data.

[0104] Specifically, match the first gene abundance data according to the control gene name to obtain the gene abundance value of each control gene; according to the gene abundance value of each control gene, calculate the average gene abundance of each group of first preset number of control genes in each microorganism to be screened to generate third gene abundance data.

[0105] S2042: Determine the gene attribute vectors of all target genes, where the gene attribute vectors are used to indicate the activation attributes or inhibition attributes of all target genes.

[0106] Among them, the gene attribute vector can be a row vector.

[0107] Exemplarily, the gene attribute vector can be [1 1 -1 …], where +1 represents that the target gene has an activation attribute, and -1 represents that the target gene has an inhibition attribute.

[0108] S2043: Calculate the cellulose degradation score of each microorganism to be screened according to the gene attribute vector, the second gene abundance data, and the third gene abundance data.

[0109] As can be seen from the above description, in the embodiments of the present application, by determining the first preset number of control genes and the third gene abundance data of each target gene according to the second gene abundance data and the first gene abundance data; calculating the cellulose degradation score of each microorganism to be screened according to the gene attribute vector, the second gene abundance data, and the third gene abundance data. The control genes are used to correct the cellulose degradation score, improving the accuracy of the cellulose degradation score, and thus improving the accuracy of screening cellulose-degrading bacteria.

[0110] In an embodiment of the present application, on the basis of the above embodiment, step S2043 also provides another implementation method, which is described in detail as follows:

[0111] Sf: Generate a first gene abundance matrix according to the second gene abundance data.

[0112] Wherein, each column of the first gene abundance matrix represents the gene abundance value of each target gene in a microorganism to be screened.

[0113] Sg: Generate a second gene abundance matrix according to the third gene abundance data.

[0114] Wherein, each column of the second gene abundance matrix represents the average gene abundance of the first preset number of control genes corresponding to each target gene in a microorganism to be screened.

[0115] Sh: Calculate a degradation score matrix according to the gene attribute vector, the first gene abundance matrix, and the second gene abundance matrix, where the degradation score matrix includes the cellulose degradation score of each microorganism to be screened.

[0116] Wherein, the gene attribute vector can be a row vector.

[0117] Wherein, the formula for calculating the degradation score matrix is:

[0118]

[0119] In the formula, D is the degradation score matrix; X is the first gene abundance matrix; B is the second gene abundance matrix; vector(G) is the gene attribute vector.

[0120] As can be seen from the above description, in the embodiments of the present application, by calculating the degradation score matrix based on the gene attribute vector, the first gene abundance matrix, and the second gene abundance matrix, the calculation process of the degradation score is simplified, providing a basis for simultaneously analyzing a large number of microorganisms to be screened.

[0121] In an embodiment of the present application, on the basis of the above embodiment, it further includes the process of obtaining a functional gene set, which is described in detail as follows:

[0122] S206: Obtain the pathway information related to cellulose degradation from a preset database.

[0123] Among them, the pathway information is the metabolic pathway of cellulose degradation, including that cellulose is gradually decomposed into cellobiose under the combined action of β-glucosidase (EC3.2.1.21), exoglucanase (cellodextrinase, EC 3.2.1.74), endoglucanase (EC 3.2.1.4), and cellobiohydrolase acting on the non-reducing end (EC 3.2.1.91), and finally converted into D-glucose under the action of β-glucosidase.

[0124] S207: Obtain the cellulases related to cellulose degradation and the genes of the cellulases according to the pathway information.

[0125] Among them, the cellulases include endoglucanase, exoglucanase (cellodextrinase), exoglucanase (cellobiohydrolase acting on the non-reducing end), exoglucanase (cellobiohydrolase acting on the reducing end), and β-glucosidase, etc.

[0126] S208: Obtain the regulatory genes for regulating cellulases from a preset database.

[0127] Among them, the regulatory genes include transcriptional activator genes and transcriptional repressor genes, etc.

[0128] Among them, the transcriptional activator genes include transcriptional factor genes that enhance the expression of cellulase genes and genes that form transcriptional activation complexes, etc.

[0129] Among them, the genes that form transcriptional activation complexes include cellulase expression activators, cellulase regulators, and transcriptional activation complexes that regulate cellulase genes in fungi.

[0130] Among them, the transcriptional repressor genes include genes that inhibit the expression of cellulase genes.

[0131] Specifically, obtain the regulatory process of cellulase gene expression from a preset database; retrieve the regulatory genes for regulating cellulases from the regulatory process of cellulase gene expression.

[0132] S209: Generate a functional gene set based on the genes of cellulase and regulatory genes.

[0133] Table 3 shows a functional gene set provided by an embodiment of the present application, as shown in Table 3:

[0134] Table 3

[0135]

[0136] As can be seen from the above description, in the embodiment of the present application, cellulase and the genes of cellulase related to cellulose degradation are obtained according to pathway information; regulatory genes for regulating cellulase are obtained from a preset database; a functional gene set is generated based on the genes of cellulase and regulatory genes. The functional gene set not only covers the genes encoding cellulase but also includes the regulatory genes that regulate the expression of cellulase genes, improving the accuracy of screening cellulose-degrading bacteria.

[0137] In an experimental example of the present application, an application of a screening method for cellulose-degrading bacteria in a sludge aerobic fermentation sample is provided, which is described in detail as follows:

[0138] The method for obtaining the sludge aerobic fermentation sample selected in this experimental example is as follows: 1.5 tons of sludge was collected from a sewage treatment facility and fully mixed with corn straw at a mass ratio of 5:1 (1.5t of sludge and 300kg of corn straw), and an aerobic composting experiment was carried out. During the high-temperature stage, samples of the mixture of sludge and auxiliary materials were taken.

[0139] In this experimental example, by determining the cellulose degradation fraction D value of microorganisms in the sludge aerobic fermentation sample, the microorganisms are sorted according to the D value from large to small, and it is determined that the microorganisms in the top 5% have strong cellulose degradation ability and can be used as candidate cellulose-degrading bacteria. For microorganisms with a D value less than or equal to 0, it is determined that they do not have the ability to degrade cellulose.

[0140] Through the above method, this experimental example screened out microorganisms with cellulose degradation ability and microorganisms that cannot degrade cellulose in sludge aerobic fermentation. Table 4 lists the D value data of typical microorganisms among them, as shown in Table 4:

[0141] Table 4

[0142]

[0143] The microorganisms with cellulose degradation ability screened out in this experimental example are specific strains of Clostridium thermocellum, Bacillus licheniformis, Streptomyces, and Bacillus subtilis. At the same time, a microorganism that does not have the ability to degrade cellulose was also identified.

[0144] In a test example of this application, a cellulose Congo red staining experiment was conducted on the microorganisms screened in the above experimental example. The experimental procedure included: selecting target bacterial colonies that were morphologically distinguishable, transferring them to shaking culture tubes, culturing them by shaking at 60 °C and 150 rpm for 48 hours, then diluting them 1000 times and spreading them on a cellulose Congo red culture medium. These culture media were placed in a constant temperature incubator and cultured at 55 °C for 48 - 60 hours, and it was observed whether a clear hydrolysis zone was formed.

[0145] Figure 3 It is a schematic diagram of the experimental results of the cellulose Congo red staining experiment in a test example of this application. From Figure 3 it can be seen that the microorganisms with cellulose degradation ability obtained in the above experimental example, numbered 1, 2, 3, and 4 respectively, formed obvious clear hydrolysis zones on the cellulose Congo red culture medium. In contrast, the microorganism without cellulose degradation ability, numbered 5, did not observe the formation of a hydrolysis zone on the same culture medium.

[0146] From Figure 3 it can be seen that the screening method for cellulose-degrading bacteria provided in the embodiments of this application has high accuracy and reliability.

[0147] Figure 4 It is a schematic structural diagram of the screening system for cellulose-degrading bacteria provided in the embodiments of this application. As Figure 4 shown, the screening system 40 for cellulose-degrading bacteria provided in this embodiment includes: a determination module 401 and a screening module 402.

[0148] The determination module 401 is used to determine all genes of all microorganisms to be screened in the sample to be detected and the first gene abundance data of all these genes, where the first gene abundance data includes the gene abundance value of each gene in each microorganism to be screened among all these genes;

[0149] The screening module 402 is used to screen out multiple target genes from all these genes based on a functional gene set, where the functional gene set includes multiple functional genes for indicating cellulose degradation function;

[0150] The determination module 401 is further used to determine the second gene abundance data of the multiple target genes according to the first gene abundance data, where the second gene abundance data includes the gene abundance value of each target gene in each microorganism to be screened among the multiple target genes;

[0151] The determination module 401 is further used to determine the cellulose degradation score of each microorganism to be screened according to the second gene abundance data;

[0152] The screening module 402 is further configured to screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation scores of each microorganism to be screened.

[0153] In a possible implementation manner, the determining module 401 is specifically configured to: determine a first preset number of control genes and third gene abundance data for each target gene according to the second gene abundance data and the first gene abundance data, where the third gene abundance data includes the average gene abundances of the first preset number of control genes for each target gene in each microorganism to be screened; determine a gene attribute vector for all the target genes, where the gene attribute vector is used to indicate the activation attribute or inhibition attribute of all the target genes; calculate the cellulose degradation scores of each microorganism to be screened according to the gene attribute vector, the second gene abundance data, and the third gene abundance data.

[0154] In a possible implementation manner, when the determining module 401 "determines a first preset number of control genes and third gene abundance data for each target gene according to the second gene abundance data and the first gene abundance data", it is specifically configured to: sum the first gene abundance data to obtain the total gene abundances of each gene in each microorganism to be screened among all the genes; sort the total gene abundances of each gene in descending order to generate an abundance sequence; divide all the genes into a second preset number of groups based on the abundance sequence; determine the groups where all the target genes are located, and randomly select the first preset number of control genes from the groups to which each target gene belongs to generate multiple groups of the first preset number of control genes; obtain the average gene abundances of each group of the first preset number of control genes in each microorganism to be screened according to the first gene abundance data to generate the third gene abundance data.

[0155] In a possible implementation manner, when the determining module 401 "calculates the cellulose degradation scores of each microorganism to be screened according to the gene attribute vector, the second gene abundance data, and the third gene abundance data", it is specifically configured to: generate a first gene abundance matrix according to the second gene abundance data; generate a second gene abundance matrix according to the third gene abundance data; calculate a degradation score matrix according to the gene attribute vector, the first gene abundance matrix, and the second gene abundance matrix, where the degradation score matrix includes the cellulose degradation scores of each microorganism to be screened.

[0156] In a possible implementation, the screening module 402 is specifically configured to: sort each microorganism to be screened according to the cellulose degradation score of each microorganism to be screened to obtain a degradation score sequence; and screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the degradation score sequence.

[0157] In a possible implementation, the determination module 401 is specifically configured to: perform metagenomic sequencing on a sample to be detected, and perform data preprocessing on the obtained sequencing data to obtain the transcripts per million (TPM) value corresponding to each contig of each microorganism to be screened in the sample to be detected; calculate the standard abundance value of each microorganism to be screened according to the TPM value corresponding to each contig of each microorganism to be screened; and calculate the gene abundance value of each gene in all the genes in each microorganism to be screened according to the standard abundance value of each microorganism to be screened to generate first gene abundance data.

[0158] In a possible design, the cellulose-degrading bacteria screening system 40 further includes:

[0159] a generation module, configured to obtain pathway information related to cellulose degradation from a preset database; obtain cellulase related to cellulose degradation and the gene of the cellulase according to the pathway information; obtain a regulatory gene for regulating the cellulase from a preset database; and generate the functional gene set according to the gene of the cellulase and the regulatory gene.

[0160] In a possible design, the regulatory gene includes a transcriptional activator gene and a transcriptional repressor gene.

[0161] In a possible design, the formula for calculating the degradation score matrix is:

[0162]

[0163] In the formula, D is the degradation score matrix; X is the first gene abundance matrix; B is the second gene abundance matrix; and vector(G) is the gene attribute vector.

[0164] In a possible design, the formula for calculating the standard abundance value of each microorganism to be screened is:

[0165]

[0166] In the formula, Bin k _abundance_value represents the standard abundance value of the microorganism to be screened k; T k,j represents the TPM value of the jth contig in the microorganism to be screened k; L k,jdenotes the length of the j-th contig in the to-be-screened microorganism k;

[0167] The formula for calculating the gene abundance value of each gene among all the genes in each to-be-screened microorganism is as follows:

[0168]

[0169] In the formula, Gene_abundance k,i denotes the gene abundance value of gene i in the to-be-screened microorganism k; mapped_gene_read_length k,i denotes the length of the sequencing data mapped to gene i; gene_length i denotes the gene length of gene i; Bin k _abundance_value denotes the standard abundance value of the to-be-screened microorganism k.

[0170] The screening system for cellulose-degrading bacteria provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0171] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0172] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0173] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0174] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0175] The division of units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0176] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0178] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.

[0179] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disks or optical disks and other various media that can store program codes.

[0180] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for screening cellulose-degrading bacteria, characterized in that, Including: Determine all genes of all microorganisms to be screened in the sample to be detected and the first gene abundance data of all the genes, where the first gene abundance data is the gene abundance value of each gene in each microorganism to be screened among all the genes; Based on the functional gene set, screen out a plurality of target genes from all the genes, where the functional gene set includes a plurality of functional genes for indicating cellulose degradation function; According to the first gene abundance data, determine the second gene abundance data of the plurality of target genes, where the second gene abundance data is the gene abundance value of each target gene in each microorganism to be screened among the plurality of target genes; According to the second gene abundance data, determine the cellulose degradation score of each microorganism to be screened; Screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation score of each microorganism to be screened; The determining the cellulose degradation score of each microorganism to be screened according to the second gene abundance data includes: According to the second gene abundance data and the first gene abundance data, determine the first preset number of control genes for each target gene and the third gene abundance data, where the third gene abundance data is the average gene abundance of the first preset number of control genes of each target gene in each microorganism to be screened; Determine the gene attribute vector of all the target genes, where the gene attribute vector is used to indicate the activation attribute or inhibition attribute of all the target genes; Calculate the cellulose degradation score of each microorganism to be screened according to the gene attribute vector, the second gene abundance data and the third gene abundance data.

2. The method according to claim 1, wherein The determining the first preset number of control genes for each target gene and the third gene abundance data according to the second gene abundance data and the first gene abundance data includes: Sum the first gene abundance data to obtain the total gene abundance value of each gene in each microorganism to be screened among all the genes; Sort the total gene abundance values of each gene according to their magnitudes to generate an abundance sequence; Based on the abundance sequence, divide all the genes into a second preset number of groups; Determine the groups where all the target genes are located, and randomly select the first preset number of control genes for each target gene from the group to which it belongs to generate multiple groups of the first preset number of control genes; According to the first gene abundance data, obtain the average gene abundance of each group of the first preset number of control genes in each microorganism to be screened to generate the third gene abundance data.

3. The method according to claim 1, wherein The calculating the cellulose degradation score of each microorganism to be screened according to the gene attribute vector, the second gene abundance data and the third gene abundance data includes: Generate a first gene abundance matrix according to the second gene abundance data; Generate a second gene abundance matrix according to the third gene abundance data; According to the gene attribute vector, the first gene abundance matrix, and the second gene abundance matrix, a degradation score matrix is calculated, where the degradation score matrix includes the cellulose degradation scores of each of the to-be-screened microorganisms.

4. The method according to any one of claims 1 to 3, characterized in that The screening of the cellulose-degrading bacteria from all the to-be-screened microorganisms according to the cellulose degradation scores of each of the to-be-screened microorganisms includes: Sorting each of the to-be-screened microorganisms according to the cellulose degradation scores of each of the to-be-screened microorganisms to obtain a degradation score sequence; Screening the cellulose-degrading bacteria from all the to-be-screened microorganisms according to the degradation score sequence.

5. The method according to any one of claims 1 to 3, characterized in that, The determination of all the genes of all the to-be-screened microorganisms in the to-be-detected sample and the first gene abundance data of all the genes includes: Performing metagenomic sequencing on the to-be-detected sample, and performing data preprocessing on the obtained sequencing data to obtain the transcripts per million (TPM) values corresponding to each contig of each of the to-be-screened microorganisms in the to-be-detected sample; Calculating the standard abundance value of each of the to-be-screened microorganisms according to the TPM values corresponding to each contig of each of the to-be-screened microorganisms; Calculating the gene abundance value of each gene in all the genes in each of the to-be-screened microorganisms according to the standard abundance value of each of the to-be-screened microorganisms to generate the first gene abundance data.

6. The method according to any one of claims 1 to 3, characterized in that, Before the determination of all the genes of all the to-be-screened microorganisms and the first gene abundance data of all the genes, it further includes: Obtaining the pathway information related to cellulose degradation from a preset database; Obtaining the cellulases related to cellulose degradation and the genes of the cellulases according to the pathway information; Obtaining the regulatory genes for regulating the cellulases from a preset database; Generating the functional gene set according to the genes of the cellulases and the regulatory genes.

7. The method according to claim 6, characterized in that, The regulatory genes include transcriptional activator genes and transcriptional repressor genes.

8. The method according to claim 3, wherein The formula for calculating the degradation score matrix is: In the formula, D is the degradation score matrix; X is the first gene abundance matrix; B is the second gene abundance matrix; vector(G) is the gene attribute vector.

9. The method according to claim 5, characterized in that, The formula for calculating the standard abundance value of each of the to-be-screened microorganisms is: Wherein, Bin k _abundance_value represents the standard abundance value of the microorganism k to be screened; T k,j represents the TPM value of the j-th contig in the microorganism k to be screened; L k,j represents the length of the j-th contig in the microorganism k to be screened; The formula for calculating the gene abundance value of each gene in all the genes in each of the to-be-screened microorganisms is: In the formula, Gene_abundance k,i represents the gene abundance value of gene i of the microorganism k to be screened; mapped_gene_read_length k,i represents the length of the sequencing data mapped to gene i; gene_length i represents the gene length of gene i; Bin k _abundance_value represents the standard abundance value of the microorganism k to be screened.

10. A screening system for cellulose-degrading bacteria, characterized in that, It includes: A determination module, configured to determine all the genes of all the to-be-screened microorganisms in the to-be-detected sample and the first gene abundance data of all the genes, where the first gene abundance data is the gene abundance value of each gene in all the genes in each of the to-be-screened microorganisms; A screening module, configured to screen out a plurality of target genes from all the genes based on the functional gene set, where the functional gene set includes a plurality of functional genes for indicating cellulose degradation functions; The determination module is further configured to determine the second gene abundance data of the plurality of target genes according to the first gene abundance data, where the second gene abundance data is the gene abundance value of each target gene in the plurality of target genes in each of the to-be-screened microorganisms; The determining module is further configured to determine the cellulose degradation score of each microorganism to be screened according to the second gene abundance data; The screening module is further configured to screen out cellulose-degrading bacteria from all the microorganisms to be screened according to the cellulose degradation score of each microorganism to be screened; Specifically, the determining module is configured to determine a first preset number of control genes of each target gene and the third gene abundance data according to the second gene abundance data and the first gene abundance data, where the third gene abundance data is the average gene abundance of the first preset number of control genes of each target gene in each microorganism to be screened; determine the gene attribute vectors of all target genes, where the gene attribute vectors are used to indicate the activation attributes or inhibition attributes of all target genes; and calculate the cellulose degradation score of each microorganism to be screened according to the gene attribute vectors, the second gene abundance data, and the third gene abundance data.

Citation Information

Patent Citations

  • Toxin gene abundance detection method based on metagenomics and annotation database construction method

    CN114621997A

  • Method for analyzing drug-resistant gene pollution of traceable soil by using metagenome

    CN114822697A