Method for Identifying Microhabitats Based on Machine Learning
The machine learning-based method addresses the challenge of identifying optimal microhabitats for microbial communities by analyzing non-linear relationships, improving identification accuracy and functionality through a microbial community evaluation model and ICE algorithm.
Patent Information
- Application Number
- JP2025011038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-10-28
- Filing Date
- 2025-01-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-27
AI Technical Summary
Existing methods fail to accurately identify optimal microhabitats for microbial communities due to their complex non-linear relationships, leading to unclear correlations between microhabitats and microbial communities.
A machine learning-based method using a microbial community evaluation model, trained with gradient boosting regression trees, to analyze the non-linear relationships between microhabitat characteristics and microbial community interactions, optimizing the mutual relationship evaluation index through the individual conditional expectation algorithm to determine optimal microhabitat characteristics.
Improves the accuracy of microhabitat identification by clarifying the relationship between microhabitats and microbial communities, enhancing the functionality of microbial communities by identifying optimal microhabitat conditions.
Smart Images

Figure 0007710648000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention belongs to the field of bioinformatics and biotechnology, and in particular to the field of microhabitat based on machine learning. It relates to a tat identification method. [Background technology]
[0002] Microbial communities are complex ecosystems that consist of various microorganisms and play important roles in material circulation and energy They play an important role in processes such as water flow, sustainable agriculture, ecosystem restoration, and human Microhabitats play an important role in areas such as human health. Microhabitats refer to the specific environment in which something lives. Factors that affect a microhabitat include physical, These factors include chemical, biological and process parameters. They not only affect the diversity and community structure of microbial communities, but also determine their ecological function and adaptive capacity. Therefore, creating suitable microhabitat conditions for microbial communities is a key They are important for stabilizing or regulating biological communities. The process of identifying microhabitats for microbial communities is called microhabitat identification. Statistical correlation analysis is usually used to quantify the relationship between microhabitats and microbial communities. Based on this, microhabitats are identified, and statistical correlation analysis, for example, principal component analysis (PCA) A), Redundancy Analysis (RDA), Canonical Correlation Analysis (CCA), Non-metric Multidimensional Scaling Analysis ( However, these statistical correlation analysis methods have some limitations. By clarifying the correlation between microhabitats and microbial communities, In reality, however, the complex relationships between microbial communities and microhabitats cannot be clearly revealed. Despite the existence of a complex non-linear relationship between them, a linear relationship is being analyzed, which has drawbacks. Under such circumstances, there is a lack of a method for analyzing the complex non-linear relationship between microbial communities and microhabitats and identifying the optimal microhabitats for microbial communities.
Summary of the Invention
[0003] The present invention provides a method for identifying microhabitats based on machine learning, which captures the complex non-linear relationship between microbial communities and microhabitats by machine learning, analyzes and identifies the values of microhabitat characteristics that optimize the mutual relationship evaluation index of microbial communities based on this, and can improve the identification accuracy of microhabitats. To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for identifying microhabitats based on machine learning, constructs a training sample set for training a microbial community evaluation model, where the training sample set includes the values of the microhabitat characteristics of each of a plurality of microbial communities in a sewage treatment plant and the mutual relationship evaluation index of each of the plurality of microbial communities. The mutual relationship evaluation index is used to indicate the mutual relationship between microorganisms in the microbial community, and the microhabitat characteristics include temperature, pH, total nitrogen, ammonia nitrogen, total phosphorus, total organic carbon, dissolved oxygen, sludge age, and hydraulic retention time. The values of the microhabitat characteristics of each microbial community in the training sample set are used as input values, the mutual relationship evaluation index of each microbial community is used as an output value, and a microbial community evaluation model is trained using a machine learning algorithm. The microbial community evaluation model is a gradient boosting regression tree GBRT. Further, the target microhabitat of the microbial community Determine the ichlohabitat characteristics and the phase of the microbial community of the above target microhabitat characteristics The degree of influence on the mutual relationship evaluation index of the microbial community is higher than that of the influence of other microhabitat characteristics of the microbial community on the mutual relationship evaluation index of the microbial community. Based on the microbial community evaluation model, Execute the individual's conditional expectation ICE algorithm for the above target microhabitat characteristics, and determine the value of the target microhabitat characteristics corresponding to the maximum value of the mutual relationship evaluation index of the microbial community. When the value of the mutual relationship evaluation index of the microbial community is maximized, Execute the individual's conditional expectation ICE algorithm for the above target microhabitat characteristics, and determine the value of the target microhabitat characteristics corresponding to the maximum value of the mutual relationship evaluation index of the microbial community. Determine the value of the target ichlohabitat characteristics. In a possible implementation manner, based on the microbial community evaluation model, execute the individual's conditional expectation ICE algorithm for the above target microhabitat characteristics, and determine the value of the target microhabitat characteristics corresponding to the maximum value of the mutual relationship evaluation index of the microbial community. This includes the following steps: Step 1: Generate a plurality of values of the target microhabitat characteristics based on the distribution of the values of the target microhabitat characteristics of the microbial community. Step 2: Input the plurality of values of the target microhabitat characteristics into the microbial community evaluation model respectively to obtain a plurality of evaluation indexes, and generate the ICE graph of the target microhabitat characteristics. The ICE graph is used to describe the relationship between the value of the microhabitat characteristics and the mutual relationship evaluation index of the microbial community. Step 3: Repeat Steps 1 and 2 to obtain a plurality of ICE graphs of the target microhabitat characteristics, calculate the average value of the plurality of ICE graphs to obtain the average ICE graph, and in the average ICE graph, determine the value of the target microhabitat characteristics corresponding to the maximum mutual relationship evaluation index of the microbial community. Generate a plurality of values of the target microhabitat characteristics. Step 2: Input the plurality of values of the target microhabitat characteristics into the microbial community evaluation model respectively to obtain a plurality of evaluation indexes, and generate the ICE graph of the target microhabitat characteristics. The ICE graph is used to describe the relationship between the value of the microhabitat characteristics and the mutual relationship evaluation index of the microbial community. Generate the ICE graph of the target microhabitat characteristics. The ICE graph is used to describe the relationship between the value of the microhabitat characteristics and the mutual relationship evaluation index of the microbial community. The ICE graph is used to describe the relationship between the value of the microhabitat characteristics and the mutual relationship evaluation index of the microbial community. Step 3: Repeat Steps 1 and 2 to obtain a plurality of ICE graphs of the target microhabitat characteristics, calculate the average value of the plurality of ICE graphs to obtain the average ICE graph, and in the average ICE graph, determine the value of the target microhabitat characteristics corresponding to the maximum mutual relationship evaluation index of the microbial community. Obtain a plurality of ICE graphs of the target microhabitat characteristics. Calculate the average value of the plurality of ICE graphs to obtain the average ICE graph. In the average ICE graph, determine the value of the target microhabitat characteristics corresponding to the maximum mutual relationship evaluation index of the microbial community. Determine the value of the target microhabitat characteristics corresponding to the maximum mutual relationship evaluation index of the microbial community. In a possible implementation manner, the above mutual relationship includes a positive correlation relationship and a negative correlation relationship. In a possible implementation, the calculation process of the mutual relationship evaluation index of each microbial community in the training sample set includes the following. For each of the above microbial communities, the community information of each microbial community is obtained. The community information includes the types included in each microbial community and the abundance of the types. Based on the types included in each microbial community and the abundance of the types, Spiec-Easi is used to construct the interaction network of each microbial community. The interaction network reflects the mutual relationship between the microorganisms in the microbial community. To use the ratio of the number of positive correlation edges to the total number of edges in the interaction network as the mutual relationship evaluation index of each microbial community. is used. In a possible implementation, determining the target microhabitat characteristics of the microbial community includes the following. For each microhabitat characteristic of the microbial community, the importance index of the microhabitat characteristic is determined, and the importance indexes of a plurality of microhabitat characteristics of the microbial community are obtained. The importance index is the absolute value of the difference between the performance index of the first microbial community evaluation model and the performance index of the second microbial community evaluation model. Here, the first microbial community evaluation model is obtained by training based on all the microhabitat characteristics of the microbial community. The second microbial community evaluation model is obtained by training based on the microhabitat characteristics after removing the microhabitat characteristics. The performance indexes of the first microbial community evaluation model and the second microbial community evaluation model are the root mean square error RMSE of the true value and the predicted value, or the mean absolute error MAE of the true value and the predicted value. The microhabitats corresponding to the first n importance indexes with large values in the importance indexes of the plurality of microhabitat characteristics are those. Here, the first microbial community evaluation model is obtained by training based on all the microhabitat characteristics of the microbial community. The second microbial community evaluation model is obtained by training based on the microhabitat characteristics after removing the microhabitat characteristics. The performance indexes of the first microbial community evaluation model and the second microbial community evaluation model are the root mean square error RMSE of the true value and the predicted value, or the mean absolute error MAE of the true value and the predicted value. The microhabitats corresponding to the first n importance indexes with large values in the importance indexes of the plurality of microhabitat characteristics are those. The characteristics are taken as the above target microhabitat characteristics, where n is 1 or more, n is an integer smaller than N, and N is the total number of microhabitat characteristics of the above microbial community. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. In a possible embodiment, obtaining the community information of each of the above microbial communities includes the following. Compare the sequence data of each microbial community with the 16S rRNA database to determine the types of microorganisms included in the above microbial community, calculate the abundance of the above types, and the abundance of the above types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the above microbial community is the sequence data obtained by determining the base sequence of each of the above microbial communities using 16S rRNA amplicon sequencing. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. The machine learning-based microhabitat identification method provided by the present invention represents the complex non-linear relationship between the microhabitat characteristics of a microbial community and the functional characteristics of the microbial community (the correlation evaluation index supports the functional characteristics of the microbial community) using a microbial community evaluation model, and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the conventional microhabitat identification method, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Embodiments for Carrying Out the Invention
[0005] First, some technical terms related to the embodiments of the present invention will be explained. 1. Microhabitat It refers to a specific environment in which microorganisms in a microbial community inhabit, and microhabitats under different conditions have different effects on the microbial community. 2. Microhabitat characteristics (also called microhabitat information) It refers to several factors that affect a microhabitat, and the factors that affect a microhabitat can be classified according to different characteristics, and can be divided into physical factors, chemical factors, biological chemical factors and process parameters. Exemplarily, physical factors include temperature (T), humidity, etc., chemical factors include pH value, etc., biological factors include the concentration of nutrients , for example, total nitrogen (abbreviation: TN, unit: mg / L), total phosphorus (TP, mg / L), total organic carbon (TOC, mg / L), etc., and process parameters include dissolved oxygen (DO, mg / L ), sludge age (SRT, h) and hydraulic retention time (HRT, h), etc. 3. Microhabitat identification In the embodiments of the present invention, microhabitat identification specifically refers to determining the value or value range of each microhabitat characteristic when a certain performance of the microbial community reaches a preset condition (for example, a certain performance is optimal), that is, determining the optimal microhabitat of the microbial community. For example, for a decomposition community for decomposing a certain pollutant, microhabitat identification refers to determining the optimal value range of microhabitat characteristics that have an important impact on the decomposition of the pollutant as described. To solve the problems existing in the background art, embodiments of the present invention provide a microhabitat identification method based on machine learning, capture the complex non-linear relationship between the microbial community and the microhabitat using the machine learning method, optimize the evaluation index of the microbial community based on the relationship between the two, analyze and identify the values of microhabitat characteristics, and can improve the identification accuracy of the microhabitat. Hereinafter, the technical solutions provided by the embodiments of the present invention will be described in detail with reference to the accompanying drawings. The main idea of the microhabitat identification method based on machine learning provided by the embodiments of the present invention is as follows: training a microbial community evaluation model based on a training sample set, identifying the values of target microhabitat characteristics that have an important impact on the microbial community based on the microbial community evaluation model, and the values of the target microhabitat characteristics can optimize the evaluation index of the microbial community. The method provided by the embodiments of the present invention is executed by an electronic device having a processing function. For example, the electronic device may be a computer, a server, etc. As shown in FIG. 1, the microhabitat identification method based on machine learning provided by the embodiments of the present invention includes the following steps: S100, constructing a training sample set for training a microbial community evaluation model. The microbial community evaluation model is a machine learning model that shows the relationship between the microhabitat characteristics of the microbial community and the evaluation index of the microbial community. Optionally, in an embodiment of the present invention, the microbial community evaluation model is a gradient boosting regression tree (Gra dient boosting regression tree, GBRT), and G BRT can continuously construct a plurality of decision tree models, gradually optimize the loss function, and improve the prediction accuracy of the entire model . Here, the evaluation index is used to evaluate a certain characteristic of the microbial community, and different evaluation indexes can evaluate different characteristics (such as correlation, diversity, difference, etc.). In an embodiment of the present invention, the evaluation index of the microbial community may be an index for evaluating the interaction between microorganisms in the microbial community (such as a network topology coefficient or other indexes), and in the following embodiments, it is referred to as the interaction evaluation index of the microbial community . It should be noted that according to actual needs, the evaluation indexes of the microbial community are also diverse. Optionally, the evaluation index of the microbial community may be an index for evaluating the species diversity of the microbial community (such as α-diversity), or an index for evaluating the difference between different microbial communities (such as β-diversity), etc. It should be understood that this is also possible . It should be noted that in an embodiment of the present invention, mainly taking the interaction evaluation index of the microbial community as an example, a microhabitat identification method based on machine learning will be described. The interaction evaluation index is obtained based on the interaction network of the microbial community, and specifically, it will be described in detail in the following embodiments . According to the description of the above embodiments, the interaction evaluation index of the microbial community is used to indicate the interaction between microorganisms in the microbial community, and the interaction between microorganisms includes a positive correlation relationship and a negative correlation relationship . For example, taking the abundance of species as an example, when the relationship between microorganism A and microorganism B is a positive correlation relationship , when the abundance of microorganism A is large, the abundance of microorganism B is also large, and the relationship between microorganism A and microorganism B ... When there is a negative correlation, if the abundance of microorganism A is large, the abundance of microorganism B is small. In an embodiment of the present invention, the above training sample set is obtained from a sewage treatment plant, and the sewage treatment plant collects a plurality of microbial communities (for example, a total of 1068 water samples or sludge samples), for example, the water samples and sludge samples are collected from each biochemical tank of 177 sewage treatment plants across the country, and the biochemical tanks include anaerobic tanks, anoxic tanks and aerobic tanks. . The microhabitat characteristics that affect the microbial community in the sewage treatment plant include the physical, chemical, biological and other factors and process operation parameters listed in the above embodiments. Specifically, , the microhabitat characteristics include temperature (abbreviation: T, unit: °C), pH, total nitrogen (TN, mg / L), ammonia nitrogen (NH4 , mg / L), total phosphorus (TP, mg / L), total organic carbon + (TOC, mg / L), dissolved oxygen (DO, mg / L), sludge age (SRT, h) and hydraulic retention time (HRT, h). The training sample set for training the above microbial community evaluation model includes the values of the microhabitat characteristics of each of the plurality of microbial communities in the sewage treatment plant and the evaluation indicators of the mutual relationship of each of the plurality of microbial communities. Therefore, the construction of the training sample set includes the following, obtaining the values of the microhabitat characteristics of the microbial community and the evaluation indicators of the mutual relationship of the microbial community. . In one embodiment, for each microhabitat characteristic, the process parameters may refer to the process parameters of the sewage plant, and the values of other microhabitat characteristics may be obtained by analyzing with water and wastewater monitoring and analysis methods. In one embodiment, the process of calculating the evaluation index of the microbial community relationship includes S101 to S102. . S101. For each microbial community among a plurality of microbial communities in the sewage treatment plant, obtain the community information of each microbial community. The community information includes the types and the abundance of the types included in each microbial community. The abundance of a certain type of microorganism is used to indicate the abundance degree of the microorganism in the microbial community. In the embodiments of the present invention, the method for obtaining the community information of each microbial community includes the following: Determine the base sequence of each microbial community using 16S rRNA amplicon sequencing, compare the obtained sequence data of each microbial community with the 16S rRNA database, determine the types of microorganisms included in the microbial community, calculate the abundance of the types, and the abundance of the types is the value obtained by dividing the number of occurrences of the types by the total number of occurrences of all types. For example, perform 16S rRNA amplicon sequencing on a sludge sample, amplify the 16S rRNA V3-V4 region of bacteria (i.e., microorganisms) in the sludge sample by polymerase chain reaction (PCR), obtain the DNA sequence in the V3-V4 region of bacteria in the sludge sample. The V3-V4 region of bacterial 16S rRNA is two highly variable regions in the 16S rRNA gene. Further, split the amplified DNA sequence into two parts for sequencing using high-throughput sequencing technology (such as the Illumina MiSeq platform) to obtain the sequence data of bacteria in the sludge sample. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. For example, perform 16S rRNA amplicon sequencing on a sludge sample, amplify the 16S rRNA V3-V4 region of bacteria (i.e., microorganisms) in the sludge sample by polymerase chain reaction (PCR), obtain the DNA sequence in the V3-V4 region of bacteria in the sludge sample. The V3-V4 region of bacterial 16S rRNA is two highly variable regions in the 16S rRNA gene. Further, split the amplified DNA sequence into two parts for sequencing using high-throughput sequencing technology (such as the Illumina MiSeq platform) to obtain the sequence data of bacteria in the sludge sample. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. For example, perform 16S rRNA amplicon sequencing on a sludge sample, amplify the 16S rRNA V3-V4 region of bacteria (i.e., microorganisms) in the sludge sample by polymerase chain reaction (PCR), obtain the DNA sequence in the V3-V4 region of bacteria in the sludge sample. The V3-V4 region of bacterial 16S rRNA is two highly variable regions in the 16S rRNA gene. Further, split the amplified DNA sequence into two parts for sequencing using high-throughput sequencing technology (such as the Illumina MiSeq platform) to obtain the sequence data of bacteria in the sludge sample. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. The 16S rRNA primers used for 16S rRNA amplicon sequencing include 341F and 806R. 341F: CCTAYGGGRBGCASCAG, as shown in SEQ ID NO: 1. 806R: GGACTACNNGGGTATCTAAT, as shown in SEQ ID NO: 2 is shown. Furthermore, the process of comparing the sequence data of each microbial community with the 16S rRNA database includes the following steps: clustering the sequence data of bacteria in the sludge sample, grouping similar sequences as taxonomic units (OTUs), then comparing the sequence data with the known 16S rRNA database to determine the classification information of each OTU, and thus grasping the species composition of bacteria in the sample. After determining the types of bacteria in the above sludge sample, count the number of occurrences of each bacterium and the total number of occurrences of all bacteria, calculate the ratio of the number of occurrences of a certain type to the total number of occurrences of all types for a certain type, and obtain the abundance of that type. In this way, the community information of the microbial community is obtained. can be obtained. After determining the types of bacteria in the sludge sample, count the number of occurrences of each bacterium and the total number of occurrences of all bacteria, calculate the ratio of the number of occurrences of a certain type to the total number of occurrences of all types for a certain type, and obtain the abundance of that type. In this way, the community information of the microbial community is obtained. After determining the types of bacteria in the sludge sample, count the number of occurrences of each bacterium and the total number of occurrences of all bacteria, calculate the ratio of the number of occurrences of a certain type to the total number of occurrences of all types for a certain type, and obtain the abundance of that type. In this way, the community information of the microbial community is obtained. After determining the types of bacteria in the sludge sample, count the number of occurrences of each bacterium and the total number of occurrences of all bacteria, calculate the ratio of the number of occurrences of a certain type to the total number of occurrences of all types for a certain type, and obtain the abundance of that type. In this way, the community information of the microbial community is obtained. S102. Based on the types and abundances of types included in each microbial community, construct an interaction network of each microbial community using Spiec-Easi, and use the ratio of the number of positive correlation edges to the total number of edges in the interaction network as the interaction evaluation index of each microbial community. The interaction network is used to reflect the interaction relationship between microorganisms in the microbial community. In the examples of the present invention, after obtaining the types and abundances of types included in the microbial community by the above S101, through literature research, statistical analysis, etc., the functional types, structural types, and co-metabolic types in the microbial community (the microbial community is used to achieve a certain processing task, for example, a decomposition task) can be determined. Here, the functional type is a type that plays a specific important function in the microbial community. For example, a microorganism that can decompose a certain pollutant is a functional type, and the presence or absence of functional types, and the presence is used. In the examples of the present invention, after obtaining the types and abundances of types included in the microbial community by the above S101, through literature research, statistical analysis, etc., the functional types, structural types, and co-metabolic types in the microbial community (the microbial community is used to achieve a certain processing task, for example, a decomposition task) can be determined. Here, the functional type is a type that plays a specific important function in the microbial community. For example, a microorganism that can decompose a certain pollutant is a functional type, and the presence or absence of functional types, and the presence Here, the functional type is a type that plays a specific important function in the microbial community. For example, a microorganism that can decompose a certain pollutant is a functional type, and the presence or absence of functional types, and the presence of functional types, and the presence The quantity directly affects the ecosystem function. The structural species are microorganisms that play a role in constructing and stabilizing the composition structure of the microbial community. For example, the microorganisms responsible for the basal metabolism of the microbial community are structural species. The cometabolic species are the types of species that decompose target pollutants by cometabolism in the microbial community. For example, the microbial community is a degradation community that degrades the pollutant sulfamethoxazole (SMX). The microbial community capable of decomposing SMX (abbreviated as SMX degradation community) includes functional species capable of decomposing SMX, structural species responsible for basal metabolism, and types of species that decompose SMX by cometabolism. In one embodiment, based on the functional species, structural species, cometabolic species, and abundance of each type in the microbial community, Spiec-Easi (Sparse InversE Covariance estimation for Ecological Association and Statistical Inference, sparse InversE covariance estimation for ecological association and statistical inference) is used to construct the interaction network of the microbial community (also called the microbial ecological network), and based on the interaction network, an evaluation index of the mutual relationship of the microbial community is calculated. Spiec-Easi is a tool for estimating the microbial ecological network from 16S rRNA amplicon resequence datasets. More details on constructing the interaction network using Spiec-Easi can be found in the prior art and are omitted here. Note that the above interaction network includes multiple nodes, each node corresponds to one type, the connection line between nodes is called an edge, and the edge represents the mutual relationship (positive correlation) between two types. indicating a (including a positive correlation relationship and a negative correlation relationship), and in this way, the edges of the interaction network include positive correlation edges and negative correlation edges. If there is no connection line between nodes, it indicates that there is no correlation between the two types. Statistically count the number of positive correlation edges and the total number of edges in the interaction network, calculate the correlation evaluation index of the microbial community, and the larger the value of the evaluation index, the more positive correlation relationships there are among the microorganisms in the microbial community, and the higher the stability and functionality of the microbial community. Using S200, take the values of the microhabitat characteristics of each microbial community in the training sample set as input values, and take the correlation evaluation index of each microbial community as the output value, and use a machine learning algorithm to train a microbial community evaluation model. In an embodiment of the present invention, input the values of the microhabitat characteristics of the microbial community into the microbial community evaluation model, and the model outputs a predicted value of the correlation evaluation index of the microbial group. Based on the predicted value and the correlation evaluation index of the microbial community in the training sample set (i.e., the true value of the evaluation index), calculate the value of the loss function (i.e., the model loss), and optimize the hyperparameters of GBRT based on the model loss until the model error meets the conditions or reaches the preset number of training times, and obtain a converged microbial community evaluation model. In one embodiment, optimize the hyperparameters of GBRT by Bayesian optimization to obtain a converged model. Bayesian optimization is an efficient global optimization method. For the process of optimizing model parameters by Bayesian optimization, refer to the relevant descriptions in the prior art. Optionally, after the training of the above microbial community evaluation model is completed, several evaluation indexes can be used to evaluate the performance of the model to verify whether the performance of the model is good. Any of the following The performance index evaluates the performance of the microbial community evaluation model. For example, the R of the true value and the predicted value 2 , the true value and the mean absolute error (MAE) of the predicted value, or the root mean square error (RMSE) of the true value and the predicted value is used. R . For details of R, MAE, and RMSE, reference may be made to prior art materials and will not be described in detail in the embodiments of the present invention. 2 Illustratively, in (a), (b), and (c) of FIG. 3, the figures in the right column are the anaerobic tank, the anoxic tank, and the aerobic tank, respectively, showing the performance indexes (R ), MAE, and RMSE) of the trained microbial community evaluation model. Taking (a) of FIG. 3 as an example, for the microbial community from the anaerobic tank, in the figure in the right column, the blue straight line shows the relationship between the predicted value and the true value for the trained microbial community evaluation model, and the black straight line shows the relationship between the predicted value and the true value under ideal conditions (the predicted 2 value is equal to the true value). S300. Determine the target microhabitat characteristics of the microbial community (i.e., the microbial community of the sewage treatment plant), and the influence degree on the mutual relationship evaluation index of the microbial community with respect to the target microhabitat characteristics is higher than the influence degree on the mutual relationship evaluation index of the microbial community with respect to other microhabitat characteristics of the microbial community. Based on the description of the microhabitat characteristics in the above embodiments, the microhabitat characteristics of the microbial community are various, and here, some microhabitat characteristics have an important impact on the overall performance of the microbial community. Optionally, as shown in FIGS. 1 and 2, determine the target microhabitat characteristics of the microbial community of the sewage treatment plant by S301 to S302. S301. For each microhabitat feature of the microbial community, determine the importance index of the microhabitat feature and obtain the importance indices of multiple microhabitat features of the microbial community. The importance index of each microhabitat feature is the degree of influence on the mutual relationship evaluation index of the microbial community of each microhabitat feature, that is, the relative importance of the microhabitat feature to the microbial community. It is used to indicate the relative importance. The importance index of the microhabitat feature is the absolute value of the difference between the performance index of the first microbial community evaluation model and the performance index of the second microbial community evaluation model. Here, the first microbial community evaluation model is obtained by training all microhabitat features of the microbial community, and the second microbial community evaluation model is obtained by training the microhabitat features after removing the microhabitat feature. Specifically, for multiple microbial communities in the training sample set, based on the training sample set including all microhabitat features of the microbial community, train the first microbial community evaluation model. Similar to the model obtained in S200 above, remove the microhabitat feature from the multiple microhabitat features of each microbial community, and based on the training sample set after removing the microhabitat feature, train the microbial community evaluation model (that is, the second microbial community evaluation model). Determine the absolute value of the difference between the performance index of the first microbial community evaluation model and the performance index of the second microbial community evaluation model. This absolute value is used to determine the performance difference between the two models. The greater the performance difference between the two (that is, the greater the absolute value), the more important the microhabitat feature. Exemplarily, for 1068 microbial community samples from the above-mentioned anaerobic tank, anoxic tank and aerobic tank Taking an example of a sample, for the microhabitat characteristics of temperature (T), pH, total nitrogen (TN), ammonia nitrogen (NH4 + ) , total phosphorus (TP), total organic carbon (TOC), dissolved oxygen (DO), sludge age (SRT) and hydraulic retention time (HRT), respectively determine the importance index (also called relative importance) of the microhabitat characteristics. Referring to Figure 3, in (a), (b) and (c) of Figure 3, the left column of figures shows the relative importance of the microhabitat characteristics of the anaerobic tank, anoxic tank and aerobic tank respectively. Taking Figure 3 (a) as an example, for the microbial community from the anaerobic tank, as can be seen from the left column of figures, total phosphorus is the most important microhabitat characteristic. Of course, the importance index of the microhabitat characteristics may be defined or calculated by other methods, for example, the importance index may be calculated by substituting the importance of the characteristics, and specifically, it may be selected according to the actual needs, and is not particularly limited in the embodiments of the present invention. S302. Use the microhabitat characteristics corresponding to the first n importance indexes with large values among the importance indexes of multiple microhabitat characteristics of the microbial community as the target microhabitat characteristics of the microbial community, where n is an integer greater than or equal to 1 and less than N, and N is the total number of microhabitat characteristics of the microbial community. In the embodiments of the present invention, the importance index of the microhabitat characteristics indicates the relative importance of the microhabitat characteristics to the microbial community, and the greater the value of the importance index, the higher the influence degree on the mutual relationship evaluation index of the microbial community of the microhabitat characteristics. After obtaining the importance indexes of each microhabitat characteristic by S302, the importance index of the microhabitat characteristics Sort the labels (for example, from largest to smallest), and for the first n importance indicators with large values Select the corresponding microhabitat characteristics as the target microhabitat characteristics. Exemplarily, the value of n may be 5. Referring to (c) in Figure 3, in the case of an aerobic tank, 5 The two target microhabitat characteristics are total phosphorus (TP), temperature (T), dissolved oxygen (DO), total nitrogen (TN), and pH. Of course, the value of n may be other values and is not particularly limited in the embodiments of the present invention herein. Based on S400 and the microbial community evaluation model, execute the individual conditional expectation ICE algorithm for the target microhabitat characteristics, and maximize the value of the microbial community correlation evaluation index to determine the value of the corresponding target microhabitat characteristic. Note that identifying the microhabitat of the microbial community means determining the value of the target microhabitat characteristic when the value of the microbial community correlation evaluation index is maximized which means, and at this time, it should be understood that the value of the target microhabitat characteristic is the value of the optimal microhabitat characteristic of the microbial community characteristic. In the embodiments of the present invention, the individual conditional expectation (Ind ividual Conditional Expectation, ICE) algorithm for the target microhabitat characteristics means that when changing the value of the target microhabitat characteristic, according to the output corresponding to the microbial community evaluation model obtain the relationship between the target microhabitat characteristic and the microbial community correlation evaluation index, and based on the relationship between the two, determine the value or value range of the target microhabitat characteristic when maximizing the value of the correlation evaluation index, where the relationship between the two is the relationship graph (ICE graph between the target microhabitat characteristic and the microbial community correlation evaluation index and determine the value or value range of the target microhabitat characteristic corresponding to maximizing the value of the correlation evaluation index based on the relationship between the two, where the relationship between the two is the relationship graph (ICE graph between the target microhabitat characteristic and the microbial community correlation evaluation index, and based on the relationship between the two, determine the value or value range of the target microhabitat characteristic when maximizing the value of the correlation evaluation index, where the relationship between the two is the relationship graph (ICE graph between the target microhabitat characteristic and the microbial community correlation evaluation index), and here, the relationship between the two is the relationship graph between the target microhabitat characteristic and the microbial community correlation evaluation index (ICE graph It is described by (abbreviation). Optionally, the specific realization process of the above S400 includes the following: Step 1: Based on the distribution of the values of the target microhabitat characteristics of the microbial community, generate multiple values of the target microhabitat characteristics. First, for the microbial community in the sewage treatment plant, through data statistics and analysis, determine the distribution of the raw data of the microbial community (i.e., data distribution characteristics), that is, determine the distribution of the values of each target microhabitat characteristic in different ranges. Next, after grasping the distribution of the values of each target microhabitat characteristic, according to the data distribution ratio of the values of the target microhabitat characteristics in different value ranges, randomly generate multiple values of the target microhabitat characteristics. Step 2: Input multiple values of the target microhabitat characteristics into the microbial community evaluation model respectively, obtain multiple evaluation indicators, and generate the ICE graph of the target microhabitat characteristics. Based on Step 1, sequentially input multiple randomly generated values of the target microhabitat characteristics into the microbial community evaluation model, and draw an ICE graph that reflects the relationship between the value of the target microhabitat characteristic and the evaluation indicators of the mutual relationship of the microbial community according to the evaluation indicators output from the model. Step 3: Repeat Steps 1 and 2, obtain multiple ICE graphs of the target microhabitat characteristics, calculate the average value of the multiple ICE graphs, obtain the average ICE graph, and determine the value of the target microhabitat characteristic corresponding to the maximum mutual relationship evaluation indicator of the microbial community from the average ICE graph. For each target microhabitat characteristic, execute the above Steps 1 to 3, and for each target The values of the microhabitat characteristics can be identified. Exemplarily, referring to FIG. 4, a certain figure is taken as an example. Each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. Further, based on the average ICE graph, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index of the microbial community is determined, and the value of the target microhabitat characteristic is the optimal value of the target microhabitat characteristic. Taking a certain figure as an example, each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. Furthermore, based on the average ICE graph, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index of the microbial community is determined, and the value of the target microhabitat characteristic is the optimal value of the target microhabitat characteristic. Taking a certain figure as an example, each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. Furthermore, based on the average ICE graph, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index of the microbial community is determined, and the value of the target microhabitat characteristic is the optimal value of the target microhabitat characteristic. Taking a certain figure as an example, each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. FIG. 4 shows the ICE graphs of different target microhabitat characteristics of different biochemical tanks. Here, the left column of FIG. 4 shows the ICE graphs of five target microhabitat characteristics (such as dissolved oxygen, pH, temperature, total nitrogen, total phosphorus) of the anaerobic tank, respectively. The middle column of FIG. 4 shows the ICE graphs of five target microhabitat characteristics of the anoxic tank, respectively. The middle column of FIG. 4 shows the ICE graphs of five target microhabitat characteristics of the aerobic tank, respectively. Taking a certain figure as an example, each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. Furthermore, based on the average ICE graph, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index of the microbial community is determined, and the value of the target microhabitat characteristic is the optimal value of the target microhabitat characteristic. Taking a certain figure as an example, each line in the figure corresponds to one ICE graph in step 2, and the blue line in FIG. 4 is the average ICE graph. Furthermore, based on the average ICE graph, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index of the microbial community is determined, and the value of the target microhabitat characteristic is the optimal value of the target microhabitat characteristic. Exemplarily, when the target microhabitat characteristics are DO, T, and pH, Table 1 shows the three target microhabitat characteristics and their optimal value ranges of different biochemical tanks. Furthermore, as can be seen from the test, the correlation evaluation index between total phosphorus (TP) and the microbial community is generally a negative relationship (the higher the total phosphorus, the lower the evaluation index), and the correlation evaluation index between total nitrogen (TN) and the microbial community is a positive relationship (the higher the total nitrogen, the higher the evaluation index). Table 1: Table of Optimal Value Ranges of Target Microhabitat Characteristics
[0006] TIFF0007710648000002.tif61152
[0007] Furthermore, as can be seen from the test, the correlation evaluation index between total phosphorus (TP) and the microbial community is generally a negative relationship (the higher the total phosphorus, the lower the evaluation index), and the correlation evaluation index between total nitrogen (TN) and the microbial community is a positive relationship (the higher the total nitrogen, the higher the evaluation index). Furthermore, as can be seen from the test, the correlation evaluation index between total phosphorus (TP) and the microbial community is generally a negative relationship (the higher the total phosphorus, the lower the evaluation index), and the correlation evaluation index between total nitrogen (TN) and the microbial community is a positive relationship (the higher the total nitrogen, the higher the evaluation index). Furthermore, as can be seen from the test, the correlation evaluation index between total phosphorus (TP) and the microbial community is generally a negative relationship (the higher the total phosphorus, the lower the evaluation index), and the correlation evaluation index between total nitrogen (TN) and the microbial community is a positive relationship (the higher the total nitrogen, the higher the evaluation index). Furthermore, as can be seen from the test, the correlation evaluation index between total phosphorus (TP) and the microbial community is generally a negative relationship (the higher the total phosphorus, the lower the evaluation index), and the correlation evaluation index between total nitrogen (TN) and the microbial community is a positive relationship (the higher the total nitrogen, the higher the evaluation index). Furthermore, the optimal value range of the microhabitat characteristics (that is, the value of the target microhabitat characteristic corresponding to the maximum value of the correlation evaluation index) The types of microbial communities (value ranges of microhabitat characteristics when [condition not specified]), and the microhabitat The worst value range of characteristics (the value range of microhabitat characteristics when minimizing the correlation evaluation index) of the microbial community, and it is possible to determine the differential microorganisms under the two ranges. Furthermore, when comparing the optimal value range and the worst value range of the target microhabitat characteristics of different biochemical tanks, since the upregulation of the abundance of functional types is much higher than the downregulation, the optimal value range of the target microhabitat characteristics obtained by the machine learning-based microhabitat identification method can effectively promote the stability and functionality of the microbial community. In summary, the machine learning-based microhabitat identification method provided by the embodiments of the present invention is based on a machine learning method, lives, and uses a microbial community evaluation model to represent the complex non-linear relationship between the microhabitat characteristics of the microbial community and the functional characteristics of the microbial community (the correlation evaluation index indicates the functional characteristics of the microbial community), and further determines the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. Compared with the existing microhabitat identification methods, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat. and further determine the value of the microhabitat characteristics (i.e., the optimal microhabitat) that improves the functionality of the microbial community based on the microbial community evaluation model and the ICE algorithm. When compared with existing microhabitat identification methods, the solution provided by the embodiments of the present invention can clarify the relationship between the microhabitat and the microbial community and improve the identification accuracy of the microhabitat.
[0008] <st26sequencelisting dtdversion="V1_3" filename="機械学習に基づくマイクロハビタ ット同定方法.xml" softwarename="WIPO Sequence" softwareversion="2.3.0" productio ndate="2025-01-20"> <applicationidentification> <ipofficecode>JP< / ipofficecode> <applicationnumbertext / > <filingdate / > < / applicationidentification> <applicantfilereference> 12100000466007458M< / applicantfilereference> <earliestpriorityapplicationidentification> <ipofficecode>CN< / ipofficecode> <applicationnumbertext>CN202411506823.1< / applicationnumbertext> <filingdate> 2024-12-10< / filingdate> < / earliestpriorityapplicationidentification> <applicantname languagecode="ja">Nanjing University< / applicantname> <applicantnamelatin>Nanjing University< / applicantnamelatin> <inventiontitle languagecode="ja">Method for Identifying Microhabitats Based on Machine Learning< / In ventionTitle> <sequencetotalquantity> 2< / sequencetotalquantity> <sequencedata sequenceidnumber="1"> <insdseq> <INSDSeq_length>17< / INSDSeq_length> <INSDSeq_moltype>RNA< / INSDSeq_moltype> <INSDSeq_division>PAT< / INSDSeq_division> <INSDSeq_feature-table> <insdfeature> <INSDFeature_key>source< / INSDFeature_key> <INSDFeature_location>1..17< / INSDFeature_location> <INSDFeature_quals> <insdqualifier> <INSDQualifier_name>mol_type< / INSDQualifier_name> <INSDQualifier_value>other RNA< / INSDQualifier_value> < / insdqualifier> <insdqualifier id="q2"> <INSDQualifier_name>organism< / INSDQualifier_name> <INSDQualifier_value>synthetic construct< / INSDQualifier_value> < / insdqualifier> < / INSDFeature_quals> < / insdfeature> < / INSDSeq_feature-table> <INSDSeq_sequence>cctaygggrbgcascag< / INSDSeq_sequence> < / insdseq> < / sequencedata> <sequencedata sequenceidnumber="2"> <insdseq> <INSDSeq_length>20< / INSDSeq_length> <INSDSeq_moltype>RNA< / INSDSeq_moltype> <INSDSeq_division>PAT< / INSDSeq_division> <INSDSeq_feature-table> <insdfeature> <INSDFeature_key>source< / INSDFeature_key> <INSDFeature_location>1..20< / INSDFeature_location> <INSDFeature_quals> <insdqualifier> <INSDQualifier_name>mol_type< / INSDQualifier_name> <INSDQualifier_value>other RNA< / INSDQualifier_value> < / insdqualifier> <insdqualifier id="q4"> <INSDQualifier_name>organism< / INSDQualifier_name> <INSDQualifier_value>synthetic construct< / INSDQualifier_value> < / insdqualifier> < / INSDFeature_quals> < / insdfeature> < / INSDSeq_feature-table> <INSDSeq_sequence>ggactacnngggtatctaat< / INSDSeq_sequence> < / insdseq> < / sequencedata> < / inventiontitle> < / st26sequencelisting>
[0009]
Claims
1. A method for identifying a microhabitat, comprising: Constructing a training sample set for training a microbial community evaluation model, wherein the training sample set includes the values of the microhabitat characteristics of each of a plurality of microbial communities in a sewage treatment plant and the mutual relationship evaluation index of each of the plurality of microbial communities, and the mutual relationship evaluation index is used to evaluate the mutual relationship between microorganisms in the microbial community, and the micro habitat characteristics include temperature, pH, total nitrogen, ammonia nitrogen, total phosphorus, total organic carbon , dissolved oxygen, sludge age, and hydraulic retention time, Including, Taking the value of the microhabitat characteristics of each microbial community in the training sample set as an input value, Taking the mutual relationship evaluation index of each microbial community as an output value, and training a microbial community evaluation model using a machine learning algorithm, wherein the microbial community evaluation model is a gradient boosting regression tree GBRT Yes, Determining the target microhabitat characteristics of the microbial community, and the influence degree of the target microhabitat characteristics on the mutual relationship evaluation index of the microbial community is higher than that of the influence degree of other micro habitat characteristics of the microbial community on the mutual relationship evaluation index of the microbial community, Based on the microbial community evaluation model, executing an individual conditional expectation ICE algorithm for the target microhabitat characteristics, and determining the value of the target microhabitat characteristics corresponding to the maximum value of the mutual relationship evaluation index of the microbial community. A method for identifying a microhabitat based on machine learning, characterized by the above.
2. Based on the microbial community evaluation model, executing an individual conditional expectation ICE algorithm for the target microhabitat characteristics, and determining the value of the target microhabitat characteristics corresponding to the maximum value of the mutual relationship evaluation index of the microbial community, which is Step 1, based on the distribution of the values of the target microhabitat characteristics of the microbial community, Generating a plurality of values of the target microhabitat characteristics; Step 2, inputting the plurality of values of the target microhabitat characteristics into the microbial community evaluation model respectively to obtain a plurality of evaluation indexes, generating an ICE graph of the target microhabitat characteristics, and the ICE graph is used to describe the relationship between the value of the microhabitat characteristics and the mutual relationship evaluation index of the microbial community. Step 3, repeatedly execute Step 1 and Step 2 to obtain multiple ICE graphs of the target microhabitat characteristics, calculate the average value of the multiple ICE graphs, obtain an average ICE graph, and in the average ICE graph, determine the value of the target microhabitat characteristic corresponding to when the mutual relationship evaluation index of the microbial community is maximized; The method according to claim 1, characterized by including the above.
3. The mutual relationship includes a positive correlation relationship and a negative correlation relationship, which is characterized by claim 1 or 2 The method described in.
4. The calculation process of the mutual relationship evaluation index of each microbial community in the training sample set includes the following. For each microbial community, obtain the community information of each microbial community, and the community information includes the types included in each microbial community and the abundance of the types, Based on the types included in each microbial community and the abundance of the types, use Spec-Ea s i to construct an interaction network of each microbial community. The interaction network reflects the mutual relationship between microorganisms in the microbial community. The ratio of the number of positive correlation edges to the total number of edges in the interaction network is used as the mutual relationship evaluation index of each microbial community. The method according to claim 3, characterized by the above.
5. Determining the target microhabitat characteristics of the microbial community includes the following. For each microhabitat characteristic of the microbial community, determine the importance index of the microhabitat characteristic, and obtain the importance indexes of multiple microhabitat characteristics of the microbial community. The importance index is the absolute value of the difference between the performance index of the first microbial community evaluation model and the performance index of the second microbial community evaluation model. Here, the first microbial community evaluation model is obtained by training based on all microhabitat characteristics of the microbial community, and the second microbial community evaluation model is obtained by training based on the microhabitat characteristics after removing the microhabitat characteristics. The performance indexes of the first microbial community evaluation model and the second microbial community evaluation model are the root mean square error RMSE of the true value and the predicted value, or the mean absolute error MAE of the true value and the predicted value. The microhabitat characteristics corresponding to the first n importance indexes with large values among the importance indexes of the multiple microhabitat characteristics are used as the target microhabitat characteristics. Here, n is a positive integer, and the value of n is determined according to the specific situation of the microbial community. The method according to claim 4, characterized by including the above.
6. The method according to claim 5, characterized in that the method for determining the importance index of the microhabitat characteristic includes: For each microhabitat characteristic of the microbial community, divide the microbial community into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the absolute value of the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model as the importance index of the microhabitat characteristic. The method according to claim 5, characterized in that the method for determining the importance index of the microhabitat characteristic includes: For each microhabitat characteristic of the microbial community, randomly sample multiple times to obtain multiple subsets of the microbial community. For each subset, divide it into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the average value of the absolute values of the differences between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model obtained from multiple samplings as the importance index of the microhabitat characteristic.
7. The method according to claim 1, characterized in that the method for obtaining the average ICE graph includes: For the multiple ICE graphs obtained in Step 3, calculate the average value of the corresponding elements at each position in the multiple ICE graphs to obtain the average ICE graph.
8. The method according to claim 1, characterized in that the method for obtaining the multiple ICE graphs includes: For each microbial community in the training sample set, use the community information of the microbial community to construct an ICE graph of the microbial community. The ICE graph reflects the relationship between the microbial community and the microhabitat characteristic.
9. The method according to claim 1, characterized in that the method for determining the value of the target microhabitat characteristic corresponding to when the mutual relationship evaluation index of the microbial community is maximized in the average ICE graph includes: Search for the position where the mutual relationship evaluation index of the microbial community reaches the maximum value in the average ICE graph. Determine the value of the target microhabitat characteristic corresponding to the position.
10. The method according to claim 1, characterized in that the method for obtaining the community information of the microbial community includes: For each microbial community in the training sample set, use a sequencing technology to sequence the microbial community to obtain the types and abundances of microorganisms included in the microbial community.
11. The method according to claim 1, characterized in that the method for constructing the interaction network of the microbial community using Spec-Ea s i includes: Based on the types and abundances of microorganisms included in the microbial community, use Spec-Ea s i to calculate the interaction strength between microorganisms. Construct an interaction network of the microbial community based on the interaction strength between microorganisms.
12. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, divide the microbial community into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the absolute value of the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model as the importance index of the microhabitat characteristic.
13. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, randomly sample multiple times to obtain multiple subsets of the microbial community. For each subset, divide it into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the average value of the absolute values of the differences between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model obtained from multiple samplings as the importance index of the microhabitat characteristic.
14. The method according to claim 1, characterized in that the method for obtaining the multiple ICE graphs includes: For each microbial community in the training sample set, use a machine learning algorithm to learn the relationship between the microbial community and the microhabitat characteristic to obtain an ICE graph of the microbial community. The ICE graph reflects the relationship between the microbial community and the microhabitat characteristic.
15. The method according to claim 1, characterized in that the method for determining the value of the target microhabitat characteristic corresponding to when the mutual relationship evaluation index of the microbial community is maximized in the average ICE graph includes: Use a search algorithm to search for the position where the mutual relationship evaluation index of the microbial community reaches the maximum value in the average ICE graph. Determine the value of the target microhabitat characteristic corresponding to the position.
16. The method according to claim 1, characterized in that the method for obtaining the community information of the microbial community includes: For each microbial community in the training sample set, use a metagenomic analysis method to analyze the microbial community to obtain the types and abundances of microorganisms included in the microbial community.
17. The method according to claim 1, characterized in that the method for constructing the interaction network of the microbial community using Spec-Ea s i includes: Based on the types and abundances of microorganisms included in the microbial community, use Spec-Ea s i to calculate the interaction probability between microorganisms. Construct an interaction network of the microbial community based on the interaction probability between microorganisms.
18. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, divide the microbial community into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the absolute value of the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model as the importance index of the microhabitat characteristic.
19. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, randomly sample multiple times to obtain multiple subsets of the microbial community. For each subset, divide it into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the average value of the absolute values of the differences between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model obtained from multiple samplings as the importance index of the microhabitat characteristic.
20. The method according to claim 1, characterized in that the method for obtaining the multiple ICE graphs includes: For each microbial community in the training sample set, use a deep learning model to learn the relationship between the microbial community and the microhabitat characteristic to obtain an ICE graph of the microbial community. The ICE graph reflects the relationship between the microbial community and the microhabitat characteristic.
21. The method according to claim 1, characterized in that the method for determining the value of the target microhabitat characteristic corresponding to when the mutual relationship evaluation index of the microbial community is maximized in the average ICE graph includes: Use an optimization algorithm to search for the position where the mutual relationship evaluation index of the microbial community reaches the maximum value in the average ICE graph. Determine the value of the target microhabitat characteristic corresponding to the position.
22. The method according to claim 1, characterized in that the method for obtaining the community information of the microbial community includes: For each microbial community in the training sample set, use a proteomic analysis method to analyze the microbial community to obtain the types and abundances of microorganisms included in the microbial community.
23. The method according to claim 1, characterized in that the method for constructing the interaction network of the microbial community using Spec-Ea s i includes: Based on the types and abundances of microorganisms included in the microbial community, use Spec-Ea s i to calculate the interaction intensity matrix between microorganisms. Construct an interaction network of the microbial community based on the interaction intensity matrix between microorganisms.
24. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, divide the microbial community into a training set and a test set according to a certain proportion. Based on the training set, train the first microbial community evaluation model and the second microbial community evaluation model. Use the test set to calculate the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model. Calculate the absolute value of the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model as the importance index of the microhabitat characteristic.
25. The method according to claim 1, characterized in that the method for determining the importance index of the microhabitat characteristic using the difference between the performance indexes of the first microbial community evaluation model and the second microbial community evaluation model includes: For each microhabitat characteristic of the microbial community, randomly sample multiple times to obtain multiple subsets of the microbial community. For each subset, divide it into a training set and a test set according to a certain proportion. Based on the training set, train the first is 1 or more, n is an integer smaller than N, and N is the total number of microhabitat characteristics of the microbiota The method according to claim 1, characterized in that.
6. Obtaining the community information of each of the microbiota includes the following: Comparing the sequence data of each of the microbiota with a 16S rRNA database to determine the types of microorganisms contained in the microbiota, calculating the abundance of the types, and the abundance of the types is the value obtained by dividing the number of occurrences of the above types by the number of occurrences of all types. The sequence data of the microbiota is sequence data obtained by determining the nucleotide sequence of each of the microbiota using 16S rRNA amplicon sequence. The method according to claim 4, characterized in that. The method according to claim 4, characterized in that. The method according to claim 4, characterized in that. The method according to claim 4, characterized in that. 。
Citation Information
Patent Citations
Habitat environment evaluation system, habitat environment evaluation method and habitat environment evaluation program
JP2015228817A
Bacterial species identification support method, multi-colony learning model generation method, bacterial species identification support device and computer program
JP2022046265A
Methods and systems for microbial pharmacogenomics
JP2022079646A
Information processing system, learning device, method, and program
JP2024010374A
Methods for predicting and generating mixtures of microbiota samples
JP2024516025A