A method for identifying the optimal microhabitat of microbial communities based on machine learning
Through machine learning methods, especially the gradient enhancement regression tree algorithm, combining microbial community and microhabitat information, the optimal range of the optimal microhabitat characteristics of microbial communities is identified, solving the problem that traditional methods cannot capture nonlinear and complex relationships, and improving the accuracy and effectiveness of microhabitat recognition.
Patent Information
- Application Number
- CN202411506823.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing traditional microhabitat recognition methods cannot effectively capture the nonlinear and complex relationships between microbial communities and microhabitat characteristics, and cannot identify the optimal microhabitat conditions.
Machine learning methods, especially gradient-enhanced regression tree algorithm (GBRT), are used to combine microbial community information and microhabitat information to identify the optimal range of key microhabitat characteristics through individual condition expectations analysis.
The capture of the nonlinear complex relationship between microbial communities and microhabitat features is achieved, and the optimal range of intuitively available key microhabitat features is identified, improving the accuracy and effectiveness of microhabitat recognition.
Smart Images

Figure CN119106344B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying the optimal microhabitat of a microbial community, and particularly to a method for identifying the optimal microhabitat of a microbial community based on machine learning, belonging to the fields of bioinformatics and biotechnology. Background Art
[0002] A microbial community is a complex ecosystem composed of different microorganisms, playing a key role in processes such as material cycling and energy flow, and playing an important role in fields such as sustainable agriculture, ecological restoration, and human health. A microhabitat refers to the specific environment in which microorganisms live, including physical, chemical, and biological factors such as temperature, pH value, humidity, and nutrients. These factors of the microhabitat not only affect the microbial diversity and community structure of the microbial community, but also determine the ecological function and adaptability of the microbial community. Therefore, identifying microhabitat conditions is of great significance for stabilizing or regulating the microbial community.
[0003] Traditional microhabitat identification methods are mainly statistical methods, such as regression analysis, analysis of variance (ANOVA), principal component analysis (PCA), structural equation analysis, etc. However, with the research on microbial communities entering the big data era, these methods can no longer meet the requirements for accurately identifying microhabitats: regression analysis cannot capture non-linear and complex relationships; analysis of variance cannot deeply explore the interactions between environmental factors; the results of principal component analysis are unreliable when dealing with non-linear or complex interactions; similarly, correlation analysis cannot capture non-linear and complex relationships, and cannot infer causal relationships based on the correlation results. Currently, there is no method that can analyze the complex non-linear relationship between the microbial community and microhabitat characteristics and identify the optimal microhabitat conditions. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a microhabitat identification method based on machine learning, which uses machine learning to capture non-linear and complex relationships in the microbial community, and identifies the optimal range of intuitive and available key microhabitats through individual conditional expectation analysis.
[0005] Technical Solution: The present invention provides a method for identifying the optimal microhabitat of a microbial community based on machine learning, including the following steps:
[0006] First step, obtain microbial community information and microhabitat information, and calculate microbial community evaluation indicators based on the microbial community information;
[0007] Second step, use the microbial community evaluation indicators as predicted values and the microhabitat information as feature values to train a machine learning model;
[0008] Third step, evaluate the performance of the machine learning model, rank the importance of microhabitat features, and screen key microhabitat features;
[0009] In the fourth step, based on the machine learning model, the optimal range of the key microhabitat features is obtained by using the individual conditional expectation.
[0010] The present invention uses machine learning to establish a microhabitat recognition method, which can capture the non-linear complex relationship between the microbial community and the microhabitat features, and obtain the key microhabitat features and the optimal range based on the analysis of the individual conditional expectation. In the first step, the microbial community information includes the species type and abundance, and the microhabitat information includes the physical and chemical factors and geographical information that affect the microbial community structure. In the second step, the microbial community evaluation indexes include commonly used indexes in the field such as α-diversity; the machine learning model includes commonly used models such as the gradient boosting regression tree algorithm. In the third step, the performance evaluation of the trained model helps to optimize the model to improve the prediction accuracy. In the fourth step, the constructed machine learning model is used to identify the optimal microhabitat features, and the optimal range of the intuitive and usable key microhabitat features is identified through the individual conditional expectation (ICE) analysis that can reflect the relationship between the predicted value of each individual and a single variable.
[0011] Preferably, in the first step, the microbial community information is obtained through open data sources or sequencing technologies. For example, 16S rRNA amplicon sequencing.
[0012] Preferably, in the first step, the microhabitat information includes the physical and chemical factors and geographical information that significantly affect the microbial community structure. For example, total nitrogen, ammonia nitrogen, nitrate nitrogen, nitrite nitrogen, water-soluble organic nitrogen, water temperature, pH, total phosphorus, organic carbon, longitude and latitude, etc.
[0013] Preferably, in the first step, the microbial community evaluation indexes include α-diversity, β-diversity or the microbial flora evaluation indexes based on the topological structure of the microbial co-occurrence network.
[0014] Preferably, in the second step, the machine learning model is the gradient boosting regression tree algorithm (GBRT). The GBRT algorithm is suitable for large datasets, can process various types of data, and has high prediction accuracy and generalization ability.
[0015] Preferably, in the second step, 75% to 85% of the predicted values and feature values are used as the training set for the machine learning model, and the rest are used as the validation set. Preferably, 75% of the data is used as the training set.
[0016] Preferably, in the second step, the machine learning model uses Bayesian optimization for hyperparameter optimization.
[0017] Preferably, in the second step, the machine learning model is subjected to several cross-validations. Preferably 10 times.
[0018] Preferably, in the third step, the evaluation includes using R 2 to evaluate the model fitness, or using the mean absolute error (MAE) and root mean square error (RMSE) for evaluation.
[0019] Preferably, in the third step, the importance ranking is obtained based on permutation feature importance (PFI). The top 5 features with preferred importance ranking are used as key microhabitat features.
[0020] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention uses a machine learning model to capture the non-linear and complex relationships in the microbial community, and identifies the optimal range of intuitive and available key microhabitat features through individual conditional expectation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic flow chart of the method for identifying the optimal microhabitat of the microbial community provided by the present invention;
[0022] Figure 2 is the identification and performance evaluation of key microhabitat features using the gradient boosting regression tree algorithm in the embodiment of the present invention;
[0023] Figure 3 is the optimal range of 5 key microhabitat features in the embodiment of the present invention;
[0024] Figure 4 is the verification result of the optimal range of key microbial features in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0026] This embodiment provides a method for identifying the optimal microhabitat of a microbial community based on machine learning. The microbial community is collected from a sewage biological treatment system, and its process parameters, water samples and sludge samples are collected from the biochemical tanks of 177 sewage treatment plants across the country, including anaerobic tanks, anoxic tanks and aerobic tanks, with a total of 1068 water samples and sludge samples.
[0027] The method for identifying the optimal microhabitat of the microbial community based on machine learning is as Figure 1 shown, and specifically includes the following steps:
[0028] In the first step, obtain the microbial community information and microhabitat information, and calculate the microbial community evaluation index based on the microbial community information.
[0029] The microbial community information, including species types and abundances, is derived from sequencing and literature research. The sludge samples are subjected to 16S rRNA amplicon sequencing to determine the V3-V4 region of bacterial 16S rRNA. The 16S primers used for sequencing are: 341F (CCTAYGGGRBGCASCAG); 806R (GGACTACNNGGGTATCTAAT). Through literature research, statistical analysis, etc., a sulfamethoxazole (SMX) degradation community composed of functional species, structural species, and co-metabolic species was determined. The microbial network topology coefficient was selected as the microbial community evaluation index, and the Spiec-Easi was used to construct a biological association network to obtain the microbial network topology coefficient.
[0030] The microhabitat information includes physical and chemical factors and geographical information that affect the microbial community structure: total nitrogen (TN), ammonia nitrogen (NH4+), temperature (T), pH, total phosphorus (TP), total organic carbon (TOC), dissolved oxygen (DO), sludge retention time (SRT), and hydraulic retention time (HRT), etc. It is mainly analyzed with reference to the Methods for Monitoring and Analysis of Water and Wastewater; the process parameters are specifically referenced from the sewage biological treatment system.
[0031] In the second step, the microbial network topology coefficient is used as the predicted value, and the microhabitat information is used as the characteristic value to train the machine learning model Gradient Boosting Regression Tree (GBRT).
[0032] Bayesian optimization and 10-fold cross-validation are used to optimize the GBRT hyperparameters. The microbial community evaluation index and water quality characteristics are input into the machine learning model according to the training set: predicted value of 75%: 25%.
[0033] In the third step, the performance of the machine learning model is evaluated, and the importance of the microhabitat characteristics is ranked to screen the key microhabitat characteristics.
[0034] Use R 2 , Mean Absolute Error (MAE), and Root Mean Square Error (RMSE) to evaluate the performance of GBRT. The top 5 microhabitat characteristics ranked by importance are called key microhabitat characteristics, as shown in Figure 2 and Figure 3 shown.
[0035] In the fourth step, based on the machine learning model, Individual Conditional Expectation (ICE) is used to obtain the optimal range of the key microhabitat characteristics.
[0036] The obtained key microhabitat characteristics and their optimal ranges are as follows: the optimal ranges of dissolved oxygen are respectively anaerobic tank: 0.36 - 0.38 mg / L, anoxic tank: 0.72 - 0.73 mg / L, aerobic tank: 4.82 - 4.89 mg / L; the optimal ranges of temperature are anaerobic tank: less than 15.73 °C, anoxic tank: less than 15.56 °C, aerobic tank: 13.60 - 15.70 °C. The optimal ranges of pH are anaerobic tank: 6.82 - 7.23, anoxic tank: 6.89 - 6.98, aerobic tank: 7.49 - 7.59. There is an overall negative relationship between total phosphorus and the microbial evaluation index, and a positive relationship between total nitrogen and the microbial evaluation index.
[0037] Based on the differences between the optimal range and the worst range of microhabitat characteristics, the optimal range of key microhabitat characteristics was verified by microorganisms. The results are as Figure 4 shown. For different biochemical tanks, the number of significantly upregulated functional species abundances in the optimal range of key microhabitat characteristics compared to the worst range is extremely significantly higher than the number of downregulated ones. The results indicate that the optimal range of key microhabitat characteristics obtained by the optimal microhabitat recognition method of microbial communities based on machine learning can effectively promote the stability and functionality of the target microbial community.
Claims
1. A method for obtaining the optimal range of microhabitat characteristics, characterized in that Calculate the network topology coefficient based on the microbial co-occurrence network topology structure, and use the network topology coefficient as an evaluation index for the microbial community to determine the complex relationship between the microbial community and the microhabitat, and then obtain the optimal range of microhabitat characteristics; The specific method is as follows: S1. Obtain microbial community information and microhabitat information, and calculate the evaluation index of the microbial community based on the microbial community information; The microhabitat information includes physicochemical factors and geographical information that affect the microbial community structure; The microbial community information can be obtained through open data sources or sequencing technologies; S2. Use the evaluation index of the microbial community as the predicted value and the microhabitat information as the feature value to train a machine learning model; The machine learning model is the gradient boosting regression tree algorithm, and 75% - 85% of the predicted values and feature values are used as the training set, and the rest are used as the validation set; The machine learning model uses Bayesian optimization for hyperparameter optimization and performs several cross-validations; S3. Evaluate the performance of the machine learning model, rank the importance of microhabitat characteristics, and screen key microhabitat characteristics; The importance ranking is obtained based on permutation feature importance; S4. Based on the machine learning model, use individual conditional expectation to obtain the optimal range of the key microhabitat characteristics; S5. Verify the optimal range of the key microhabitat characteristics based on the differential microorganisms between the optimal range and the worst range of the microhabitat characteristics: The number of significantly up-regulated functional species abundances in the optimal range of the key microhabitat characteristics of different biochemical pools is extremely significantly higher than the number of down-regulated ones compared to the worst range. The optimal range of the key microhabitat characteristics obtained by the machine learning-based method for identifying the optimal microhabitat of the microbial community can effectively promote the stability and functionality of the target microbial community.
Citation Information
Patent Citations
Machine learning method for predicting diversity of soil rhizosphere microorganisms
CN117198395A