A method for predicting river water microbial community stability using watershed land use patch density

By setting up a buffer zone within the watershed to calculate the density of land use patches and constructing a machine learning model, the shortcomings of traditional methods in predicting microbial community stability were addressed, and accurate prediction of the stability of microbial communities in river water bodies was achieved, supporting watershed ecological protection and planning.

CN120452556BActive Publication Date: 2025-09-12HOHAI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510940064.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-12
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Traditional methods find it difficult to effectively integrate landscape fragmentation parameters and microbial multidimensional data, cannot accurately predict the microbial stability of river water bodies, and consume a lot of manpower and material resources.

Method used

By setting up a buffer zone in the watershed, calculating the density of land use patches, and combining high-throughput sequencing technology to analyze the microbial 16SrRNA gene, linear regression, quadratic polynomial regression, and random forest regression models were constructed to predict the stability of microbial communities.

Benefits of technology

It has achieved accurate prediction of the stability of microbial communities in river water bodies, provided a quantitative basis for watershed ecological protection, and guided land use planning and ecological management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452556B_ABST
    Figure CN120452556B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the stability of microbial communities in river water bodies by utilizing the patch density of land use in a watershed. The present invention relates to the technical field of microbial community stability prediction and watershed water ecology research, and addresses the deficiencies of traditional research methods in revealing the association and prediction between land use and microbial community stability. The method comprises obtaining land use data of the target watershed, calculating the patch density of each sampling point in the buffer zone; collecting samples to determine the composition of microbial species; preprocessing the data, and performing hierarchical clustering grouping based on the patch density; calculating the average variation dissimilarity of the groups to characterize the stability of the microbial community; conducting statistical analysis and performing a spatial autocorrelation test; comparing the prediction performance of different regression models, and selecting the best model to predict the stability of the microbial community in unsampled areas. This method significantly improves the prediction efficiency and accuracy, and provides reliable technical support for watershed ecological management and land planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microbial community stability prediction and watershed water ecology research, and in particular to a method for predicting the stability of river water microbial communities by utilizing watershed land use patch density. Background Art

[0002] River basin ecosystems are facing landscape fragmentation due to drastic changes in land use patterns caused by expanding urbanization, agricultural development, and intensified industrial activity. At the watershed scale, human activities have significantly altered the migration paths of natural runoff. Pollutants generated throughout the entire watershed are significantly retained due to topographic barriers and soil adsorption, resulting in direct impacts on river aquatic ecosystems primarily in the river's immediate vicinity. Relevant studies analyzing the relationship between land use and water ecology at different scales have found that buffer zone scale plays a significant role in influencing mechanisms. Land use changes within buffer zones can fragment natural habitats into discrete patches. Alterations in the number, morphology, and distribution of these patches indirectly impact the habitat of river microbial communities by altering the physical and chemical properties of water and soil, material exchange processes, and spatial connectivity. As core carriers of ecological functions, microbial community stability directly determines key processes such as material circulation, pollution degradation, and system balance, making it a crucial indicator of watershed health.

[0003] However, current assessments of river water microbial stability lack the ability to reveal land-use-driven mechanisms at the landscape scale. Traditional methods rely on intensive field sampling and laboratory analysis, which consumes significant human and material resources. Furthermore, limited by the lack of representative local data, it is difficult to systematically analyze the spatial associations between land-use patterns and river water microbial stability. Traditional statistical models have limited ability to characterize complex nonlinear relationships and are unable to effectively integrate landscape fragmentation parameters with multidimensional microbial data, resulting in insufficient prediction accuracy. Therefore, it is urgent to develop a technical framework that integrates landscape pattern quantification, machine learning modeling, and spatial analysis to predict river water microbial stability based on land-use patch density, providing a quantifiable decision-making basis for watershed ecological protection and land planning. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for predicting the stability of river water microbial communities using the density of land use patches in a watershed, so as to address the shortcomings of traditional research methods in revealing the association and prediction between land use and microbial community stability.

[0005] The present invention provides a method for predicting the stability of river water microbial communities using the density of land use patches in a river basin, comprising:

[0006] Step 1: Within the target watershed, sampling points are arranged based on topography, land use type, water system distribution, and ecosystem heterogeneity, and the latitude and longitude coordinates of each sampling point are recorded. Using remote sensing images and geographic information systems, the river between the source and the sampling point is divided, a buffer zone is set, land use data within the buffer zone is extracted, and patch density is calculated.

[0007] Step 2: Collect water samples at each sampling point and store them in the laboratory for processing. Use high-throughput sequencing technology to analyze the 16S rRNA gene of microorganisms, generate an operational taxonomic unit abundance table, and perform standardization.

[0008] Step 3: Standardize the patch density and divide the sampling points into groups with similar patch density characteristics based on the Euclidean distance matrix and hierarchical clustering method;

[0009] Step 4: Calculate the Bray-Curtis distance matrix of the microbial communities within each group with similar patch density characteristics, and calculate the average variation dissimilarity based on the Bray-Curtis distance matrix. Use the mean patch density of the group as the independent variable and the average variation dissimilarity as the dependent variable to construct linear regression models, quadratic polynomial regression models, and random forest regression models, respectively. Evaluate model performance and select the optimal prediction model.

[0010] Step 5: Construct a spatial weight matrix of sampling points and test spatial autocorrelation. Based on the optimal model and watershed patch density data, predict and verify the stability of the microbial community.

[0011] Furthermore, in step 1, the formula for calculating plaque density is:

[0012] , where PD is the patch density, N is the number of patches, and A is the landscape area;

[0013] The grid spacing of the grid points is dynamically adjusted according to the basin area to ensure that the sampling points cover the gradient of different land use types; the land use classification adopts the maximum likelihood method, and the classification results are verified by the Kappa coefficient, Kappa>0.80, to ensure that the interpretation accuracy meets the modeling requirements.

[0014] Furthermore, in step 2, 1 L of mixed surface water samples at a depth of 0-20 cm were collected at each sampling point. The water samples were vacuum filtered on site using a 0.22 μm microfiltration membrane. After the filter membrane was sealed, the samples were transported back to the laboratory at 4°C and stored in a -80°C refrigerator.

[0015] Illumina high-throughput sequencing technology was used to analyze the microbial 16S rRNA gene and generate an operational taxonomic unit abundance table. The R language vegan package was used to remove low-abundance operational taxonomic units with a total abundance of <0.01%. The operational taxonomic unit data were standardized using the Hellinger transformation. The Hellinger transformation was performed using the decostand function of the vegan package, and the transformed operational taxonomic unit data were directly used to calculate the Bray-Curtis distance matrix. The conversion formula is:

[0016] ,in, is the converted data, is the original data, is the number of species, is the sample size.

[0017] Furthermore, in step 3, the plaque density is normalized using the formula:

[0018] ,in, is the standardized value, is the original value, is the mean, is the standard deviation;

[0019] The number of groups was determined by the inflection points of the cluster dendrogram; the number of hierarchical clustering groups was verified by the silhouette coefficient method to ensure the consistency of plaque density characteristics within the group and maximize the heterogeneity between groups.

[0020] Furthermore, in step 4, the Bray-Curtis distance matrix of the microbial communities in each group was calculated using the R language vegan package, and the average variation dissimilarity was calculated according to the following formula:

[0021] ,in, is the average variation dissimilarity, is the number of samples within the group, For samples and samples The Bray-Curtis distance between them.

[0022] Furthermore, in step 4, the stats toolkit of R language was used to construct linear regression and quadratic polynomial regression models with the mean of patch density as the independent variable and the mean variation dissimilarity as the dependent variable. At the same time, the randomForest toolkit was used to construct a random forest regression model with the parameter ntree=500. After the model was constructed, the mean square error, mean absolute error and determination coefficient were used as indicators to evaluate the model performance and select the optimal prediction model.

[0023] Furthermore, in step 4, before building the random forest regression model, the sample size needs to be judged: if the sample size is less than 30, in order to ensure the reliability of the model evaluation, the leave-one-out cross-validation method is adopted, and the validation method is set to leave-one-out cross-validation through the trainControl function of the caret package, and the model is trained using the train function; if the sample size is greater than or equal to 30, the data is divided into a training set and a test set in a ratio of 7:3, and the model is trained using the training set, and the model is evaluated using the test set.

[0024] Furthermore, in step five, the spdep package was used to construct a k-nearest-neighbor spatial weight matrix with k=3-5. The model residuals were subjected to Moran's I test with a significance threshold of p<0.05, and the spatial autocorrelation was assessed using Moran's scatter plot. Based on the optimal model and watershed patch density data, the stability of the microbial community in the river water of the unsampled area was predicted, and a spatial distribution map was generated using ArcGIS to guide the identification of ecologically vulnerable areas and land use planning and regulation.

[0025] Furthermore, in step 5, if the Moran index test shows significant spatial autocorrelation, p < 0.05, the spatial error model is used to correct the prediction results. The model is implemented by the errorsarlm function of the spdep package.

[0026] Furthermore, in step five, when the prediction results are compared and verified with the field sampling data, the prediction error is required to be ≤15%, otherwise the model is re-optimized or the sampling data is supplemented.

[0027] The present invention has the following beneficial effects: The method of the present invention for predicting the stability of river water microbial communities using the density of watershed land use patches establishes a connection between the density of watershed land use patches and the stability of river water microbial communities, and uses scientific and rigorous methods to explore the intrinsic relationship between the two, opening up new ideas and approaches for watershed ecological research. By quantifying the impact of land use patch density on the stability of river water microbial communities and combining it with machine learning modeling to construct a prediction model, an effective prediction can be made of the trend of changes in microbial community stability. This not only provides strong technical support for the monitoring of watershed ecosystems, but also provides a quantitative basis for watershed land use planning and ecological management. Based on the prediction results, managers can adjust land use strategies in a timely manner before abnormal fluctuations occur in the stability of water body microbial communities, take targeted ecological protection measures, and effectively prevent the imbalance and degradation of watershed ecosystems, which has extremely important practical guiding value for promoting the sustainable development of watershed ecology. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0029] Figure 1 It is a schematic flow diagram of the present invention;

[0030] Figure 2 It is the visualization of polynomial regression of average patch density and river water microbial community stability in the Min-Tuo River Basin;

[0031] Figure 3 It is the random forest regression prediction of the average patch density and river water microbial community stability in the Min-Tuo River Basin. DETAILED DESCRIPTION

[0032] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. The technical solutions provided by each embodiment of the present invention are described in detail below in conjunction with the drawings.

[0033] See also Figures 1 to 3 The method provided by the present invention for predicting the stability of river water microbial communities by using the density of land use patches in a river basin comprises the following steps:

[0034] 1. Sampling point determination and data acquisition

[0035] Based on technical specifications such as the "Technical Guidelines for River Ecological Security Survey and Assessment," "Technical Specifications for River Ecological Survey," "Technical Guidelines for Ecological Environmental Protection of Rivers, Lakes, and Lakes," "Technical Guidelines for River Water Ecological Environmental Quality Monitoring and Evaluation," and "Technical Guidelines for Ecological Security Survey and Assessment of Lakes," basic data was collected for the Chengdu section of the Mintuo River. This included the natural environmental characteristics of the river and lake, socioeconomic characteristics, hydrological and water system characteristics, and land use characteristics. Geographic Information System (GIS) technology was used to overlay and analyze various data layers based on land use types within the Mintuo River basin, including cultivated land, forest land, construction land, and water areas; topographical zoning such as mountainous, plain, and hilly areas; and hydrological characteristics such as flow velocity, discharge, water depth, and river channel width. Fifty-three sampling points were identified, and water samples were collected during the flood season. High-precision GPS positioning equipment was used to measure and record the latitude and longitude coordinates of each sampling point to six decimal places.

[0036] Landsat ETM+ remote sensing imagery of the basin was obtained from the USGS website. Image acquisition was performed during the vegetation growing season, during clear, cloud-free weather, to ensure image quality. The remote sensing imagery was preprocessed using the Environment for Visualizing Images (ENVI) remote sensing image processing platform, performing radiometric calibration, atmospheric correction, and geometric correction. Radiometric calibration converts sensor-recorded digital numbers (DN) into actual surface radiance values. Atmospheric correction uses specialized algorithms to remove the effects of atmospheric scattering and absorption. For geometric correction, 45 evenly distributed ground points on a topographic map were selected as ground control points. Quadratic polynomials and cubic convolution interpolation were used to geometrically correct the imagery to ensure spatial accuracy.

[0037] Using a remote sensing image processing platform, supervised classification was performed on preprocessed remote sensing images using the maximum likelihood method. Referring to the "Current Land Use Classification" (GB / T21010-2017), land use types were subdivided into nine categories: cultivated land, forest land, construction land, water area, grassland, wetland, unused land, transportation land, and special land. After classification, 500 ground truth sample points were randomly selected. The accurate land use types of these sample points were confirmed through field surveys and historical data inquiries. The Kappa coefficient was calculated using the accuracy verification tool in ENVI software. A Kappa coefficient greater than 0.80 indicates that the classification results are reliable and accurately reflect the land use status of the watershed.

[0038] Previous studies have shown that at the watershed scale, the impact of the buffer zone on the relationship between land use and water ecology is more prominent. Therefore, this study used ArcGIS software to divide the river between the source and the sampling point, set a buffer zone of 1000m, and extracted land use data within the buffer zone. The processed remote sensing images were imported into Fragstats, and the formula was used: , calculate patch density (PD), where N is the number of patches and A is the landscape area in hectares (hm²). Land use patch classification results are validated using the Kappa coefficient, with a Kappa value > 0.80 to ensure interpretation accuracy meets subsequent modeling requirements.

[0039] 2. Microbial Data Collection and Processing

[0040] At the selected sampling point, a sterile sampling device was used to collect 1 L of mixed surface water samples at a depth of 0-20 cm. The water samples were vacuum filtered on-site using a 0.22 μm microfiltration membrane. After the filter membrane was sealed, the samples were transported back to the laboratory at 4°C and stored in a -80°C refrigerator pending subsequent microbial composition determination to ensure that the integrity and activity of the microbial samples were not affected.

[0041] Illumina high-throughput sequencing technology was used to analyze the 16S rRNA gene of the microorganisms on the filter membrane. First, microbial DNA was extracted from the filter membrane using a specialized deoxyribonucleic acid (DNA) extraction kit. The extraction process strictly followed the kit instructions to ensure that the extracted DNA quality and purity met sequencing requirements. After extraction, DNA quality and concentration were tested by agarose gel electrophoresis and a nucleic acid concentration analyzer. The extracted DNA was then amplified by polymerase chain reaction (PCR). Primers specific for the 16S rRNA gene were selected, and the amplification system and conditions were optimized to ensure specificity and efficiency. After purification, sequencing libraries were constructed and high-throughput sequencing was performed on the Illumina sequencing platform. The sequencing data were processed using a professional bioinformatics analysis pipeline, including quality control, sequence assembly, denoising, and species annotation, ultimately generating an operational taxonomic unit (OTU) abundance table.

[0042] The OTU abundance table was processed using the vegan package in R language to remove low-abundance OTUs with a total abundance of <0.01%. The filtered OTU data were subjected to Hellinger transformation using the decostand function of the vegan package. The transformation formula is: ,in, is the converted data, is the original data, is the number of species, The converted OTU data were directly used to calculate the Bray-Curtis distance matrix.

[0043] 3. PD Characteristic Analysis and Grouping

[0044] The PD values ​​calculated at each sampling point are imported into the R language environment and standardized using the stats toolkit. ,in, is the standardized value, is the original value, is the mean, is the standard deviation.

[0045] The dist function is used to calculate the Euclidean distance between the standardized PD values ​​and construct a distance matrix. The hclust function is used to perform hierarchical cluster analysis using the Ward method to generate a cluster dendrogram. The number of groups is preliminarily determined by observing the inflection point of the cluster dendrogram. At the same time, the silhouette coefficient method is used to verify the preliminarily determined number of groups, and the silhouette coefficient under different grouping numbers is calculated. The number of groups corresponding to the maximum silhouette coefficient is selected as the final number of groups. According to the determined number of groups, the sampling points are divided into PD feature similarity groups. The embodiment is divided into 8 groups.

[0046] 4. Quantification and Modeling of Microbial Stability

[0047] The Bray-Curtis distance matrix of the microbial community in each group was calculated using the R language vegan package. , calculate the average variation dissimilarity (AVD) of the microbial community in each group, where is the number of samples within the group, For samples and samples After the calculation is completed, the AVD value, the corresponding group information, the PD average value, etc. are integrated into a data frame.

[0048] The stats toolkit of R language was used to construct the regression model. The mean of group PD was used as the independent variable and AVD was used as the dependent variable to construct a linear regression model and a quadratic polynomial regression model. The lm function was used for model fitting. The formula for constructing the linear regression model is: , the quadratic polynomial regression model formula is .

[0049] At the same time, the random forest regression model is constructed with the help of the randomForest toolkit, the parameter ntree=500 is set, and the randomForest function is used to construct the model. Setting ntree to 500 means using 500 trees to build the model, which helps improve model stability and predictive accuracy. Before building a random forest regression model, it is necessary to determine the sample size: if the sample size is less than 30, to ensure the reliability of model evaluation, use Leave-One-Out Cross Validation (LOOCV) by setting the validation method to LOOCV using the trainControl function in the caret package, and use the train function to train the model; if the sample size is greater than or equal to 30, divide the data into a training set and a test set in a ratio of 7:3. Use the training set to train the model, and the test set to evaluate the model.

[0050] To evaluate the performance of each model, metrics such as mean square error (MSE), mean absolute error (MAE), and coefficient of determination (R²) were used. The MSE and MAE functions in the metrics package were used to calculate the MSE and MAE, respectively, for each model. For linear regression and quadratic polynomial regression models, R² was obtained using the summary function. For random forest regression models, if leave-one-out cross-validation was used, R² was calculated using the postResample function based on the predicted and true values. If a 7:3 data split was used, the same calculation was performed on the test set. The MSE measures the average of the squared errors between the model's predicted values ​​and the true values. A smaller MSE value indicates a smaller deviation between the model's predicted values ​​and the true values. The MAE reflects the average of the absolute errors between the model's predicted values ​​and the true values. A smaller MAE value indicates a smaller overall model prediction error. R² is used to assess the goodness of fit of the model to the data; an R² closer to 1 indicates a better fit. By comparing these indicators of different models, we can intuitively understand the prediction error and fitting effect of each model, and then select the model with the smallest prediction error and the best fitting effect as the optimal prediction model for subsequent prediction analysis of the stability of water microbial communities.

[0051] 5. Spatial autocorrelation analysis and prediction verification

[0052] The spdep package of R language is used to construct a k-nearest neighbor spatial weight matrix based on the latitude and longitude coordinate information of the sampling points, and the value of k is 3-5. For example, when k=5, the 5 nearest neighbor points of each sampling point are calculated by the knearneigh function, the knn2nb function converts the neighbor relationship into a network object, and the nb2listw function converts the network object into a spatial weight matrix. The residuals of the linear regression model are tested with the Moran's Index (Moran'sI) and the moran.test function is used for testing. The Moran'sI test results in the embodiment are: I=-0.033, p=0.2803. If the test result shows p<0.05, it indicates that there is significant spatial autocorrelation, and the spatial error model (SpatialErrorModel, SEM) is used to correct the prediction results, and the errorsarlm function of the spdep package is used to fit the spatial error model.

[0053] Based on an optimal model, such as quadratic polynomial regression (poly_model), and known watershed PD data, the stability of water microbial communities in unsampled areas is predicted. The prediction results are imported into ArcGIS software, and spatial analysis functions are used to generate a spatial distribution map of river water microbial community stability. Field sampling is performed at specific verification points in unsampled areas, and the AVD values ​​of these verification points are measured. The prediction results are compared with the field sampling data, and the mean absolute error (MAE) is calculated. If the MAE is ≤15%, the prediction result is considered reliable. If the prediction error exceeds this threshold, the model parameters are re-optimized or additional sampling data is added, and the prediction and verification are repeated until the accuracy requirements are met.

[0054] Based on the prediction results and spatial distribution maps, ecologically fragile areas with low stability of water microbial communities in the basin are identified, providing a scientific basis for land use planning and regulation, such as restricting development activities in ecologically fragile areas and strengthening ecological protection measures.

[0055] The above-described embodiments of the present invention do not limit the protection scope of the present invention.

Claims

1. A method for predicting the stability of river water microbial communities using the density of land use patches in a watershed, characterized in that: include: Step 1: Within the target watershed, sampling points are arranged based on topography, land use type, water system distribution, and ecosystem heterogeneity, and the latitude and longitude coordinates of each sampling point are recorded. Using remote sensing images and geographic information systems, the river between the source and the sampling point is divided, a buffer zone is set, land use data within the buffer zone is extracted, and patch density is calculated. Step 2: Collect water samples at each sampling point and store them in the laboratory for processing. Use high-throughput sequencing technology to analyze the 16S rRNA gene of microorganisms, generate an operational taxonomic unit abundance table, and perform standardization. Step 3: Standardize the patch density and divide the sampling points into groups with similar patch density characteristics based on the Euclidean distance matrix and hierarchical clustering method; Step 4: Calculate the Bray-Curtis distance matrix of the microbial communities within each group with similar patch density characteristics, and calculate the average variation dissimilarity based on the Bray-Curtis distance matrix. Use the mean patch density of the group as the independent variable and the average variation dissimilarity as the dependent variable to construct linear regression models, quadratic polynomial regression models, and random forest regression models, respectively. Evaluate model performance and select the optimal prediction model. Step 5: Construct a spatial weight matrix of sampling points and test spatial autocorrelation. Based on the optimal model and watershed patch density data, predict and verify the stability of the microbial community.

2. The method for predicting the stability of river water microbial communities using watershed land use patch density according to claim 1, wherein: In step 1, the formula for calculating plaque density is: , where PD is the patch density, N is the number of patches, and A is the landscape area; The grid spacing of the grid points is dynamically adjusted according to the basin area to ensure that the sampling points cover the gradient of different land use types; the land use classification adopts the maximum likelihood method, and the classification results are verified by the Kappa coefficient, Kappa>0.80, to ensure that the interpretation accuracy meets the modeling requirements.

3. The method for predicting the stability of river water microbial communities using watershed land use patch density according to claim 1, wherein: In step 2, 1 L of surface water mixed sample at a depth of 0-20 cm was collected at each sampling point. The water sample was vacuum filtered on site using a 0.22 μm microfiltration membrane. After the filter membrane was sealed, it was transported back to the laboratory at 4°C and stored in a -80°C refrigerator. Illumina high-throughput sequencing technology was used to analyze the microbial 16S rRNA gene and generate an operational taxonomic unit abundance table. The R language vegan package was used to remove low-abundance operational taxonomic units with a total abundance of <0.01%. The operational taxonomic unit data were standardized using the Hellinger transformation. The Hellinger transformation was performed using the decostand function of the vegan package, and the transformed operational taxonomic unit data were directly used to calculate the Bray-Curtis distance matrix. The conversion formula is: ,in, is the converted data, is the original data, is the number of species, is the sample size.

4. The method for predicting the stability of river water microbial communities using watershed land use patch density according to claim 1, wherein: In step 3, the plaque density is normalized using the formula: ,in, is the standardized value, is the original value, is the mean, is the standard deviation; The number of groups was determined by the inflection points of the cluster dendrogram; the number of hierarchical clustering groups was verified by the silhouette coefficient method to ensure the consistency of plaque density characteristics within the group and maximize the heterogeneity between groups.

5. The method for predicting the stability of river water microbial communities using watershed land use patch density according to claim 1, wherein: In step 4, the Bray-Curtis distance matrix of the microbial communities in each group was calculated using the R language vegan package, and the average variation dissimilarity was calculated according to the following formula: ,in, is the average variation dissimilarity, is the number of samples within the group, For samples and samples The Bray-Curtis distance between them.

6. The method for predicting river water microbial community stability using watershed land use patch density according to claim 1, wherein: In step 4, the stats toolkit of R language was used to construct linear regression and quadratic polynomial regression models with the mean of patch density as the independent variable and the mean variation dissimilarity as the dependent variable. At the same time, the randomForest toolkit was used to construct a random forest regression model with the parameter ntree=500. After the model was constructed, the mean square error, mean absolute error, and determination coefficient were used as indicators to evaluate the model performance and select the optimal prediction model.

7. The method for predicting river water microbial community stability using watershed land use patch density according to claim 1, wherein: In step 4, before building the random forest regression model, the sample size needs to be judged: if the sample size is less than 30, to ensure the reliability of the model evaluation, the leave-one-out cross-validation method is used. The validation method is set to leave-one-out cross-validation through the trainControl function of the caret package, and the model is trained using the train function; if the sample size is greater than or equal to 30, the data is divided into a training set and a test set in a ratio of 7:

3. The model is trained using the training set and evaluated using the test set.

8. The method for predicting river water microbial community stability using watershed land use patch density according to claim 1, wherein: In step five, the spdep package was used to construct a k-nearest-neighbor spatial weight matrix with k=3-5. The model residuals were subjected to a Moran's I test with a significance threshold of p<0.05, and spatial autocorrelation was assessed using a Moran's scatter plot. Based on the optimal model and watershed patch density data, the stability of the microbial community in the river water of the unsampled area was predicted, and a spatial distribution map was generated using ArcGIS to guide the identification of ecologically vulnerable areas and land use planning and regulation.

9. The method for predicting river water microbial community stability using watershed land use patch density according to claim 1, wherein: In step 5, if the Moran index test shows significant spatial autocorrelation, p < 0.05, the spatial error model is used to correct the prediction results. The model is implemented by the errorsarlm function of the spdep package.

10. The method for predicting river water microbial community stability using watershed land use patch density according to claim 1, wherein: In step five, when the prediction results are compared and verified with the field sampling data, the prediction error is required to be ≤15%, otherwise the model is re-optimized or the sampling data is supplemented.

Citation Information

Patent Citations

  • Basin ecological health threshold evaluation method based on landscape fragmentation and microbial network modularity

    CN120015131A

  • Method of Determining River Basin Ecological Risk with Spatial Aggregation Level and Regional Imbalance Level

    US20250172535A1