A method for identifying interactive effects of land use and water quality on stream benthic animals
By combining the random forest and SHAP methods with the causal forest, the interactive effects of land use and water quality on river benthic animals are identified, which solves the problem that traditional methods cannot accurately identify nonlinear interactive effects, achieves accurate identification of interactive effects and causal verification, and improves the scientificity and practicality of watershed ecological assessment.
Patent Information
- Application Number
- CN202510977306.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Traditional methods have difficulty in accurately identifying the complex nonlinear interactive effects of land use and water quality on river benthic animals, resulting in quantification bias and the inability to clarify the causal directionality of the interactive effects.
The random forest and SHAP methods are combined with the causal forest technology. Data are obtained through field sampling and remote sensing interpretation, and a regression model is constructed. The SHAP method is used to quantify the interaction effect, and the causal directionality is verified through the causal forest to identify the interaction type between land use and water quality.
It has achieved accurate identification of the interactive effects of land use and water quality, broken through the limitations of traditional linear models, and improved the scientificity and practicality of watershed ecological impact assessment.
Smart Images

Figure CN120493090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental science and technology, and in particular to a method for identifying the interactive effects of land use and water quality on river benthic animals. Background Art
[0002] As a vital component of aquatic ecosystems, the community structure and diversity of river benthic fauna are key indicators of river health. Land use types (such as cultivated land, forest land, and construction land) indirectly influence water quality through changes in watershed hydrological processes and pollutant inputs. Changes in water quality, in turn, directly impact the survival and reproduction of benthic fauna. Therefore, clarifying the interactive mechanisms between land use and water quality on benthic fauna is crucial for watershed ecological protection and water quality management. However, traditional research methods (such as multivariate linear regression, principal component analysis, and correlation analysis) are mostly based on linear assumptions, treating the effects of land use and water quality on benthic fauna as independent or simply additive linear relationships. This approach fails to capture the complex nonlinear interactions between the two. In reality, the impact of land use on water quality can exhibit threshold effects (e.g., when intensive construction land development exceeds the carrying capacity of a watershed, water quality deteriorates significantly). Water quality can also exhibit nonlinear responses on benthic fauna (e.g., logarithmic or exponential relationships between pollutant concentrations and biodiversity). This leads to biases in the quantification of these interactions using traditional methods.
[0003] In recent years, the random forest-based SHAP (SHapley Additive exPlanations) method, as an interpretable machine learning technique, has provided new insights into nonlinear interactions in complex systems. By decomposing model predictions, this method can quantify the marginal contribution of each characteristic (such as land use percentage and water quality indicators) to the output variable (benthic biodiversity) and identify interactions between characteristics. It excels at handling high-dimensional data and complex nonlinear relationships, making it more suitable than traditional statistical methods for analyzing the complex mechanisms by which land use and water quality influence benthic animals. However, the SHAP method inherently reveals only statistical associations and cannot clearly identify the causal direction of interactions (e.g., whether land use changes are the direct cause of water quality changes that affect benthic animals). This can lead to misjudgment of the type of interaction (e.g., synergistic, antagonistic, or independent). Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the embodiment of the present invention aims to provide a method for identifying the interactive effects of land use and water quality on river benthic animals, so as to solve the problems in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for identifying the interactive effects of land use and water quality on stream benthic fauna includes the following steps:
[0007] Step S1: Collect benthic biomass data, water quality data, elevation data, and land use data within the study area through a combination of field sampling and remote sensing interpretation. Calculate biodiversity indicators based on benthic samples, calculate the area proportions of various land use types using sampling buffers as units, construct a dataset of input features and output variables, and form an analysis sample after standardization using a standard score (Z-score).
[0008] Step S2: Using the standardized land use ratio and water quality index as input features and the benthic biodiversity index as output variables, the random forest algorithm was used to construct a regression model. The training set and test set were divided into a ratio of 7:3, and the determination coefficient ( R ²) and the root mean square error ( RMSE ), select the model with the smallest error as the optimal relationship;
[0009] Step S3: Based on the optimal random forest model, the SHAP method is used to quantify the interactive effects of land use and water quality characteristics. The interactive SHAP value of each pair of characteristics at all sample points is calculated by decomposing the model prediction results. The average value of the interactive SHAP value of each pair of characteristics is calculated. If the average value is greater than 0, it is judged as a synergistic effect. If it is less than 0, it is an antagonistic effect. If it is close to 0, it is an independent effect. This completes the identification of the interactive effect type at the statistical level.
[0010] Step S4: Convert the continuous land use proportion and water quality index into binary treatment variables to generate land use-water quality interaction treatment variables. Apply the causal forest algorithm, using the original feature data as input, the binary treatment variables as intervention conditions, and the biodiversity index as output, to calculate the average treatment effect (ATE) of land use, water quality, and their interaction terms. By comparing the interaction term ATE with the sum of the single factor ATE, the interaction type at the causal level is determined: if the interaction term ATE is significantly greater than the sum of the two, it is a synergistic effect; if it is significantly less than the sum, it is an antagonistic effect; otherwise, it is an independent effect, thus verifying the causal directionality of the statistical interaction effect.
[0011] Step S5: Perform consistency check on the statistical interaction type obtained by the SHAP method and the causal interaction type verified by the causal forest. When the two judgment results are consistent, the interaction type is determined to be a valid interaction. If they are inconsistent, the interaction is determined to be an invalid interaction, and then the final conclusion is formed.
[0012] As a further solution of the present invention, step S1 specifically includes:
[0013] S11: Research area and data scope
[0014] Area selection: 115°35′-117°18′E, 26°29′-27°58′N, covering the main stream (length 346km) and 12 major tributaries, divided into 50 sub-basins (area 50-200 km²);
[0015] Time range: From January of the current year to December of the following year, benthic animal and water quality data were collected monthly, and water quality data were supplemented monthly, for a total of 24 time points;
[0016] S12: Data Collection Methods
[0017] Sampling point layout: One sampling point was set up in each sub-basin (a total of 50 points). A Sober net (mesh diameter 30 cm, pore size 0.5 mm) was used to randomly collect samples three times in the rapids area of the middle and upper reaches of the riverbed. The samples were combined and sorted through a 40-mesh sieve.
[0018] Laboratory analysis: The specimens were fixed with 75% ethanol, and the Shannon-Wiener index ( H '), , p i No. i The proportion of species individuals; Pielou evenness index ( J ), , S is the total number of species; Margalef richness index ( D ), , N is the total number of individuals;
[0019] Water quality data detection indicators: dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen and conductivity, three parallel samples were collected simultaneously at each sampling point;
[0020] Elevation data and land use data were obtained from the National Basic Geographic Information Center at a resolution of 30 m. The land use categories were: cultivated land, forest land, grassland, construction land, water area, and unused land. A 5 km circular buffer zone was set around the sampling point, and the land use percentage within the buffer zone was used to represent the land use situation at the sampling point.
[0021] S13: Data Preprocessing
[0022] Dataset construction: Taking sub-basin as unit, construct feature matrix: input features are 6 types of land use proportion and 6 water quality indicators, output features are 3 biodiversity indicators ( H '、 J 、 D ), a total of 1200 samples (50 sub-basins × 24 time points).
[0023] As a further solution of the present invention, the standardization process is to perform Z-score standardization on the land use ratio, water quality index and biodiversity index, and the standardization steps are as follows:
[0024] Calculate the mean of a data set: ,in, n is the number of samples in the dataset, x i For the i The characteristic value of each sample;
[0025] Calculate the standard deviation of a data set: ;
[0026] Standardized transformations: ,in, x i * is the standardized value.
[0027] As a further solution of the present invention, step S2 specifically includes:
[0028] S21: Model construction and parameter setting
[0029] For each standardized biodiversity indicator (H', J, D), a separate random forest model was constructed with 12 input features (6 standardized land use proportions + 6 standardized water quality indicators). The mathematical expression is: ,in, T is the number of decision trees, Indicates the t regression trees, For the t The parameters of the tree, x is the input feature vector, Represents the model's response to input features x The predicted value of
[0030] S22: Model training and evaluation
[0031] Data partitioning: training set and test set are divided into 7:3 ratio;
[0032] Evaluation index: Determination coefficient (R 2 ), ,in, y i is the true value, is the predicted value, is the true mean value, n is the sample size of the test set; the root mean square error ( RMSE ), ;
[0033] Model screening: For each biodiversity indicator, the random forest model is calculated on the test set. R 2 and RMSE , select the test set R 2 >0.8 and RMSE The smallest model is taken as the optimal relationship.
[0034] As a further solution of the present invention, step S3 specifically includes:
[0035] S31: Interactive SHAP value calculation
[0036] In the SHAP method, the concept of Shapley value is used to quantify the interaction between land use and water quality characteristics (a total of 36 pairs). x , land use characteristics i and water quality characteristics j Interaction SHAP value of It can be calculated as follows;
[0037]
[0038] in, F is the set of all features; n is the total number of all features; S Does not contain features i and j Any subset of features; x S Representing feature subsets S Take the actual value of the sample and the baseline value of the other features x 0 (training set mean); | S | for subset S The number of features in f (⋅) is the optimal relationship selected in step S2;
[0039] S32: Interaction type determination
[0040] Statistics calculation: For these 36 pairs of land use-water quality characteristics, calculate the interactive SHAP value of each pair of characteristics at all sample points , and find its average :
[0041] ,in, m For the current sample, M is the total number of samples;
[0042] based on , determine the type of interaction impact, the criteria are, synergistic effect: >0; antagonistic effect: <0; Independent effect: neither of the above two applies.
[0043] As a further solution of the present invention, step S4 specifically includes:
[0044] S41: Processing variable construction
[0045] In causal analysis, continuous land use ratios and water quality indicators are converted into binary variables. Binary variables can clearly indicate whether a factor is in a specific "high impact" state, facilitating the construction and analysis of subsequent causal models.
[0046] Land use treatment variables T LU,i,l : For each land use type l ,set up x LU,i,l For the i In the sample l The proportion of land use types, Median l For all samples l The median of the proportion of land use types, the land use treatment variable Defined as: ;
[0047] Water quality treatment variables T WQ,i,q :For each water quality indicator q , x WQ,i,q For the i In the sample q The concentration of water quality indicators, Standard q It is the third water quality standard in Class III. q The threshold value of the indicator, then the water quality treatment variable T WQ,i,q Defined as: ;
[0048] Interaction Processing Variables T LU×WQ,i,l,q : For each land use type l and each water quality indicator q Combination of interactive processing variables T LU×WQ,i,l,q Defined as land use treatment variable T LU,i,l and water quality treatment variables T WQ,i,q The product of , that is: ;
[0049] S42: Causal forest modeling and interaction type determination
[0050] Model input features: including original land use percentage data x LU,i,l and water quality index values x WQ,i,q ,These continuous data retain rich information, which helps the model to more accurately capture the relationship between different factors and biodiversity indicators;
[0051] Treatment variables: including land use treatment variables T LU,i,l , water quality treatment variables T WQ,i,q and interaction variables T LU×WQ,i,l,q , they specify the treatment status of different factors and are used to distinguish different experimental conditions;
[0052] Output: Biodiversity indicators in the optimal relationship Y i ;
[0053] Treatment effect index calculation: average treatment effect ( ATE ) measures the average effect of the treatment variable on the outcome variable. For land use treatment variables T LU , water quality treatment variables T WQ and interaction variables T LU×WQ , the average treatment effect is calculated as follows:
[0054] Land use type l of ATE ( ): ,in, n is the sample size, Indicates the i The biodiversity index value of each sample when the land use treatment variable is 1. Indicates the value when it is 0;
[0055] Water quality indicators q of ATE ( ): ,in, Indicates the i The biodiversity index value of each sample when the water quality treatment variable is 1, Indicates the value when it is 0;
[0056] Land use land water quality indicators q The interaction ATE ( ): in, Indicates the i The biodiversity index value of a sample when the current interaction variable is 1, Indicates the value when it is 0;
[0057] The criteria for determining the interaction type are: synergy effect, > ATE LU,l +ATE WQ,q ; antagonistic effect, < ATE LU,l +ATE WQ,q ; Independent effect, neither of the above two belongs to it.
[0058] As a further solution of the present invention, step S5 specifically includes:
[0059] S51: Interaction type consistency determination
[0060] When the interaction type of the SHAP method is consistent with that of the causal forest, the interaction type is determined to be the interactive impact type of the land use and water quality characteristics and is identified as a valid interaction; if not, it is identified as an invalid interaction.
[0061] In summary, the embodiments of the present invention have the following beneficial effects compared with the prior art:
[0062] This invention breaks through the limitations of traditional linear models by integrating the random forest-SHAP method with the causal forest. It can not only analyze complex nonlinear statistical interaction effects, but also verify the type of interaction influence from the causal relationship level, filling the technical gap in the "statistical association-causal inference" dual analysis, and provides an effective way to accurately identify the interaction between land use and water quality on benthic animals, significantly improving the scientificity and practicality of watershed ecological impact assessment.
[0063] In order to more clearly illustrate the structural features and effects of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Flowchart of an embodiment of the invention.
[0065] Figure 2 Flowchart for data collection constructed in an embodiment of the invention.
[0066] Figure 3 This is a flow chart of random forest model fitting and optimization in an embodiment of the invention.
[0067] Figure 4 This is a flow chart of the SHAP method for quantifying statistical interaction effects in an embodiment of the invention.
[0068] Figure 5 This is a flow chart of verifying causal interaction effects in a causal forest according to an embodiment of the invention.
[0069] Figure 6 This is a flow chart of verifying causal interaction effects in a causal forest according to an embodiment of the invention. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0071] The specific implementation of the present invention is described in detail below with reference to specific embodiments.
[0072] In one embodiment, a method for identifying the interactive effects of land use and water quality on stream benthic fauna is described in Figures 1 to 6 , including the following steps:
[0073] Step S1: Collect benthic biomass data, water quality data, elevation data, and land use data within the study area through a combination of field sampling and remote sensing interpretation. Calculate biodiversity indicators based on benthic samples, calculate the area proportions of various land use types using sampling buffers as units, construct a dataset of input features and output variables, and form an analysis sample after standardization using a standard score (Z-score).
[0074] Step S2: Using the standardized land use ratio and water quality index as input features and the benthic biodiversity index as output variables, the random forest algorithm was used to construct a regression model. The training set and test set were divided into a ratio of 7:3, and the determination coefficient ( R ²) and the root mean square error ( RMSE ), select the model with the smallest error as the optimal relationship;
[0075] Step S3: Based on the optimal random forest model, the SHAP method is used to quantify the interactive effects of land use and water quality characteristics. The interactive SHAP value of each pair of characteristics at all sample points is calculated by decomposing the model prediction results. The average value of the interactive SHAP value of each pair of characteristics is calculated. If the average value is greater than 0, it is judged as a synergistic effect. If it is less than 0, it is an antagonistic effect. If it is close to 0, it is an independent effect. This completes the identification of the interactive effect type at the statistical level.
[0076] Step S4: Convert the continuous land use proportion and water quality index into binary treatment variables to generate land use-water quality interaction treatment variables. Apply the causal forest algorithm, using the original feature data as input, the binary treatment variables as intervention conditions, and the biodiversity index as output, to calculate the average treatment effect (ATE) of land use, water quality, and their interaction terms. By comparing the interaction term ATE with the sum of the single factor ATE, the interaction type at the causal level is determined: if the interaction term ATE is significantly greater than the sum of the two, it is a synergistic effect; if it is significantly less than the sum, it is an antagonistic effect; otherwise, it is an independent effect, thus verifying the causal directionality of the statistical interaction effect.
[0077] Step S5: Perform consistency check on the statistical interaction type obtained by the SHAP method and the causal interaction type verified by the causal forest. When the two judgment results are consistent, the interaction type is determined to be a valid interaction. If they are inconsistent, the interaction is determined to be an invalid interaction, and then the final conclusion is formed.
[0078] For further information, see Figures 1 to 6 , the step S1 specifically includes:
[0079] S11: Research area and data scope
[0080] Area selection: 115°35′-117°18′E, 26°29′-27°58′N, covering the main stream (length 346km) and 12 major tributaries, divided into 50 sub-basins (area 50-200 km²);
[0081] Time range: From January of the current year to December of the following year, benthic animal and water quality data were collected monthly, and water quality data were supplemented monthly, for a total of 24 time points;
[0082] S12: Data Collection Methods
[0083] Sampling point layout: One sampling point was set up in each sub-basin (a total of 50 points). A Sober net (mesh diameter 30 cm, pore size 0.5 mm) was used to randomly collect samples three times in the rapids area of the middle and upper reaches of the riverbed. The samples were combined and sorted through a 40-mesh sieve.
[0084] Laboratory analysis: The specimens were fixed with 75% ethanol, and the Shannon-Wiener index ( H '), , p i No. i The proportion of species individuals; Pielou evenness index ( J ), , S is the total number of species; Margalef richness index ( D), , N is the total number of individuals;
[0085] Water quality data detection indicators: dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen and conductivity, three parallel samples were collected simultaneously at each sampling point;
[0086] Elevation data and land use data were obtained from the National Basic Geographic Information Center at a resolution of 30 m. The land use categories were: cultivated land, forest land, grassland, construction land, water area, and unused land. A 5 km circular buffer zone was set around the sampling point, and the land use percentage within the buffer zone was used to represent the land use situation at the sampling point.
[0087] S13: Data Preprocessing
[0088] Dataset construction: Taking sub-basin as unit, construct feature matrix: input features are 6 types of land use proportion and 6 water quality indicators, output features are 3 biodiversity indicators ( H '、 J 、 D ), a total of 1200 samples (50 sub-basins × 24 time points).
[0089] For further information, see Figures 1 to 6 The standardization process is to perform Z-score standardization on land use ratio, water quality index and biodiversity index. The standardization steps are as follows:
[0090] Calculate the mean of a data set: ,in, n is the number of samples in the dataset, x i For the i The characteristic value of each sample;
[0091] Calculate the standard deviation of a data set: ;
[0092] Standardized transformations: ,in, x i * is the standardized value.
[0093] For further information, see Figures 1 to 6 , the step S2 specifically includes:
[0094] S21: Model construction and parameter setting
[0095] For each standardized biodiversity indicator (H', J, D), a separate random forest model was constructed with 12 input features (6 standardized land use proportions + 6 standardized water quality indicators). The mathematical expression is: ,in, T is the number of decision trees, Indicates the t regression trees, For the t The parameters of the tree, x is the input feature vector, Represents the model's response to input features x The predicted value of
[0096] S22: Model training and evaluation
[0097] Data partitioning: training set and test set are divided into 7:3 ratio;
[0098] Evaluation index: Determination coefficient (R 2 ), ,in, y i is the true value, is the predicted value, is the true mean value, n is the sample size of the test set; the root mean square error ( RMSE ), ;
[0099] Model screening: For each biodiversity indicator, the random forest model is calculated on the test set. R 2 and RMSE , select the test set R 2 >0.8 and RMSE The smallest model is taken as the optimal relationship.
[0100] For further information, see Figures 1 to 6 , the step S3 specifically includes:
[0101] S31: Interactive SHAP value calculation
[0102] In the SHAP method, the concept of Shapley value is used to quantify the interaction between land use and water quality characteristics (a total of 36 pairs). x , land use characteristics i and water quality characteristics j Interaction SHAP value of It can be calculated as follows;
[0103]
[0104] in, F is the set of all features; n is the total number of all features; S Does not contain features i andj Any subset of features; x S Representing feature subsets S Take the actual value of the sample and the baseline value of the other features x 0 (training set mean); | S | for subset S The number of features in f (⋅) is the optimal relationship selected in step S2;
[0105] S32: Interaction type determination
[0106] Statistics calculation: For these 36 pairs of land use-water quality characteristics, calculate the interactive SHAP value of each pair of characteristics at all sample points , and find its average :
[0107] ,in, m For the current sample, M is the total number of samples;
[0108] based on , determine the type of interaction impact, the criteria are, synergistic effect: >0; antagonistic effect: <0; Independent effect: neither of the above two applies.
[0109] For further information, see Figures 1 to 6 , the step S4 specifically includes:
[0110] S41: Processing variable construction
[0111] In causal analysis, continuous land use ratios and water quality indicators are converted into binary variables. Binary variables can clearly indicate whether a factor is in a specific "high impact" state, facilitating the construction and analysis of subsequent causal models.
[0112] Land use treatment variables T LU,i,l : For each land use type l ,set up x LU,i,l For the i In the sample l The proportion of land use types, Median l For all samples l The median of the proportion of land use types, the land use treatment variable Defined as: ;
[0113] Water quality treatment variables T WQ,i,q :For each water quality indicator q , x WQ,i,q For the i In the sample q The concentration of water quality indicators, Standard q It is the third water quality standard in Class III. q The threshold value of the indicator, then the water quality treatment variable T WQ,i,q Defined as: ;
[0114] Interaction Processing Variables T LU×WQ,i,l,q : For each land use type l and each water quality indicator q Combination of interactive processing variables T LU×WQ,i,l,q Defined as land use treatment variable T LU,i,l and water quality treatment variables T WQ,i,q The product of , that is: ;
[0115] S42: Causal forest modeling and interaction type determination
[0116] Model input features: including original land use percentage data x LU,i,l and water quality index values x WQ,i,q ,These continuous data retain rich information, which helps the model to more accurately capture the relationship between different factors and biodiversity indicators;
[0117] Treatment variables: including land use treatment variables T LU,i,l , water quality treatment variables T WQ,i,q and interaction variables T LU×WQ,i,l,q , they specify the treatment status of different factors and are used to distinguish different experimental conditions;
[0118] Output: Biodiversity indicators in the optimal relationship Y i ;
[0119] Treatment effect index calculation: average treatment effect ( ATE ) measures the average effect of the treatment variable on the outcome variable. For land use treatment variables T LU , water quality treatment variablesT WQ and interaction variables T LU×WQ , the average treatment effect is calculated as follows:
[0120] Land use type l of ATE ( ): ,in, n is the sample size, Indicates the i The biodiversity index value of each sample when the land use treatment variable is 1. Indicates the value when it is 0;
[0121] Water quality indicators q of ATE ( ): ,in, Indicates the i The biodiversity index value of each sample when the water quality treatment variable is 1, Indicates the value when it is 0;
[0122] Land use l and water quality indicators q The interaction ATE ( ): in, Indicates the i The biodiversity index value of a sample when the current interaction variable is 1, Indicates the value when it is 0;
[0123] The criteria for determining the interaction type are: synergy effect, > ATE LU,l +ATE WQ,q ; antagonistic effect, < ATE LU,l +ATE WQ,q ; Independent effect, neither of the above two belongs to it.
[0124] For further information, see Figures 1 to 6 , the step S5 specifically includes:
[0125] S51: Interaction type consistency determination
[0126] When the interaction type of the SHAP method is consistent with that of the causal forest, the interaction type is determined to be the interactive impact type of the land use and water quality characteristics and is identified as a valid interaction; if not, it is identified as an invalid interaction.
[0127] In this embodiment, S11: Research area and data range
[0128] Regional selection: Fuhe River Basin in Jiangxi Province (115°35′-117°18′E, 26°29′-27°58′N), covering the main stream (length 346 km) and 12 major tributaries, divided into 50 sub-basins (area 50-200 km²).
[0129] Time range: From January of the current year to December of the following year, benthic animal and water quality data are collected on a monthly basis, and water quality data are supplemented on a monthly basis, for a total of 24 time nodes.
[0130] S12: Data Collection Methods
[0131] Sampling point layout: One sampling point was set up in each sub-basin (a total of 50 points). A Sober net (mesh diameter 30 cm, pore size 0.5 mm) was used to randomly sample the rapids area in the middle and upper reaches of the riverbed three times, and the samples were combined and sorted through a 40-mesh sieve.
[0132] Laboratory analysis: Specimens were fixed with 75% ethanol and identified to family level (referring to the Chinese Freshwater Benthic Animals). The number of individuals and biomass (wet weight, g / m²) were counted, and biodiversity indicators were calculated. Shannon-Wiener index ( H '), , p i No. i The proportion of species individuals; Pielou evenness index ( J ), , S is the total number of species; Margalef richness index ( D ), , N is the total number of individuals.
[0133] Water quality data detection indicators: dissolved oxygen (DO, electrochemical method), chemical oxygen demand (COD, potassium dichromate method), ammonia nitrogen (NH3-N, Nessler's reagent method), total phosphorus (TP, molybdenum antimony anti-spectrophotometry method), total nitrogen (TN, alkaline potassium persulfate digestion ultraviolet spectrophotometry method), electrical conductivity (EC, electrode method), 3 parallel samples were collected simultaneously at each sampling point.
[0134] Elevation and land use data were obtained from the National Geographic Information Center at a 30 m resolution. Land use categories include cultivated land, forest land, grassland, construction land, water area, and unused land. A 5 km circular buffer zone was established around the sampling point, and the land use percentage within the buffer zone was used to represent the land use situation at the sampling point.
[0135] S13: Data Preprocessing
[0136] Dataset construction: Taking sub-basin as unit, construct feature matrix: input features are 6 types of land use proportion and 6 water quality indicators, output features are 3 biodiversity indicators ( H '、 J 、 D ), a total of 1200 samples (50 sub-basins × 24 time points).
[0137] Standardization: Z-score standardization is performed on land use ratio, water quality index and biodiversity index. The standardization steps are as follows:
[0138] 1. Calculate the mean of the data set: ,in, n is the number of samples in the dataset, x i For the i The characteristic values of the samples.
[0139] 2. Calculate the standard deviation of the data set: .
[0140] 3. Standardized conversion: ,in, x i * is the standardized value.
[0141] Step S2: Random Forest Model Fitting and Optimization
[0142] S21: Model construction and parameter setting
[0143] For each standardized biodiversity indicator (H', J, D), a separate random forest model was constructed with 12 input features (6 standardized land use proportions + 6 standardized water quality indicators). The mathematical expression is: ,in, T is the number of decision trees, Indicates the t regression trees, For the t The parameters of the tree, x is the input feature vector, Represents the model's response to input features x The predicted value of .
[0144] S22: Model training and evaluation
[0145] Data partitioning: The training set and test set are divided into 7:3 ratios.
[0146] Evaluation index: Determination coefficient (R2 ), ,in, y i is the true value, is the predicted value, is the true mean value, n is the sample size of the test set; the root mean square error ( RMSE ), .
[0147] Model screening: For each biodiversity indicator, the random forest model is calculated on the test set. R 2 and RMSE , select the test set R 2 >0.8 and RMSE The smallest model is taken as the optimal relationship.
[0148] Step S3: SHAP method to quantify statistical interaction effects
[0149] S31: Interactive SHAP value calculation
[0150] In the SHAP method, the concept of Shapley value is used to quantify the interaction between land use and water quality characteristics (36 pairs in total). x , land use characteristics i and water quality characteristics j Interaction SHAP value of It can be calculated as follows.
[0151]
[0152] in, F is the set of all features; n is the total number of all features; S Does not contain features i and j Any subset of features; x S Representing feature subsets S Take the actual value of the sample and the baseline value of the other features x 0 (training set mean); | S | for subset S The number of features in f (⋅) is the optimal relationship selected in step S2.
[0153] S32: Interaction type determination
[0154] Statistics calculation: For these 36 pairs of land use-water quality characteristics, calculate the interactive SHAP value of each pair of characteristics at all sample points , and find its average :
[0155] ,in, m For the current sample, M is the total number of samples.
[0156] based on , determine the type of interaction. The criteria for determination are, synergistic effect: >0; antagonistic effect: <0; Independent effect: neither of the above two applies.
[0157] Step S4: Causal forest verification of causal interaction effects
[0158] S41: Processing variable construction
[0159] In causal analysis, converting continuous land use and water quality indicators into binary variables simplifies complex data relationships and more clearly explores the causal impacts of different factors on biodiversity indicators. Binary variables clearly indicate whether a factor is in a specific "high impact" state, facilitating the construction and analysis of subsequent causal models.
[0160] Land use treatment variables T LU,i,l : For each land use type l ,set up x LU,i,l For the i In the sample l The proportion of land use types, Median l For all samples l The median of the proportion of land use types. The land use treatment variable Defined as:
[0161] Water quality treatment variables T WQ,i,q :For each water quality indicator q , x WQ,i,q For the i In the sample q The concentration of water quality indicators, Standard q It is the third water quality standard in Class III. q The threshold value of the indicator. Then the water quality treatment variable T WQ,i,q Defined as:
[0162] Interaction Processing Variables TLU×WQ,i,l,q : For each land use type l and each water quality indicator q Combination of interactive processing variables T LU×WQ,i,l,q Defined as land use treatment variable T LU,i,l and water quality treatment variables T WQ,i,q The product of , that is: .
[0163] S42: Causal forest modeling and interaction type determination
[0164] Model input features: including original land use percentage data x LU,i,l and water quality index values x WQ,i,q These continuous data retain rich information and help the model more accurately capture the relationship between different factors and biodiversity indicators.
[0165] Treatment variables: including land use treatment variables T LU,i,l , water quality treatment variables T WQ,i,q and interaction variables T LU×WQ,i,l,q , which clarify the treatment status of different factors and are used to distinguish different experimental conditions.
[0166] Output: Biodiversity indicators in the optimal relationship Y i .
[0167] Treatment effect index calculation: average treatment effect ( ATE ) measures the average effect of the treatment variable on the outcome variable. T LU , water quality treatment variables T WQ and interaction variables T LU×WQ The formula for calculating the average treatment effect is as follows.
[0168] Land use type l of ATE ( ): ,in, n is the sample size, Indicates the i The biodiversity index value of each sample when the land use treatment variable is 1. Indicates the value when it is 0.
[0169] Water quality indicators q of ATE ( ): ,in, Indicates the i The biodiversity index value of each sample when the water quality treatment variable is 1, Indicates the value when it is 0.
[0170] Land use l and water quality indicators q The interaction ATE ( ): in, Indicates the i The biodiversity index value of a sample when the current interaction variable is 1, Indicates the value when it is 0.
[0171] The criteria for determining the interaction type are: synergy effect, > ATE LU,l +ATE WQ,q ; antagonistic effect, < ATE LU,l +ATE WQ,q ; Independent effect, neither of the above two belongs to it.
[0172] Step S5: Interaction type consistency check
[0173] S51: Interaction type consistency determination
[0174] When the interaction type of the SHAP method is consistent with that of the causal forest, the interaction type is determined to be the interactive impact type of the land use and water quality characteristics and is identified as a valid interaction; if not, it is identified as an invalid interaction.
[0175] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying the interactive effects of land use and water quality on river benthic animals, characterized in that: The following steps are involved: Step S1: Collect benthic biomass data, water quality data, elevation data, and land use data within the study area through a combination of field sampling and remote sensing interpretation. Calculate biodiversity indicators based on benthic samples, calculate the area proportions of various land use types using sampling buffers as units, construct a dataset of input features and output variables, and form an analysis sample after standardization using standard scores. Step S2: Using the standardized land use ratio and water quality index as input features and the benthic biodiversity index as the output variable, the random forest algorithm was used to construct a regression model. The training set and test set were divided into a 7:3 ratio. The coefficient of determination and root mean square error were calculated, and the model with the smallest error was selected as the optimal relationship. Step S3: Based on the optimal random forest model, the SHAP method is used to quantify the interactive effects of land use and water quality characteristics. The interactive SHAP value of each pair of characteristics at all sample points is calculated by decomposing the model prediction results. The average value of the interactive SHAP value of each pair of characteristics is calculated. If the average value is greater than 0, it is judged as a synergistic effect. If it is less than 0, it is an antagonistic effect. If it is close to 0, it is an independent effect. This completes the identification of the interactive effect type at the statistical level. Step S4: Convert the continuous land use ratio and water quality index into binary treatment variables to generate land use-water quality interaction treatment variables. Apply the causal forest algorithm, with the original feature data as input, the binary treatment variables as intervention conditions, and the biodiversity index as output, to calculate the average treatment effect (ATE) of land use, water quality, and their interaction terms. By comparing the interaction term ATE with the sum of the single factor ATE, the interaction type at the causal level is determined: if the interaction term ATE is significantly greater than the sum of the two, it is a synergistic effect; if it is significantly less than the sum, it is an antagonistic effect; otherwise, it is an independent effect, thereby verifying the causal directionality of the statistical interaction effect. Step S5: Perform consistency check on the statistical interaction type obtained by the SHAP method and the causal interaction type verified by the causal forest. When the judgment results of the two are consistent, the interaction type is determined to be a valid interaction. If they are inconsistent, the interaction is determined to be an invalid interaction, and then the final conclusion is formed.
2. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 1, characterized in that: The step S1 specifically includes: S11: Research area and data scope Area selection: 115°35′-117°18′ east longitude, 26°29′-27°58′ north latitude, covering the main stream and 12 major tributaries, divided into 50 sub-basins; Time range: From January of the current year to December of the following year, benthic animal and water quality data were collected monthly, and water quality data were supplemented monthly, for a total of 24 time points; S12: Data Collection Methods Sampling point layout: One sampling point was set up in each sub-basin. Three random samples were taken from the rapids in the middle and upper reaches of the riverbed using a Sober net. The samples were combined and sorted using a 40-mesh sieve. Laboratory analysis: The specimens were fixed with 75% ethanol, and the Shannon-Wiener index was H ', , p i No. i The proportion of species individuals; Pielou evenness index J , , S is the total number of species; Margalef richness index D , , N is the total number of individuals; Water quality data detection indicators: dissolved oxygen, chemical oxygen demand, ammonia nitrogen, total phosphorus, total nitrogen and conductivity, three parallel samples were collected simultaneously at each sampling point; Elevation data and land use data were obtained at a resolution of 30 m. The land use categories were: cultivated land, forest land, grassland, construction land, water area, and unused land. A 5 km circular buffer zone was set around the sampling point, and the land use percentage within the buffer zone was used to represent the land use situation at the sampling point. S13: Data Preprocessing Dataset construction: Taking sub-basin as unit, construct feature matrix: input features are 6 types of land use proportion and 6 water quality indicators, output features are 3 biodiversity indicators ( H '、 J 、 D ), a total of 1200 samples.
3. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 2, characterized in that: The standardization process is to perform Z-score standardization on land use ratio, water quality index and biodiversity index. The standardization steps are as follows: Calculate the mean of a data set: ,in, n is the number of samples in the dataset, x i For the i The characteristic value of each sample; Calculate the standard deviation of a data set: ; Standardized transformations: ,in, x i * is the standardized value.
4. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 3, characterized in that: The step S2 specifically includes: S21: Model construction and parameter setting For each standardized biodiversity indicator (H', J, D), an independent random forest model is constructed with a 12-dimensional input feature. The mathematical expression is: ,in, T is the number of decision trees, Indicates the t Regression trees, For the t The parameters of the tree, x is the input feature vector, Represents the model's response to input features x The predicted value of S22: Model training and evaluation Data division: the training set and the test set are divided into 7:3 ratios; Evaluation indicator: Coefficient of determination R 2 , ,in, y i is the true value, is the predicted value, is the true mean value, n is the sample size of the test set; the root mean square error RMSE , ; Model screening: For each biodiversity indicator, the random forest model is calculated on the test set. R 2 and RMSE , select the test set R 2 >0.8 and RMSE The smallest model is taken as the optimal relationship.
5. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 4, characterized in that: The step S3 specifically includes: S31: Interactive SHAP value calculation In the SHAP method, the interaction between land use and water quality characteristics is quantified by Shapley value. x , land use characteristics i and water quality characteristics j Interaction SHAP value of It can be calculated as follows; ; in, F is the set of all features; n is the total number of all features; S Does not contain features i and j Any subset of features; x S Representing feature subsets S Take the actual value of the sample and the baseline value of the remaining features x 0 , the baseline value is the mean of the training set; | S | for subset S The number of features in f (⋅) is the optimal relationship selected in step S2; S32: Interaction type determination Statistics calculation: For land use-water quality characteristics, calculate the interactive SHAP value of each pair of characteristics at all sample points , and find its average : ,in, m For the current sample, M is the total number of samples; based on , determine the type of interaction impact, the criteria are, synergistic effect: >0; antagonistic effect: <0; Independent effect: neither of the above two applies.
6. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 5, characterized in that: The step S4 specifically includes: S41: Processing variable construction In causal analysis, continuous land use ratios and water quality indicators are converted into binary variables. Binary variables can clearly indicate whether a factor is in a specific high-impact state, facilitating the construction and analysis of subsequent causal models. Land use treatment variables T LU,i,l : For each land use type l ,set up x LU,i,l For the i In the sample l The proportion of land use types, Median l For all samples l The median of the proportion of land use types, the land use treatment variable Defined as: ; Water quality treatment variables T WQ,i,q :For each water quality indicator q , x WQ,i,q For the i In the sample q The concentration of water quality indicators, Standard q It is the third water quality standard in Class III. q The threshold value of the indicator, then the water quality treatment variable T WQ,i,q Defined as: ; Interaction Processing Variables T LU×WQ,i,l,q : For each land use type l and each water quality indicator q Combination of interactive processing variables T LU×WQ,i,l,q Defined as land use treatment variable T LU,i,l and water quality treatment variables T WQ,i,q The product of , that is: ; S42: Causal forest modeling and interaction type determination Model input features: including original land use percentage data x LU,i,l and water quality index values x WQ,i,q ; Treatment variables: including land use treatment variables T LU,i,l , water quality treatment variables T WQ,i,q and interaction variables T LU×WQ,i,l,q , used to distinguish different experimental conditions; Output: Biodiversity indicators in the optimal relationship Y i ; Treatment effect index calculation: For land use treatment variables T LU , water quality treatment variables T WQ and interaction variables T LU×WQ , the average treatment effect is calculated as follows: Land use type l of ATE ( ): ,in, n is the sample size, Indicates the i The biodiversity index value of each sample when the land use treatment variable is 1. Indicates the value when it is 0; Water quality indicators q of ATE ( ): ,in, Indicates the i The biodiversity index value of each sample when the water quality treatment variable is 1, Indicates the value when it is 0; Land use l and water quality indicators q Interaction : in, Indicates the i The biodiversity index value of a sample when the current interaction variable is 1, Indicates the value when it is 0; The criteria for determining the interaction type are: synergy effect, > ATE LU,l +ATE WQ,q ; antagonistic effect, < ATE LU,l +ATE WQ,q ; Independent effect, neither of the above two belongs to it.
7. The method for identifying the interactive effects of land use and water quality on river benthic animals according to claim 6, characterized in that: The step S5 specifically includes: S51: Interaction type consistency determination When the interaction type of the SHAP method is consistent with that of the causal forest, the interaction type is determined to be the interactive impact type of the land use and water quality characteristics and is identified as a valid interaction; if not, it is identified as an invalid interaction.
Citation Information
Patent Citations
Soil heavy metal pollution driving factor identification method based on Catboost-SHAP model
CN117290727A
Interpretable forest above-ground biomass remote sensing estimation method and device
CN117611997A