A method for predicting river water ecological integrity with multiple inputs and multiple outputs
By combining hydrodynamic water quality models and random forest models, a multi-input multi-output river water ecological integrity prediction method is constructed, which solves the problem of establishing multi-input multi-output response relationships, realizes the simultaneous prediction of multiple water ecological integrity indicators, and improves prediction accuracy and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINESE RES ACAD OF ENVIRONMENTAL SCI
- Filing Date
- 2026-05-19
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies cannot effectively establish multi-input and multi-output response relationships, cannot simultaneously predict multiple aquatic ecological integrity indicators, and cannot meet the needs of multiple biological protection objectives.
By combining hydrodynamic and water quality models and random forest models, and by constructing a multivariate joint splitting criterion and a cross-index residual feature enhancement mechanism, we can achieve multi-input and multi-output prediction of river water ecological integrity. We use the physical, chemical, and biological indicators output by the hydrodynamic and water quality models as inputs to the random forest models to predict water ecological integrity indicators.
It achieves simultaneous output of multiple water ecological integrity indicators, solving the problem that traditional methods can only output a single target, and improving the accuracy and reliability of water ecological prediction.
Smart Images

Figure CN122452437A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water ecology, specifically relating to a multi-input, multi-output method for predicting the integrity of river water ecology. Background Technology
[0002] Chinese patent application CN119090093A discloses a method, device, and storage medium for generating a water ecology prediction model that couples a mechanistic model and a machine learning model. This method belongs to the technical field of prediction data processing methods and solves the problem of insufficient spatial resolution in water ecology prediction. The water ecology prediction model generation method includes: acquiring sampled data; calculating aquatic biological indicators of water samples; training the model to obtain a selected basic water ecology prediction model; using current hydrological data as input and calibrating the distributed hydrological model of the watershed using hydrodynamic conditions measured data from hydrological stations; using current water pollution data and watershed hydrodynamic conditions-driven water quality models as input and calibrating the watershed water quality model using monitoring data of water environment quality indicators; and generating a water ecology prediction model based on the above models. This invention improves the spatiotemporal resolution and accuracy of water ecology prediction and can predict the impact of governance measures and water conservancy projects on water ecology. The above technical solution emphasizes using the results of hydrological, hydrodynamic, and soil utilization mechanism models as input boundaries and aquatic biological indicators as output to establish a multi-input, single-output machine learning model to obtain water ecology prediction data. However, the above technical solutions have the following shortcomings: the water ecological protection target is not just a single biological indicator, but includes multiple biological indicators and is oriented towards multiple output indicators. The above technical methods cannot solve the problem of establishing response relationships between multiple drivers and multiple protection targets in the field of water ecology.
[0003] River water ecological integrity encompasses multiple biological indicators and is also affected by various driving factors. Establishing the response relationship between multiple input factors and multiple output indicators to predict river water ecological integrity is of great significance for the protection and restoration of river water ecological integrity. Summary of the Invention
[0004] Technical Challenge: The challenge lies in establishing multi-input and multi-output response relationships during aquatic ecosystem integrity prediction. Specifically, how can multiple inputs be used to simultaneously simulate and predict multiple targets, addressing the many-to-many technical problem? Many current prediction models rely on mechanistic models as inputs and single biological indicators as outputs to establish mechanistic models and solve the many-to-single problem. However, aquatic ecosystem integrity is a comprehensive indicator encompassing multiple biological protection targets. Therefore, how to utilize multiple physicochemical indicators output from physical models to predict multiple biological indicators is a pressing technical challenge.
[0005] Purpose of the invention: The present invention addresses the problems existing in the prior art by disclosing a multi-input, multi-output method for predicting the integrity of river water ecosystems.
[0006] The method coupled in this invention includes the construction of a hydrodynamic water quality model and a random forest (RF) model. The hydrodynamic water quality model is mainly used to simulate and output the physical, chemical and biological indicators of the water body, and uses them as input indicators for the random forest model. The output indicators of the random forest model are water ecological integrity indicators, thereby achieving the prediction of water ecological integrity protection targets.
[0007] Technical solution: A multi-input, multi-output method for predicting the integrity of river water ecology, the steps of which are as follows: (1) Construction of hydrodynamic water quality model; (2) Output of hydrodynamic water quality model Based on the evaluation points of the water ecological integrity index, the corresponding grid data is extracted from the hydrodynamic and water quality model of the target river obtained in step (1). The grid data includes hydrological and hydrodynamic index data, chemical index data and biological index data. (3) Construction, validation and evaluation of the random forest model (31) Data Allocation: The hydrological and hydrodynamic index data, chemical index data, and biological index data output from the hydrodynamic and water quality model in step (2) are used as input boundaries, and multiple indicators of aquatic ecological integrity are used as output indicators. Training and validation sets are allocated, and a random forest model is constructed, wherein: The random forest model is implemented using R language. The workflow of the random forest model includes calling the tool package in R language, importing the dataset, splitting the dataset, building the model and evaluating its performance. (32) Validation and evaluation of the random forest model: Based on the data allocation results, the simulated data and measured data of multiple water ecological integrity indicators output by the random forest model are compared to calibrate and verify the accuracy and effectiveness of the model.
[0008] Furthermore, the steps of step (1) are as follows: (11) Collect data on the target river. The data to be collected for the target river includes river boundary data, river topography data, flow data, water level data, water quality concentration data, and water temperature data, among which: The river boundary data was drawn from the Ovi map and is used to define the boundaries for constructing the hydrodynamic and water quality model. The river topographic data is obtained by secondary processing based on DEM data extraction. The flow data was obtained based on hydrological yearbooks; The water level data was obtained based on hydrological yearbooks; The water quality concentration data was obtained from the National Surface Water Quality Automatic Monitoring Real-time Data Release System. The water temperature data was obtained based on a hydrological yearbook. (12) Grid generation Obtain a high-resolution satellite image of the target river and import it into ArcGIS software. Then, using the built-in editing tools of ArcGIS, outline the river boundary of the target river according to the range of the target river in the high-resolution satellite image. Import the depicted river boundary into globalmapper and convert its format to one that Delft3D can recognize. Import the converted river boundary file into Delft3D, divide the river into meshes in Delft3D, and export it as a .grid format mesh data file required for model building. (13) Terrain processing Import the mesh data file generated in step (12) into the EFDC model and obtain the x and y coordinates of the four corners of the mesh; By taking the x-coordinates and y-coordinates of the four corners of the grid, and calculating the average of the x-coordinates of the four corners and the average of the four x-coordinates, we can obtain the x-coordinates and y-coordinates of the grid center point. Based on the upstream and downstream order of the grid, and using a small amount of measured cross-sectional data and the average topographic slope, the elevation of the center point of each grid is calculated. The calculated grid points are input into the EFDC model, and after processing, the river topography data required for model construction is obtained. (14) Boundary Input Input the river topography data of the target river obtained in step (13), the flow data, water quality concentration and water level data of the target river obtained in step (11) into the EFDC model to obtain the construction of the hydrodynamic water quality model of the target river. (15) Model simulation By comparing the measured data with the simulated data obtained from the hydrodynamic and water quality model of the target river obtained in step (14), the calibration and verification of the hydrodynamic and water quality model of the target river is completed. Based on the calibrated and verified hydrodynamic and water quality model, it is possible to obtain flow data, water level data, water quality data and flow velocity data of any grid and any cross section.
[0009] Furthermore, the water quality concentration data mentioned in step (11) includes chemical oxygen demand, ammonia nitrogen concentration, total phosphorus concentration and nitrate concentration.
[0010] Furthermore, the high-resolution satellite image mentioned in step (12) refers to a satellite image with a resolution higher than 10m.
[0011] Furthermore, in step (2), the hydrological and hydrodynamic index data includes flow rate data, water level data, flow velocity data, and water temperature data. The chemical indicator data include chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus; The biometric data includes chlorophyll a concentration.
[0012] The reason why this invention can simultaneously output multiple indicators of aquatic ecological integrity, solving the problem that traditional methods can only achieve single-target output, is as follows: By constructing a decision tree-based learner with a multivariate joint splitting criterion and combining it with a cross-index residual feature enhancement mechanism, an end-to-end collaborative mapping from multi-source environmental factor inputs to multiple water ecological integrity evaluation indicators was achieved, overcoming the ecological logic inconsistency problem caused by traditional independent modeling.
[0013] Beneficial Effects: The multi-input, multi-output river water ecological integrity prediction method disclosed in this invention has the following beneficial effects: This invention can simultaneously output multiple indicators of aquatic ecological integrity, solving the problem that traditional methods can only achieve a single target output. Attached Figure Description
[0014] Figure 1 This is a flowchart of a multi-input, multi-output method for predicting the ecological integrity of rivers, as disclosed in this invention.
[0015] Figure 2 This is a schematic diagram of the river boundary data obtained from the Ovi map in Example 1.
[0016] Figure 3 This is a schematic diagram of the river topographic data obtained by secondary processing based on DEM data extraction in Example 1.
[0017] Figure 4 This is a flow change diagram of the Xiantao hydrological station in the middle and lower reaches of the Han River in Example 1.
[0018] Figure 5 This is a flow change diagram of the Luoshan hydrological station in the middle and lower reaches of the Hanjiang River in Example 1.
[0019] Figure 6 This is a schematic diagram of the water level data in Example 1.
[0020] Figures 7-10 This is a schematic diagram of the water quality concentration data in Example 1.
[0021] Figure 11 This is a schematic diagram of the water temperature data in Example 1.
[0022] Figure 12This is a schematic diagram of the grid after it was drawn in Example 1.
[0023] Figure 13 This is a schematic diagram of the terrain after processing in Example 1.
[0024] Figure 14 This is a schematic diagram of the hydrodynamic boundary treatment in Example 1.
[0025] Figure 15 This is a schematic diagram of the water quality boundary treatment in Example 1.
[0026] Figure 16 This is a schematic diagram comparing the measured and simulated water levels in Hanchuan in Example 1.
[0027] Figure 17 This is a schematic diagram comparing the measured and simulated water levels at Guishan in Example 1.
[0028] Figures 18-23 This is a schematic diagram comparing the simulated and measured results of water quality indicators in Example 1.
[0029] Figures 24-33 This is a schematic diagram of the input and output indicators of the hydrodynamic and water quality model of the middle and lower reaches of the Han River in Example 1. The black curve represents the input variable, and the red curve represents the output variable.
[0030] Figure 34 This is a schematic diagram of a coupled model for the integrity of the water ecology in the middle and lower reaches of the Han River.
[0031] Figure 35 This is a schematic diagram comparing the training and prediction results of paraplankton algal density based on radiofrequency (RF).
[0032] Figure 36 This is a schematic diagram comparing the training and prediction results of the phytoplankton diversity index based on radiometrics (RF).
[0033] Figure 37 This is a schematic diagram comparing the training and prediction results of the RF-based phytoplankton richness index.
[0034] Figure 38 This is a schematic diagram illustrating the validation results of a coupled model for aquatic ecological integrity, using phytoplankton density as an indicator.
[0035] Figure 39 This is a schematic diagram illustrating the validation results of a coupled model for aquatic ecological integrity, using the phytophyte diversity index as an indicator.
[0036] Figure 40 This is a schematic diagram illustrating the validation results of a coupled aquatic ecosystem integrity model using the phytoplankton richness index as an indicator. Detailed Implementation
[0037] The specific embodiments of the present invention are described in detail below.
[0038] The "range" disclosed in this invention is defined by a lower limit and an upper limit. A given range is defined by selecting a lower limit and an upper limit, which define the boundaries of a particular range. Ranges defined in this way can include or exclude endpoints and can be arbitrarily combined; that is, any lower limit can be combined with any upper limit to form a range. For example, if a range of 10–50 is listed for a specific parameter, it is also expected that ranges of 10–40 and 20–50 are also included. Furthermore, if the minimum range values are 1 and 2, and the maximum range values are 3, 4, and 5, then the following ranges are all expected: 1–3, 1–4, 1–5, 2–3, 2–4, and 2–5. In this application, unless otherwise stated, the numerical range "a–b" represents a shortened representation of any combination of real numbers between a and b, where a and b are real numbers. For example, the numerical range "0–5" means that all real numbers between "0–5" have been listed herein; "0–5" is merely a shortened representation of these numerical combinations.
[0039] Unless otherwise specified, all embodiments and optional embodiments of this application can be combined to form new technical solutions.
[0040] Unless otherwise specified, all technical features and optional technical features of this application may be combined to form new technical solutions.
[0041] Unless otherwise specified, all steps in this application may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), indicating that the method may include steps (a) and (b) performed sequentially, or it may include steps (b) and (a) performed sequentially. For example, the mention that the method may also include step (c) indicates that step (c) may be added to the method in any order. For example, the method may include steps (a), (b), and (c), or it may include steps (a), (c), and (b), or it may include steps (c), (a), and (b), etc.
[0042] Unless otherwise specified, the terms "comprising" and "including" as used in this application can be open-ended or closed-ended. For example, "comprising" and "including" can mean that other components not listed may also be included, or that only the listed components may be included.
[0043] Unless otherwise specified, the reaction will proceed under normal temperature and pressure conditions.
[0044] Unless otherwise specified, all parts or percentages are by weight or by weight percentage.
[0045] In this invention, all the substances used are known substances that can be purchased or synthesized by known methods.
[0046] In this invention, all the devices or equipment used are conventional devices or equipment known in the art and are readily available.
[0047] Example 1 The technical solution of the present invention is verified by taking the middle and lower reaches of the Han River as an example.
[0048] Region Determination: Since the sampling points for the assessment of the water ecological integrity of the middle and lower reaches of the Han River are located downstream of Xiantao and upstream of the confluence with the Yangtze River, the simulated river region is the section of the Han River from Xiantao to Hankou, which is 150 km long. In the initial modeling stage, the boundary outline of the river section was accurately drawn using Google Earth, and the grid was then divided based on this outline. A multi-input, multi-output method for predicting the ecological integrity of rivers comprises the following steps: (1) Construction of hydrodynamic water quality model (11) Collect data on the target river. The data to be collected for the target river includes river boundary data, river topography data, flow data, water level data, water quality concentration data, and water temperature data, among which: The river boundary data was drawn from Ovi Maps and is used to define the boundaries for constructing the hydrodynamic and water quality model. A schematic diagram is shown below. Figure 2 As shown; The river channel topographic data is obtained by secondary processing based on DEM data extraction, as illustrated below. Figure 3 As shown; The flow data is based on hydrological yearbooks, and the flow change graph of Xiantao Hydrological Station is shown below. Figure 4 As shown in the diagram, the flow rate changes at the Luoshan Hydrological Station are as follows: Figure 5 As shown; The water level data is based on hydrological yearbooks, and its schematic diagram is shown below. Figure 6 As shown; The water quality concentration data was obtained from the National Surface Water Quality Automatic Monitoring Real-time Data Release System, as shown in the diagram below. Figure 7-10 As shown; The water temperature data is based on hydrological yearbooks, and its schematic diagram is shown below. Figure 11 As shown; (12) Mesh generation: Orthogonal curve meshes were generated using Delft3D software. After optimization, the model was divided into 9720 meshes, with the largest mesh size being 391×33.5 m and the smallest mesh size being 14.5×18.8 m. A detailed schematic diagram is shown below. Figure 12 As shown, the specific steps of step (12) are as follows: Obtain a high-resolution satellite image of the target river and import it into ArcGIS software. Then, using the built-in editing tools of ArcGIS, outline the river boundary of the target river according to the range of the target river in the high-resolution satellite image. Import the depicted river boundary into globalmapper and convert its format to one that Delft3D can recognize. Import the converted river boundary file into Delft3D, divide the river into meshes in Delft3D, and export it as a .grid format mesh data file required for model building. (13) Terrain Processing – Based on the riverbed topographic data from discrete elevation sampling points, the riverbed elevation is generated through EFDC interpolation to complete the terrain construction. A specific schematic diagram is shown below. Figure 13 As shown, the specific steps of step (13) are as follows: Import the mesh data file generated in step (12) into the EFDC model and obtain the x and y coordinates of the four corners of the mesh; By taking the x-coordinates and y-coordinates of the four corners of the grid, and calculating the average of the x-coordinates of the four corners and the average of the four x-coordinates, we can obtain the x-coordinates and y-coordinates of the grid center point. Based on the upstream and downstream order of the grid, and using a small amount of measured cross-sectional data and the average topographic slope, the elevation of the center point of each grid is calculated. The calculated grid points are input into the EFDC model, and after processing, the river topography data required for model construction is obtained. (14) Boundary Input Input the river topography data of the target river obtained in step (13), the flow data, water quality concentration and water level data of the target river obtained in step (11) into the EFDC model to obtain the construction of the hydrodynamic water quality model of the target river. Specifically, in this embodiment, the specific steps of step (14) are as follows: (141) Boundary input (hydrodynamic boundary): such as Figure 14 As shown, the simulation area is located at the confluence of two rivers, with a total of four inflow boundaries and one outflow boundary. The specific settings are as follows: the inflow boundary of the Hanjiang main stream uses daily flow monitoring data from the Xiantao hydrological station; the tributary boundary incorporates flow data from the Hanchuan and Xigou sluices on the Hanbei River; the inflow boundary of the Yangtze River is defined by the daily flow sequence from the Luoshan hydrological station. Model validation and parameter calibration selected the Guishan and Hanchuan hydrological stations as key nodes. The outflow boundary of the entire simulation area is set as the daily average water level observation value from the Hankou hydrological station on the Yangtze River. (142) Boundary Input (Water Quality Boundary): such as Figure 15As shown, the water quality model is configured with the following boundary conditions and simulation parameters: (1421) Boundary Condition Setting: For the main stream of the Han River, the pollution load from the upstream flow of the Xiantao hydrological station is used as the upstream water quality boundary input, and the water quality monitoring data of Hannan Village near the Xiantao station is selected as the specific value basis. For the northern tributary, the water quality data of the Xigou Sluice Gate of the Hanbei River is used as the boundary condition for the northern surface source pollution and tributary input of the simulation area. For the main stream of the Yangtze River, the water quality monitoring section data of Yangsigang near Luoshan station is used as the input boundary of the Yangtze River section.
[0049] (1422) Simulation Indicators and Scales: The main water quality indicators simulated by the model include COD. Mn The data included DO, TP, NH3-N, and water temperature. All input and output timescales for water quality data were set to monthly. Algal dynamics simulation relied on algal density data from the Han River main stream and Chla field monitoring data, with carbon content of cyanobacteria, green algae, and diatoms calculated from measured algal densities.
[0050] (15) Model simulation By comparing the measured data with the simulated data obtained from the hydrodynamic and water quality model of the target river obtained in step (14), the calibration and verification of the hydrodynamic and water quality model of the target river is completed. Based on the calibrated and verified hydrodynamic and water quality model, it is possible to obtain flow data, water level data, water quality data and flow velocity data of any grid and any cross section; Specifically, in this embodiment, step (15) is as follows: (151) Model Simulation (Hydrodynamic Simulation Calibration and Verification): Based on daily-scale simulation data from 2014 to 2015, the parameters of the hydrodynamic model were calibrated. The Hanjiang River water level at Hanchuan Station was used as the verification basis, while the Yangtze River water level was assessed by interpolating data from Shijitou and Hankou Stations to obtain a representative water level near Guishan. The simulation results are shown in Figure 1. like Figure 16 As shown, the simulated and measured values of Hanchuan Station in 2014 showed good agreement in terms of time variation trend and consistent fluctuation pattern. Their root mean square error (RMSE) and relative error (ARE) were 0.22% and 0.88% respectively (within 5%), and the Nash efficiency coefficient (NSE) reached 0.99 (above 90%). like Figure 17 As shown, the simulated water level at Guishan Station in 2014 also performed well, with RMSE and ARE of 0.23 and 0.88% respectively, and NSE of 0.99. The simulation and measured results for both stations met the model accuracy requirements and demonstrated good performance. In 2015, the RMSE and ARE of the water level at Hanchuan Station were 0.33% and 1.30%, respectively, with an NSE of 0.97; the RMSE and ARE of the water level at Guishan Station in 2015 were 0.20% and 0.80%, respectively, with an NSE of 0.97. Further verification of the water levels at Hanchuan and Guishan Stations was conducted using simulated data from 2016. The results showed that in 2016, the RMSE and ARE of Hanchuan Station were 0.28% and 1.12%, respectively, with an NSE of 0.99; while the RMSE and ARE of Guishan Station in 2016 were 0.21% and 0.76%, respectively, with an NSE of 0.99.
[0051] The results above indicate that the simulated and measured water levels at the two stations from 2014 to 2016 remain highly consistent, meeting the accuracy requirements of the model simulation and can be used for subsequent simulation analysis under different scenarios.
[0052] Table 1: Verification Results of Hydrodynamic Module (152) Model simulation (water quality simulation and calibration verification): Monthly water quality monitoring data of Zongguan section from 2014 to 2015 were used to calibrate and verify the model parameters. The main water quality parameters are water temperature, DO, and COD. Mn TP and NH3-N. Water temperature is crucial to the response rate of various water quality indicators in rivers, directly affecting the concentration of each water quality parameter. The accuracy of water temperature simulation directly impacts the accuracy of other water quality simulations. Relative error (RE) and root mean square error (RMSE) are used to evaluate the accuracy of the water quality simulation results. Initial values for model parameters such as algal metabolism, growth, and sedimentation are provided in relevant literature, and the final parameter values are obtained through model calibration. The calibration results for each water quality parameter are shown below. Figure 18-23 As shown, the model demonstrates high accuracy in simulating all key water quality indicators, with relative errors (REs) all controlled below 20%. Specifically, the RMSE for NH3-N is 0.054, with a relative error of 15.3%; the RMSE for DO is 0.465, with a relative error of 0.7%; and the RMSE for COD is... Mn The RMSE for TP was 0.456 with a relative error of 5.9%; the RMSE for TP was 0.024 with a relative error of 6.8%; and the RMSE for water temperature was 0.82 with a relative error of 0.13%. Water temperature was the best-performing indicator in the water quality module. In summary, the deviations between the predicted and measured values for each indicator were within acceptable ranges, demonstrating the effectiveness and reliability of the model.
[0053] (2) Output of hydrodynamic water quality model Based on the evaluation points of the aquatic ecological integrity index, corresponding grid data are extracted from the hydrodynamic and water quality model of the target river obtained in step (1). The grid data includes hydrological and hydrodynamic index data, chemical index data, and biological index data. Specifically, in this embodiment, the specific steps of step (2) are as follows: Using water level, flow velocity, water temperature, and nutrients output from the hydrodynamic water quality model as physical and chemical explanatory variables, and Chla output from the phytoplankton module of the hydrodynamic water quality model as a biological explanatory variable, all three types of explanatory variables are incorporated into the random forest model system as input boundaries. Multiple indicators of aquatic ecological integrity are used as outputs, including phytoplankton algal density (FYZW_CELL), phytoplankton diversity (FYZW_SW), and richness index (FYZW_M). A multi-input multi-output response relationship is established to construct a coupled model of aquatic ecological integrity in the middle and lower reaches of the Han River.
[0054] (3) Construction and validation of the random forest (RF) model, as shown in the diagram below. Figure 34 As shown.
[0055] (31) Data allocation: The physical, chemical and biological indicators output by the hydrodynamic water quality model are used as input boundaries, and the water ecological integrity indicators are used as outputs. The RF model data is divided into a training set of 70% and a test set of 30% to construct a random forest model. (32) Model validation: Based on the data allocation results, the simulated data and measured data of multiple water ecological integrity indicators output by the random forest model are compared to calibrate and validate the model's accuracy and effectiveness.
[0056] The code for building the random forest model is shown below: library(randomForest) library(dplyr) library(ggplot2) library(caret) library(reshape2) # Set the working directory (you need to specify the path if the file is not in the current directory) setwd("D: / R / RF") # Use forward slash / or double backslash \\[5,11](@ref) as the path separator # Import CSV file data <- read.csv( file = "WEI.csv", # Filename (including path) header = TRUE, # The first row contains column names (default TRUE) sep = ",", # The separator is a comma (default) fileEncoding = "UTF-8", # Chinese files need to specify encoding [4,7](@ref) na.strings = c("NA","") # Convert missing value markers to NA[7](@ref) ) head(data) # Divide the dataset into training and test sets (80% training, 20% test). train_idx <- createDataPartition(data$WEI, p = 0.7, list = FALSE) train_data <- data[train_idx, ] test_data <- data[-train_idx, ] cat("\nNumber of training set samples:",nrow(train_data)) cat("\nTest set sample count:",nrow(test_data)) # Building a Random Forest Model # Use default parameters rf_model <- randomForest( WEI ~ ., data = train_data, importance = TRUE, # Calculate feature importance ntree = 500, # Number of trees mtry = 4, # Number of features used per tree nodesize = 5, # Minimum size of terminal nodes do.trace = 100 # Display progress every 100 trees ) cat(" Random Forest Model Summary: ") print(rf_model) # Make predictions on the test set predictions <- predict(rf_model, newdata = test_data) # Model Evaluation actual <- test_data$WEI mse <- mean((predictions - actual)^2) rmse <- sqrt(mse) r2 <- 1 - sum((actual - predictions)^2) / sum((actual - mean(actual))^2) cat(" Model performance evaluation: ") cat("Mean Squared Error (MSE):", mse, ") ") cat("Root Mean Square Error (RMSE):", rmse, ") ") cat("R squared(R 2 ):", r2, " ") cat("Mean Absolute Error:", mean(abs(predictions - actual)), " ").
[0057] like Figures 35-40 As shown, the measured and predicted values of each indicator show a relatively consistent trend. We use R... 2 The model was evaluated using NSE and RMSE. Overall, the R-value of phytoplankton algal density was [value missing]. 2 The NSE and RMSE were 0.81, 0.79, and 0.256, respectively, and the RSE of the phytoplankton diversity index was... 2 The NSE and RMSE values were 0.93, 0.91, and 0.181, respectively. The RSE for phytoplankton richness was... 2 The NSE and RMSE were 0.84, 0.81, and 0.344, respectively, indicating that the simulation and prediction effects of the three indicators were all good. The model validation set data showed that the simulated and measured values of phytoplankton algal density, diversity index, and richness index were consistent with the observed values. 2 The values are 0.94, 0.95, and 0.89 respectively, further demonstrating that the model has a certain degree of credibility and generalization ability.
[0058] Example 1 uses the middle and lower reaches of the Han River as a case to verify the technical solution of the present invention, proving that the technical solution of the present invention can effectively realize the simulation and prediction of multiple indicators of water ecological integrity.
[0059] The embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments, and various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A multi-input, multi-output method for predicting the ecological integrity of rivers, characterized in that, The steps are as follows: (1) Construction of hydrodynamic water quality model; (2) Output of hydrodynamic water quality model Based on the evaluation points of the water ecological integrity index, the corresponding grid data is extracted from the hydrodynamic and water quality model of the target river obtained in step (1). The grid data includes hydrological and hydrodynamic index data, chemical index data and biological index data. (3) Construction, validation and evaluation of the random forest model (31) Data allocation: The hydrological and hydrodynamic index data, chemical index data and biological index data output by the hydrodynamic and water quality model in step (2) are used as input boundaries, and multiple indicators of water ecological integrity are used as output indicators. The training set and the validation set are allocated to construct a random forest model. (32) Validation and evaluation of the random forest model: Based on the data allocation results, the simulated data and measured data of multiple water ecological integrity indicators output by the random forest model are compared to calibrate and verify the accuracy and effectiveness of the model.
2. The multi-input multi-output method for predicting river water ecological integrity as described in claim 1, characterized in that, The steps of step (1) are as follows: (11) Collect data on the target river. The data to be collected for the target river includes river boundary data, river topography data, flow data, water level data, water quality concentration data, and water temperature data, among which: The river boundary data was drawn from the Ovi map and is used to define the boundaries for constructing the hydrodynamic and water quality model. The river topographic data is obtained by secondary processing based on DEM data extraction. The flow data was obtained based on hydrological yearbooks; The water level data was obtained based on hydrological yearbooks; The water quality concentration data was obtained from the National Surface Water Quality Automatic Monitoring Real-time Data Release System. The water temperature data was obtained based on a hydrological yearbook. (12) Grid generation Obtain a high-resolution satellite image of the target river and import it into ArcGIS software. Then, using the built-in editing tools of ArcGIS, outline the river boundary of the target river according to the range of the target river in the high-resolution satellite image. Import the depicted river boundary into globalmapper and convert its format to one that Delft3D can recognize. Import the converted river boundary file into Delft3D, divide the river into meshes in Delft3D, and export it as a .grid format mesh data file required for model building. (13) Terrain processing Import the mesh data file generated in step (12) into the EFDC model and obtain the x and y coordinates of the four corners of the mesh. By taking the x-coordinates and y-coordinates of the four corners of the grid, and calculating the average of the x-coordinates of the four corners and the average of the four x-coordinates, we can obtain the x-coordinates and y-coordinates of the grid center point. Based on the upstream and downstream order of the grid, and using a small amount of measured cross-sectional data and the average topographic slope, the elevation of the center point of each grid is calculated. The calculated grid points are input into the EFDC model, and after processing, the river topography data required for model construction is obtained. (14) Boundary Input Input the river topography data of the target river obtained in step (13), the flow data, water quality concentration and water level data of the target river obtained in step (11) into the EFDC model to obtain the construction of the hydrodynamic water quality model of the target river. (15) Model simulation By comparing the measured data with the simulated data obtained from the hydrodynamic and water quality model of the target river obtained in step (14), the calibration and verification of the hydrodynamic and water quality model of the target river is completed.
3. The multi-input multi-output method for predicting the integrity of river water ecology as described in claim 2, characterized in that, The water quality concentration data mentioned in step (11) include chemical oxygen demand, ammonia nitrogen concentration, total phosphorus concentration and nitrate concentration.
4. The multi-input multi-output method for predicting river water ecological integrity as described in claim 2, characterized in that, The high-resolution satellite image mentioned in step (12) refers to a satellite image with a resolution higher than 10m.
5. The multi-input multi-output method for predicting river water ecological integrity as described in claim 1, characterized in that, In step (2), the hydrological and hydrodynamic index data includes flow rate data, water level data, flow velocity data, and water temperature data. The chemical indicator data include chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus; The biometric data includes chlorophyll a concentration.
6. The multi-input multi-output method for predicting river water ecological integrity as described in claim 1, characterized in that, The random forest model described in step (31) is implemented using R language. The workflow of the random forest model includes calling the tool package in R language, importing the dataset, splitting the dataset, and establishing a model performance evaluation.
Citation Information
Patent Citations
CN119090093A