A method for simulating nutrient concentration in lakes and reservoirs based on a coupling mechanism model and deep learning
By coupling the water dynamic water quality mechanism model of lake reservoirs and deep learning methods, the balance problem between accuracy, interpretability and efficiency of traditional simulation methods is solved, and high-precision and interpretability of lake reservoir nutrient concentration simulation and prediction are achieved.
Patent Information
- Application Number
- CN202510447526.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional lake reservoir nutrient concentration simulation methods are difficult to balance high precision, interpretability and computational efficiency. The mechanism model is affected by parameter uncertainty and computational complexity. The deep learning model lacks physical constraints and the ‘black box’ characteristics hinder the identification of key driver factors.
Using the method of coupling mechanism model and deep learning, the water quality mechanism model of the lake reservoir water dynamics is constructed, and the internal circulation process flux of nutrients is calculated, and it is used as the input of the deep learning model. The physical information neural network is regularly constructed through physical guidance to realize the learning and replacement of the mechanism model by deep learning.
The simulation accuracy and interpretability of nutrient concentration in the lake reservoir are improved, key driver factors are identified, the driving mechanism of nutrient changes is revealed, and the generalization ability of the model and the retention of physical laws are enhanced.
Smart Images

Figure CN119962408B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simulation and prediction of nutrient concentration of lakes and reservoirs, and specifically relates to a method for simulating nutrient concentration of lakes and reservoirs by coupling a mechanism model and deep learning. Background Art
[0002] Changes in nutrient concentrations in lake ecosystems are driven by complex biogeochemical processes, including interactions between hydrological dynamics, climate change, human activities, and biological processes. Traditional nutrient concentration simulation methods mainly include mechanism-based process models and data-driven statistical or machine learning models. However, these two types of methods have their own limitations, and it is difficult to strike a balance between high accuracy, interpretability, and computational efficiency. Although water quality models based on physical and chemical equations (such as EFDC) can analyze the intrinsic mechanism of nutrient cycling, they are limited by high parameter uncertainty and high computational complexity, making it difficult to achieve efficient simulation at a regional scale; while machine learning models (such as LSTM) have powerful nonlinear time series modeling capabilities, the lack of physical constraints causes the simulation results to be out of touch with ecological laws, and the "black box" characteristics hinder the identification of key driving factors, and the generalization performance is significantly reduced in data-scarce scenarios.
[0003] The implementation methods for coupling the mechanism model and the deep learning model include feedforward / feedback correction model, embedding based on physical constraints, etc. Feedforward correction refers to the mechanism model as the base layer to generate preliminary results to correct the residuals of the deep learning model. At present, most studies use state variables generated by the mechanism model as the input of the data-driven model. However, the state variables cannot reflect the migration and transformation process of nutrients in lakes and reservoirs. At the same time, the method of embedding the deep learning model based on physical constraints has better interpretability, but its advantage of retaining the causal interpretation of the mechanism model has not been effectively utilized. Therefore, there is an urgent need for a method that can characterize the migration and transformation process of nutrients in lakes and reservoirs, and has high-precision simulation capabilities and interpretability, to efficiently simulate and predict the nutrient concentration of lakes and reservoirs. Summary of the invention
[0004] To solve the above problems, the present invention proposes a method for simulating the nutrient concentration of lakes and reservoirs by coupling a mechanism model with deep learning. This method can better simulate the migration and transformation process of nutrients, making it closer to the actual situation, thereby improving the simulation effect of the nutrient concentration of lakes and reservoirs; in addition, compared with only using a mechanism model, the coupling method can better play the advantages of explainable machine learning, identify key driving factors, reveal the driving mechanism of nutrient changes, and improve the simulation and prediction accuracy of the nutrient concentration of lakes and reservoirs.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for simulating nutrient concentration in lakes and reservoirs by coupling mechanism model and deep learning, comprising the following steps:
[0007] S1. Construct a model of lake and reservoir hydrodynamics and water quality mechanism, calculate the internal circulation process flux of nutrients, and obtain secondary data of flux indicators and nutrient concentrations;
[0008] The specific steps of step S1 are:
[0009] S11, collecting basic data of lake and reservoir areas, and preprocessing the basic data to obtain a basic database;
[0010] S12, setting boundary conditions based on the preprocessed data, then performing grid division based on the geographic information of the lake and reservoir, and setting initial model parameters to construct a lake and reservoir hydrodynamic and water quality mechanism model;
[0011] S13. Automatic monitoring data and flux measurement data are used to evaluate the simulation effect of the lake and reservoir hydrodynamic and water quality mechanism model. The Sobol method is used to determine the sensitive parameters, and then the Bayesian optimization is used to calibrate the sensitive parameters of the lake and reservoir hydrodynamic and water quality mechanism model;
[0012] S14. Use mean square error and Nash efficiency coefficient to evaluate the accuracy of the lake and reservoir hydrodynamic and water quality mechanism model, so as to make the output of the lake and reservoir hydrodynamic and water quality mechanism model conform to the actual hydrological and water quality laws;
[0013] S15. Run the calibrated lake and reservoir hydrodynamic and water quality mechanism model to generate secondary internal circulation data of flux indicators and nutrient concentrations;
[0014] S16, aligning the secondary inner cycle data with the basic data in time and constructing a new database for subsequent deep learning model training;
[0015] S2. By constructing a physical information neural network through physical guidance regularization of the loss function, and using the Bayesian optimization algorithm for hyperparameter optimization, deep learning is used to learn and replace the lake and reservoir hydrodynamic and water quality mechanism model;
[0016] The specific process of step S2 is:
[0017] S21, taking the automatically monitored water quality index and the secondary internal circulation flux index generated in step S1 as input, and taking the secondary nutrient concentration output by the lake and reservoir hydrodynamic water quality mechanism model as output;
[0018] S22. When building a deep learning model, the mass conservation of nutrient changes is characterized by introducing a physics-guided regularization term in the mean square error loss;
[0019] S23, pre-training the deep learning model so that the deep learning model can initially grasp the physical laws, and a self-paced learning strategy is adopted in the pre-training stage;
[0020] S24, with the goal of minimizing the error of the validation set, dynamically search for the optimal combination of learning rate, regularization coefficient, and number of hidden layer nodes;
[0021] S25. Use the Bayesian optimization method to find the best hyperparameter combination process, then use the Gaussian process to model the hyperparameter space, and iteratively update the posterior distribution to improve convergence efficiency;
[0022] S26. Compare the deep learning simulation results with the output of the lake and reservoir hydrodynamic and water quality mechanism model, so that the deep learning model can achieve efficient substitution while retaining the physical laws;
[0023] S3, setting some weights of the physical information neural network as variable parameters, fine-tuning the deep learning model established in step S2, and realizing the simulation of the actual observed nutrient concentration;
[0024] The specific process of step S3 is:
[0025] S31, fix the main structure of the physical information neural network, set the weights of the last two fully connected layers as adjustable parameters to adapt to the distribution characteristics of the actual monitoring data; wherein the main structure includes a GRU layer and an attention module;
[0026] S32, merging the water quality index and the secondary internal circulation process flux index obtained by automatic monitoring into a fine-tuning data set, and then taking the nutrient concentration obtained by automatic monitoring as the output, balancing the data ratio through weighted sampling;
[0027] S33. Apply transfer learning to optimize the training process so that the deep learning model can adapt to the distribution characteristics of the actual monitoring data;
[0028] S34, fine-tuning the simulation accuracy of the deep learning model to the actual concentration through the Nash efficiency coefficient and the root mean square error quantification;
[0029] S4. Use deep learning interpretation methods to identify the main influencing factors of nutrient concentration fluctuations, analyze the driving mechanism of nutrient concentration evolution, and simulate and predict the nutrient concentration of lakes and reservoirs.
[0030] Preferably, the internal circulation process of nutrients in step S1 includes atmospheric deposition of nitrogen and phosphorus, sedimentation and resuspension of nitrogen and phosphorus at the interface between water and sediment, and nitrification and denitrification of nitrogen.
[0031] Preferably, the basic data in step S11 include water quality data, hydrological data and meteorological data; the water quality data include nitrogen concentration, phosphorus concentration, pH value, chlorophyll a concentration, dissolved oxygen content and turbidity; the hydrological data include water level, flow rate, flow, water temperature and water depth; the meteorological data include rainfall and wind speed.
[0032] Preferably, the specific process of data preprocessing in step S11 is:
[0033] S111. Eliminate outliers from the collected basic data based on statistical methods;
[0034] S112, data normalization was performed using Z-score standardization;
[0035] S113. Use time series interpolation method to fill missing values in water quality data.
[0036] Preferably, in step S12, the boundary conditions include inflow, outflow and meteorological forcing, the geographical information of the lake includes topography and water depth, and the model parameters include turbulence coefficient, diffusion coefficient and reaction rate constant.
[0037] Preferably, the specific method of embedding the physical constraints in step S22 is:
[0038] S221, by introducing the flux conservation term, based on the mass balance condition, calculating the residual and adding the loss function; wherein the mass balance condition is the sum of the absolute values of the difference between the predicted nutrient concentration and the observed nutrient concentration;
[0039] S222. The physical relationship between variables is defined through partial dependence diagrams, and the lake hydrodynamic and water quality mechanism model is forced to learn the response through regularization.
[0040] Preferably, the specific process of step S4 is:
[0041] S41. In the fine-tuned deep learning model, the DeepSHAP method was applied to calculate the contribution of each input feature to the simulation results of nutrient concentration;
[0042] S42. Based on the size and distribution of SHAP values calculated by the DeepSHAP method, identify the factors that have the greatest impact on the change in nutrient concentration;
[0043] S43. By drawing partial dependence diagrams, we can identify the response relationship between nutrient concentration and driving mechanism, and simulate and predict the nutrient concentration of lakes and reservoirs.
[0044] After adopting the above technical scheme, the present invention has the following beneficial effects: the coupling mechanism model of the present invention and the deep learning method for simulating the nutrient concentration of lakes and reservoirs can make full use of the physical constraints of the hydrodynamic and water quality mechanism model of lakes and reservoirs and the efficient simulation ability of the deep learning method, and the flux data output by the hydrodynamic and water quality mechanism model of lakes and reservoirs is used as the input of the deep learning model to learn the time variation pattern of nutrient concentration, and at the same time, the generalization ability of the hydrodynamic and water quality mechanism model of lakes and reservoirs can be improved. In addition, the loss function based on physical laws is used to constrain the learning process of the deep learning model, so that it can better simulate the migration and transformation process of nutrients, which is closer to the actual situation, thereby improving the simulation effect of the nutrient concentration of lakes and reservoirs. In addition, the application of interpretable machine learning can identify key driving factors and their response relationships. Therefore, the present invention not only provides a high-precision simulation method for the nutrient concentration of lakes and reservoirs, but also the coupling method can better play the advantages of interpretable machine learning, provide effective support for the management of nutrients in lakes and reservoirs, and has significant advantages and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flow chart of the present invention;
[0046] Figure 2 It is a flowchart of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] like Figure 1 and Figure 2 As shown, a method for simulating nutrient concentration in lakes and reservoirs by coupling mechanism model and deep learning includes the following steps:
[0049] S1. Construct a model of lake and reservoir hydrodynamics and water quality mechanism, calculate the internal circulation process flux of nutrients, and obtain secondary data of flux indicators and nutrient concentrations;
[0050] The internal circulation process of nutrients in step S1 includes atmospheric deposition of nitrogen and phosphorus, sedimentation and resuspension of nitrogen and phosphorus at the interface between water and sediment, and nitrification and denitrification of nitrogen;
[0051] The specific steps of step S1 are:
[0052] S11, collecting basic data of lake and reservoir areas, and preprocessing the basic data to obtain a basic database;
[0053] The basic data in step S11 include water quality data, hydrological data and meteorological data; the water quality data include nitrogen concentration, phosphorus concentration, pH value, chlorophyll a concentration, dissolved oxygen content and turbidity; the hydrological data include water level, flow rate, flow, water temperature and water depth; the meteorological data include rainfall and wind speed;
[0054] The specific process of data preprocessing in step S11 is:
[0055] S111. Eliminate outliers from the collected basic data based on statistical methods;
[0056] S112, data normalization was performed using Z-score standardization;
[0057] S113. Use time series interpolation method to fill missing values in water quality data;
[0058] S12, setting boundary conditions based on the preprocessed data, then performing grid division based on the geographic information of the lake and reservoir, and setting initial model parameters to construct a lake and reservoir hydrodynamic and water quality mechanism model;
[0059] In step S12, the boundary conditions include inflow, outflow and meteorological forcing, the geographical information of the lake includes topography and water depth, and the model parameters include turbulence coefficient, diffusion coefficient and reaction rate constant;
[0060] S13. Automatic monitoring data and flux measurement data are used to evaluate the simulation effect of the lake and reservoir hydrodynamic and water quality mechanism model. The Sobol method is used to determine the sensitive parameters, and then the Bayesian optimization is used to calibrate the sensitive parameters of the lake and reservoir hydrodynamic and water quality mechanism model;
[0061] S14. Use mean square error and Nash efficiency coefficient to evaluate the accuracy of the lake and reservoir hydrodynamic and water quality mechanism model, so as to make the output of the lake and reservoir hydrodynamic and water quality mechanism model conform to the actual hydrological and water quality laws;
[0062] S15. Run the calibrated lake and reservoir hydrodynamic and water quality mechanism model to generate secondary internal circulation data of flux indicators and nutrient concentrations;
[0063] S16, aligning the secondary inner cycle data with the basic data in time and constructing a new database for subsequent deep learning model training;
[0064] S2. By constructing a physical information neural network through physical guidance regularization of the loss function, and using the Bayesian optimization algorithm for hyperparameter optimization, deep learning is used to learn and replace the lake and reservoir hydrodynamic and water quality mechanism model;
[0065] The specific process of step S2 is:
[0066] S21, taking the automatically monitored water quality index and the secondary internal circulation flux index generated in step S1 as input, and taking the secondary nutrient concentration output by the lake and reservoir hydrodynamic water quality mechanism model as output;
[0067] S22. When building a deep learning model, the mass conservation of nutrient changes is characterized by introducing a physics-guided regularization term in the mean square error loss;
[0068] The specific method of embedding the physical constraints in step S22 is:
[0069] S221, by introducing the flux conservation term, based on the mass balance condition, calculating the residual and adding the loss function; wherein the mass balance condition is the sum of the absolute values of the difference between the predicted nutrient concentration and the observed nutrient concentration;
[0070] S222, defining the physical relationship between variables through partial dependence diagrams, and forcing the lake and reservoir hydrodynamic and water quality mechanism model to learn responses through regularization;
[0071] S23, pre-training the deep learning model so that the deep learning model can initially grasp the physical laws, and a self-paced learning strategy is adopted in the pre-training stage;
[0072] S24, with the goal of minimizing the error of the validation set, dynamically search for the optimal combination of learning rate, regularization coefficient, and number of hidden layer nodes;
[0073] S25. Use the Bayesian optimization method to find the best hyperparameter combination process, then use the Gaussian process to model the hyperparameter space, and iteratively update the posterior distribution to improve convergence efficiency;
[0074] S26. Compare the deep learning simulation results with the output of the lake and reservoir hydrodynamic and water quality mechanism model, so that the deep learning model can achieve efficient substitution while retaining the physical laws;
[0075] S3, setting some weights of the physical information neural network as variable parameters, fine-tuning the deep learning model established in step S2, and realizing the simulation of the actual observed nutrient concentration;
[0076] The specific process of step S3 is:
[0077] S31, fix the main structure of the physical information neural network, set the weights of the last two fully connected layers as adjustable parameters to adapt to the distribution characteristics of the actual monitoring data; wherein the main structure includes a GRU layer and an attention module;
[0078] S32, merging the water quality index and the secondary internal circulation process flux index obtained by automatic monitoring into a fine-tuning data set, and then taking the nutrient concentration obtained by automatic monitoring as the output, balancing the data ratio through weighted sampling;
[0079] S33. Apply transfer learning to optimize the training process so that the deep learning model can adapt to the distribution characteristics of the actual monitoring data;
[0080] S34, fine-tuning the simulation accuracy of the deep learning model to the actual concentration through the Nash efficiency coefficient and the root mean square error quantification;
[0081] S4. Use deep learning interpretation methods to identify the main influencing factors of nutrient concentration fluctuations, analyze the driving mechanism of nutrient concentration evolution, and simulate and predict lake and reservoir nutrient concentrations;
[0082] The specific process of step S4 is:
[0083] S41. In the fine-tuned deep learning model, the DeepSHAP method was applied to calculate the contribution of each input feature to the simulation results of nutrient concentration;
[0084] S42. Based on the size and distribution of SHAP values calculated by the DeepSHAP method, identify the factors that have the greatest impact on the change in nutrient concentration;
[0085] S43. By drawing partial dependence diagrams, we can identify the response relationship between nutrient concentration and driving mechanism, and simulate and predict the nutrient concentration of lakes and reservoirs.
[0086] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for simulating nutrient concentration in lakes and reservoirs by coupling mechanism model and deep learning, characterized in that: The following steps are involved: S1. Construct a model of lake and reservoir hydrodynamics and water quality mechanism, calculate the internal circulation process flux of nutrients, and obtain secondary data of flux indicators and nutrient concentrations; The specific steps of step S1 are: S11, collecting basic data of lake and reservoir areas, and preprocessing the basic data to obtain a basic database; S12, setting boundary conditions based on the preprocessed data, then performing grid division based on the geographic information of the lake and reservoir, and setting initial model parameters to construct a lake and reservoir hydrodynamic and water quality mechanism model; S13. Automatic monitoring data and flux measurement data are used to evaluate the simulation effect of the lake and reservoir hydrodynamic and water quality mechanism model. The Sobol method is used to determine the sensitive parameters, and then the Bayesian optimization is used to calibrate the sensitive parameters of the lake and reservoir hydrodynamic and water quality mechanism model. S14. Use mean square error and Nash efficiency coefficient to evaluate the accuracy of the lake and reservoir hydrodynamic and water quality mechanism model, so as to make the output of the lake and reservoir hydrodynamic and water quality mechanism model conform to the actual hydrological and water quality laws; S15. Run the calibrated lake and reservoir hydrodynamic and water quality mechanism model to generate secondary internal circulation data of flux indicators and nutrient concentrations; S16, aligning the secondary inner cycle data with the basic data in time and constructing a new database for subsequent deep learning model training; S2. By physically guiding the regularization of the loss function, a physical information neural network is constructed, and the Bayesian optimization algorithm is used for hyperparameter optimization to achieve deep learning and replacement of the lake and reservoir hydrodynamic and water quality mechanism model; The specific process of step S2 is: S21, taking the automatically monitored water quality index and the secondary internal circulation flux index generated in step S1 as input, and taking the secondary nutrient concentration output by the lake and reservoir hydrodynamic water quality mechanism model as output; S22. When building a deep learning model, the mass conservation of nutrient changes is characterized by introducing a physics-guided regularization term in the mean square error loss; S23, pre-training the deep learning model so that the deep learning model can initially grasp the physical laws, and a self-paced learning strategy is adopted in the pre-training stage; S24, with the goal of minimizing the error of the validation set, dynamically search for the optimal combination of learning rate, regularization coefficient, and number of hidden layer nodes; S25. Use the Bayesian optimization method to find the best hyperparameter combination process, then use the Gaussian process to model the hyperparameter space, and iteratively update the posterior distribution to improve convergence efficiency; S26. Compare the deep learning simulation results with the output of the lake and reservoir hydrodynamic and water quality mechanism model, so that the deep learning model can achieve efficient substitution while retaining the physical laws; S3, setting some weights of the physical information neural network as variable parameters, fine-tuning the deep learning model established in step S2, and realizing the simulation of the actual observed nutrient concentration; The specific process of step S3 is: S31, fix the main structure of the physical information neural network, set the weights of the last two fully connected layers as adjustable parameters to adapt to the distribution characteristics of the actual monitoring data; wherein the main structure includes a GRU layer and an attention module; S32, merging the water quality index and the secondary internal circulation process flux index obtained by automatic monitoring into a fine-tuning data set, and then taking the nutrient concentration obtained by automatic monitoring as the output, balancing the data ratio through weighted sampling; S33. Apply transfer learning to optimize the training process so that the deep learning model can adapt to the distribution characteristics of the actual monitoring data; S34. Fine-tune the simulation accuracy of the deep learning model to the actual concentration through the Nash efficiency coefficient and the root mean square error quantification; S4. Use deep learning interpretation methods to identify the main influencing factors of nutrient concentration fluctuations, analyze the driving mechanism of nutrient concentration evolution, and simulate and predict the nutrient concentration of lakes and reservoirs.
2. The method for simulating nutrient concentration in lakes and reservoirs by coupling a mechanism model and deep learning as claimed in claim 1, characterized in that: The internal circulation process of nutrients in step S1 includes atmospheric deposition of nitrogen and phosphorus, sedimentation and resuspension of nitrogen and phosphorus at the interface between water and sediment, and nitrification and denitrification of nitrogen.
3. The method for simulating nutrient concentration in lakes and reservoirs by coupling a mechanism model and deep learning as claimed in claim 1, characterized in that: The basic data in step S11 include water quality data, hydrological data and meteorological data; the water quality data include nitrogen concentration, phosphorus concentration, pH value, chlorophyll a concentration, dissolved oxygen content and turbidity; the hydrological data include water level, flow rate, flow, water temperature and water depth; the meteorological data include rainfall and wind speed.
4. The method for simulating nutrient concentration of lakes and reservoirs by coupling mechanism model and deep learning as claimed in claim 3, characterized in that: The specific process of data preprocessing in step S11 is: S111. Eliminate outliers from the collected basic data based on statistical methods; S112, data normalization was performed using Z-score standardization; S113. Use time series interpolation method to fill missing values in water quality data.
5. The method for simulating nutrient concentration in lakes and reservoirs by coupling a mechanism model and deep learning as claimed in claim 1, characterized in that: In step S12, the boundary conditions include inflow, outflow and meteorological forcing, the geographical information of the lake includes topography and water depth, and the model parameters include turbulence coefficient, diffusion coefficient and reaction rate constant.
6. The method for simulating nutrient concentration of lakes and reservoirs by coupling mechanism model and deep learning as claimed in claim 1, characterized in that: The specific method of embedding the physical constraints in step S22 is: S221, by introducing the flux conservation term, based on the mass balance condition, calculating the residual and adding the loss function; wherein the mass balance condition is the sum of the absolute values of the difference between the predicted nutrient concentration and the observed nutrient concentration; S222. The physical relationship between variables is defined through partial dependence diagrams, and the lake hydrodynamic and water quality mechanism model is forced to learn the response through regularization.
7. The method for simulating nutrient concentration in lakes and reservoirs by coupling a mechanism model and deep learning as claimed in claim 1, characterized in that: The specific process of step S4 is: S41. In the fine-tuned deep learning model, the DeepSHAP method was applied to calculate the contribution of each input feature to the simulation results of nutrient concentration; S42. Based on the size and distribution of SHAP values calculated by the DeepSHAP method, identify the factors that have the greatest impact on the change in nutrient concentration; S43. By drawing partial dependence diagrams, the response relationship between nutrient concentration and driving mechanism can be identified, and the nutrient concentration of lakes and reservoirs can be simulated and predicted.
Citation Information
Patent Citations
Lake and reservoir cyanobacterial bloom prediction method based on adaptive dynamic programming
CN110532646A
Water body remote sensing image expansion and eutrophication prediction method based on atmosphere-water quality multi-modal information
CN118735752A