A method for formulating nutrient criteria for regional lakes and reservoirs based on Bayesian hierarchical models
By integrating lake and reservoir data using Bayesian hierarchical model, the problem of difficulty in considering lake and reservoir heterogeneity in the existing technology is solved, and accurate nutritional benchmarking is achieved, which promotes the protection and restoration of lake and reservoir ecosystems.
Patent Information
- Application Number
- CN202510312875.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-17
AI Technical Summary
When formulating nutrient bases for lake and reservoirs, it is difficult for the existing technology to effectively consider the heterogeneity of lake and reservoirs in the region, resulting in insufficient accuracy in the benchmark setting, affecting the protection and restoration of lake and reservoir ecosystems.
Using a Bayesian hierarchical model method, the data of different lakes and reservoirs were integrated, and the differences within individual lakes and reservoirs were taken into account and the similarities between lakes and reservoirs were similar between regions were simulated through Bayesian hierarchical model, and the nutrient standard value was derived based on the posterior distribution.
When quantifying the balance between regional commonality and differences, we can accurately and scientifically formulate regional lake and reservoir nutrient benchmarks, providing more effective support for lake and reservoir water environment management.
Smart Images

Figure CN119848026B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water quality monitoring and water body management, and particularly relates to a method for formulating nutrient criteria for regional lakes and reservoirs based on a Bayesian hierarchical model. Background Art
[0002] The phenomenon of eutrophication in lakes and reservoirs is an important challenge in global water environment problems. The increase in the concentrations of nutrients such as nitrogen and phosphorus not only easily leads to the degradation of the water body ecosystem, but also may cause water quality problems such as algal blooms and lack of dissolved oxygen, thereby affecting biodiversity and threatening the safety of water resources. Therefore, establishing scientific and reasonable nutrient criteria, especially their application in the water quality management of regional lakes and reservoirs, has become an important task for water resources management and environmental protection in various countries.
[0003] Currently, the general methods for formulating nutrient criteria for lakes and reservoirs usually adopt the reference state method and the stress-response model. This method often faces the "similarity dilemma" due to the heterogeneity of regional lakes and reservoirs. Due to the differences in natural conditions and human activity intensities among lakes and reservoirs within a region, there are significant spatial variations in the nutrient background values. However, traditional methods rely on a unified reference lake or fixed quantiles, which easily neglect the specificities of sub-regions. As an emerging method, the Bayesian hierarchical model can integrate data from different lakes and reservoirs, not only considering the differences within a single lake but also paying attention to the similarities among lakes and reservoirs in different regions.
[0004] In summary, the present invention proposes a method for formulating nutrient criteria for regional lakes and reservoirs based on a Bayesian hierarchical model, aiming to overcome the limitations in the prior art and more accurately set nutrient standard values suitable for each region, so as to promote the protection and restoration of the lake and reservoir ecosystem. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a method for formulating nutrient criteria for regional lakes and reservoirs based on a Bayesian hierarchical model. This method is directly applicable and easy to use, and has obvious advantages in quantifying the balance between regional commonalities and differences. It can accurately and scientifically formulate nutrient criteria for regional lakes and reservoirs, provide effective support for the water environment management of lakes and reservoirs, and has significant advantages and application values.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for formulating nutrient criteria for regional lakes and reservoirs based on a Bayesian hierarchical model, comprising the following steps:
[0008] S1. Collect monitoring data of different forms of nitrogen and phosphorus indicators of regional lakes and reservoirs, and construct a database for formulating nutrient criteria for regional lakes and reservoirs;
[0009] S2. For a specific form of nutrient, establish a Bayesian hierarchical model for simulating the nutrient concentration in lakes and reservoirs;
[0010] The specific process of step S2 is as follows:
[0011] S21. Perform logarithmic transformation on the concentration values in the monitoring data; among them, for zero-value data, according to the water quality monitoring standard, replace the zero-value data with half of the detection limit;
[0012] S22. Use the data in the database for formulating regional lake and reservoir nutrient criteria to train the Bayesian hierarchical model. The formula of the model is: , where is the th monitoring value of a specific nutrient in the lake or reservoir after logarithmic transformation; ; is the lake or reservoir site; is the order of monitoring values of a specific lake or reservoir; is the regional overall mean of the concentration of a specific form of nutrient; is the deviation of the lake or reservoir relative to the regional overall mean; is the residual of the lake or reservoir at the th monitoring;
[0013] S23. Use the monitoring data of different forms of nitrogen and phosphorus indicators as input variables to simulate the nutrient concentration in the lake or reservoir, and realize the construction of the Bayesian hierarchical model;
[0014] S3. According to the monitoring data in the database for formulating regional lake and reservoir nutrient criteria, perform parameter estimation on the Bayesian hierarchical model to obtain the posterior distribution of each parameter in the model;
[0015] S4. According to the parameter estimation results, obtain the distribution of the regional lake and reservoir nutrient concentration;
[0016] S5. Derive the nutrient standard value according to the quantiles set by the reference state method;
[0017] The specific process of step S5 is as follows:
[0018] S51. Select the reference quantiles from the probability density function of the overall distribution of the nutrient concentration in the regional lakes and reservoirs; among them, select the 25% reference quantile for the areas with pollution greater than the threshold, and select the 75% reference quantile for the areas with protection greater than the threshold;
[0019] S52. Use the posterior distribution of the Bayesian hierarchical model, extract samples from it and calculate the concentration values of the selected quantiles to obtain the nutrient concentration standard values of the regional lakes and reservoirs.
[0020] Preferably, the specific process of step S1 is as follows:
[0021] S11. Collect the monitoring data of different forms of nitrogen and phosphorus indicators of multiple regional lakes and reservoirs;
[0022] S12. Perform outlier processing on the collected monitoring data, and then use the Kalman filtering method to improve the accuracy of state estimation by fusing measurement values and state estimation, and complete the filling of missing values.
[0023] S13. Store and manage the data to construct a database for formulating regional lake and reservoir nutrient benchmarks.
[0024] Preferably, the monitoring data in step S1 includes total nitrogen data, ammonia nitrogen data, nitrate nitrogen data, nitrite nitrogen data, organic nitrogen data, total phosphorus data, dissolved phosphorus data, particulate phosphorus data, and organic phosphorus data.
[0025] Preferably, the outlier processing in step S12 includes removing non-numerical data and duplicate data; the non-numerical data includes characters and null values; the removal of duplicate data is to remove data with repeated timestamps.
[0026] Preferably, the Kalman filtering method in step S12 includes establishing a state space model to describe the concentration change process, setting the covariance matrices of process noise and observation noise, and realizing the dynamic interpolation of missing values through a prediction-update cycle.
[0027] Preferably, the specific process of step S3 is as follows:
[0028] S31. Set the prior distribution for the parameters in the Bayesian hierarchical model. The prior distribution of the parameters is: , , , where and are constants; is a variable and follows a Gamma distribution;
[0029] , where and are prior hyperparameters of known distributions;
[0030] S32. Define the prior distribution as: , where is the th monitoring value of a specific nutrient in the lake or reservoir after logarithmic transformation;
[0031] S33. Use the Markov chain Monte Carlo algorithm for parameter estimation. Through the Markov chain Monte Carlo algorithm, sample from the defined prior distribution to construct a Markov chain, which is used to make the stationary distribution of the Markov chain the posterior distribution of the parameters.
[0032] Preferably, the specific process of step S33 is as follows:
[0033] S331. For each parameter , , and , update the value of the parameter according to the current value and the prior distribution by acceptance-rejection sampling or conditional distribution sampling;
[0034] S332. Generate a large number of samples from the prior distribution through multiple samplings. The goal of the Markov chain Monte Carlo algorithm is to gradually approximate the posterior distribution of the target parameter through continuous iterative sampling;
[0035] S333. Use diagnostic tools to check the convergence of the Markov chain; after multiple samplings until the Markov chain converges, use the posterior distribution for parameter estimation and inference; the diagnostic tools include Gelman-Rubin diagnostics and trace plots.
[0036] Preferably, the specific process of step S4 is as follows:
[0037] S41. Use the posterior distribution samples in the Bayesian hierarchical model to calculate the predicted value of the nutrient concentration of each lake reservoir through the Bayesian hierarchical model formula ;
[0038] S42. For the concentration distribution samples of all lake reservoirs, based on the posterior distribution of the Bayesian method, perform repeated calculations to generate concentration prediction values under multiple posterior samples, which are used to reflect the different performances of nutrient concentrations under different environmental conditions and constitute multi-dimensional samples of the concentration distribution;
[0039] S43. Integrate the concentration distributions obtained from all posterior samples to obtain the overall distribution of the nutrient concentration of the regional lake reservoirs; through the overall distribution of the nutrient concentration of the regional lake reservoirs, reflect the concentration change range and uncertainty of each lake reservoir in the region under different times and environmental conditions, and provide a basis for the formulation of regional nutrient standards.
[0040] After adopting the above technical solution, the present invention has the following beneficial effects: The method for formulating regional lake nutrient benchmarks based on the Bayesian hierarchical model of the present invention is directly applicable and easy to use. During the data quality control stage, through outlier identification and missing value imputation, the impact of abnormal data on the model construction and accuracy evaluation process is reduced; by performing logarithmic transformation on the concentration and replacing zero values with 1 / 2 of the detection limit, the skewed distribution problem in the data can be solved and the rationality of statistical analysis can be ensured; as a probability generation model, the Bayesian hierarchical model is significantly superior to traditional frequentist statistical models (such as linear regression, mixture distribution) and static empirical methods in benchmark formulation, and directly outputs benchmark values in combination with adjustable quantiles (such as 25% or 75%), breaking through the dependence of traditional methods on single distribution assumptions and data independence. Therefore, the method for formulating regional lake nutrient benchmarks based on the Bayesian hierarchical model of the present invention has obvious advantages in quantifying the balance of regional commonalities and differences, can accurately and scientifically formulate regional lake nutrient benchmarks, provide effective support for lake water environment management, and has significant advantages and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of the present invention;
[0042] Figure 2 It is a flow block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] As Figure 1 and Figure 2 shown, a method for formulating regional lake nutrient benchmarks based on the Bayesian hierarchical model includes the following steps:
[0045] S1. Collect monitoring data of different forms of nitrogen and phosphorus indicators in regional lakes and reservoirs, and construct a database for formulating regional lake nutrient benchmarks;
[0046] The monitoring data in step S1 includes total nitrogen data, ammonia nitrogen data, nitrate nitrogen data, nitrite nitrogen data, organic nitrogen data, total phosphorus data, dissolved phosphorus data, particulate phosphorus data, and organic phosphorus data;
[0047] The specific process of step S1 is as follows:
[0048] S11. Collect monitoring data of different forms of nitrogen and phosphorus indicators in multiple regional lakes and reservoirs;
[0049] S12. Process the collected monitoring data for outliers, and then use the Kalman filtering method to improve the accuracy of state estimation by fusing measurement values and state estimates, and complete the filling of missing values.
[0050] The outlier processing described in step S12 includes removing non-numerical data and duplicate data; the non-numerical data includes characters and null values; the removal of duplicate data is to remove the data with duplicate timestamps.
[0051] The Kalman filtering method described in step S12 includes establishing a state space model to describe the concentration change process, setting the covariance matrices of process noise and observation noise, and realizing the dynamic imputation of missing values through a prediction-update loop.
[0052] S13. Store and manage the data to construct a database for formulating regional lake and reservoir nutrient benchmarks.
[0053] S2. For specific forms of nutrients, establish a Bayesian hierarchical model to simulate the nutrient concentrations in lakes and reservoirs.
[0054] The specific process of step S2 is as follows:
[0055] S21. Perform a logarithmic transformation on the concentration values in the monitoring data; among them, for zero-value data, according to the water quality monitoring standards, replace the zero-value data with half of the detection limit.
[0056] S22. Use the data in the database for formulating regional lake and reservoir nutrient benchmarks to train the Bayesian hierarchical model. The formula of the model is: , where is the th monitoring value of the specific nutrient after logarithmic transformation in the lake or reservoir ; is the lake or reservoir location; is the monitoring value sequence of the specific lake or reservoir; is the regional overall mean of the concentration of the specific form of nutrient; is the deviation of the lake or reservoir relative to the regional overall mean; is the residual of the lake or reservoir at the th monitoring;
[0057] S23. Use the monitoring data of different forms of nitrogen and phosphorus indicators as input variables to simulate the nutrient concentrations in lakes and reservoirs, and realize the construction of the Bayesian hierarchical model.
[0058] S3. According to the monitoring data in the database for formulating regional lake and reservoir nutrient benchmarks, perform parameter estimation on the Bayesian hierarchical model to obtain the posterior distribution of each parameter in the model.
[0059] The specific process of step S3 is as follows:
[0060] S31. Set the prior distribution for the parameters in the Bayesian hierarchical model. The prior distribution of the parameters is as follows: , , , where and are constants; is a variable and follows a Gamma distribution;
[0061] , where and are the prior hyperparameters of the known distribution;
[0062] S32. The defined prior distribution is: , where is the logarithmically transformed value of the specific nutrient in the lake / reservoir at the th monitoring;
[0063] S33. Use the Markov chain Monte Carlo algorithm for parameter estimation. Through the Markov chain Monte Carlo algorithm, sample from the defined prior distribution to construct a Markov chain, so that the stationary distribution of the Markov chain is the posterior distribution of the parameters;
[0064] The specific process of step S33 is as follows:
[0065] S331. For each parameter , , and , update the value of the parameter according to the current value and the prior distribution through acceptance-rejection sampling or conditional distribution sampling;
[0066] S332. Generate a large number of samples from the prior distribution through multiple samplings. The goal of the Markov chain Monte Carlo algorithm is to gradually approximate the posterior distribution of the target parameter through continuous iterative sampling;
[0067] S333. Use diagnostic tools to check the convergence of the Markov chain; after multiple samplings until the Markov chain converges, use the posterior distribution for parameter estimation and inference; the diagnostic tools include Gelman-Rubin diagnostics and trace plots;
[0068] S4. Obtain the distribution of nutrient concentrations in regional lakes / reservoirs according to the parameter estimation results;
[0069] The specific process of step S4 is as follows:
[0070] S41. Use the posterior distribution samples in the Bayesian hierarchical model to calculate the nutrient concentration of each lake / reservoir through the Bayesian hierarchical model formula Predicted value;
[0071] S42. For the concentration distribution samples of all lakes and reservoirs, based on the posterior distribution of the Bayesian method, perform repeated calculations to generate concentration predicted values under multiple posterior samples, which are used to reflect the different performances of nutrient concentrations under different environmental conditions and constitute multi-dimensional samples of the concentration distribution;
[0072] S43. Integrate the concentration distributions obtained from all posterior samples to obtain the overall distribution of nutrient concentrations in regional lakes and reservoirs; through the overall distribution of nutrient concentrations in regional lakes and reservoirs, reflect the concentration change ranges and uncertainties of each lake and reservoir in the region under different times and environmental conditions, and provide a basis for the formulation of regional nutrient standards;
[0073] S5. Deduce the nutrient standard value according to the quantiles set by the reference state method;
[0074] The specific process of step S5 is as follows:
[0075] S51. Select reference quantiles from the probability density function of the overall distribution of nutrient concentrations in regional lakes and reservoirs; among them, select the 25% reference quantile for the area where the pollution is greater than the threshold, and select the 75% reference quantile for the area where the protection is greater than the threshold;
[0076] S52. Use the posterior distribution of the Bayesian hierarchical model to extract samples therefrom and calculate the concentration values of the selected quantiles to obtain the nutrient concentration standard values of regional lakes and reservoirs.
[0077] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for establishing a regional lake nutrient benchmark based on a Bayesian hierarchical model, characterized in that: The following steps are involved: S1. Collect monitoring data on different forms of nitrogen and phosphorus indicators in regional lakes and reservoirs, and build a database for the formulation of nutrient salt benchmarks in regional lakes and reservoirs; S2. For specific forms of nutrients, a Bayesian hierarchical model is established to simulate the nutrient concentration in lakes and reservoirs; The specific process of step S2 is: S21, performing logarithmic transformation on the concentration values in the monitoring data; wherein, for zero-value data, the zero-value data is replaced with half of the detection limit according to the standard of water quality monitoring; S22. Use the regional lake nutrient benchmark database data to train the Bayesian hierarchical model. The model formula is: ,in, is the specific nutrient in lakes and reservoirs after logarithmic transformation No. Secondary monitoring value; It is the location of lakes and reservoirs; The sequence of monitoring values for a particular lake or reservoir; is the regional overall mean of the concentration of a specific form of nutrient salt; For lakes and reservoirs Deviation from the overall mean of the region; For lakes and reservoirs In the The residual error of the second monitoring; S23. Use monitoring data of different forms of nitrogen and phosphorus indicators as input variables to simulate the nutrient concentration of lakes and reservoirs and realize the construction of Bayesian hierarchical model; S3. According to the monitoring data in the database of regional lake nutrient salt benchmark, the parameters of the Bayesian hierarchical model are estimated to obtain the posterior distribution of each parameter in the model; S4. Based on the parameter estimation results, the distribution of nutrient concentrations in regional lakes and reservoirs is obtained; S5. Derive the standard value of nutrient salts according to the quantiles set by the reference state method; The specific process of step S5 is: S51. Select reference quantiles from the probability density function of the overall distribution of nutrient concentrations in regional lakes and reservoirs; where the 25% reference quantile is selected for areas with pollution greater than the threshold, and the 75% reference quantile is selected for areas with protection greater than the threshold; S52. Using the posterior distribution of the Bayesian hierarchical model, samples are extracted and the concentration values of the selected quantiles are calculated to obtain the standard values of nutrient concentrations in regional lakes and reservoirs.
2. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 1, characterized in that: The specific process of step S1 is: S11. Collect monitoring data on different forms of nitrogen and phosphorus indicators in lakes and reservoirs in multiple regions; S12, processing the outliers of the collected monitoring data, and then using the Kalman filter method to improve the accuracy of the state estimation by fusing the measured value and the state estimation, and completing the missing value filling; S13. Store and manage data to build a regional lake and reservoir nutrient benchmark database.
3. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 1, characterized in that: The monitoring data in step S1 include total nitrogen data, ammonia nitrogen data, nitrate nitrogen data, nitrite nitrogen data, organic nitrogen data, total phosphorus data, dissolved phosphorus data, particulate phosphorus data and organic phosphorus data.
4. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 2, characterized in that: The abnormal value processing in step S12 includes eliminating non-numeric data and eliminating duplicate data; the non-numeric data includes characters and null values; and the elimination of duplicate data is to eliminate data with duplicate timestamps.
5. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 2, characterized in that: The Kalman filtering method described in step S12 includes establishing a state space model to describe the concentration change process, setting the covariance matrix of process noise and observation noise, and realizing dynamic interpolation of missing values through a prediction-update cycle.
6. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 1, characterized in that: The specific process of step S3 is: S31. Set the prior distribution for the parameters in the Bayesian hierarchical model. The prior distribution of the parameters is: , , ,in, and is a constant; is a variable and follows the Gamma distribution; ,in, and is a priori hyperparameter of a known distribution; S32. The prior distribution defined is: ,in, is the specific nutrient in lakes and reservoirs after logarithmic transformation No. Secondary monitoring value; S33. Using the Markov Chain Monte Carlo algorithm to estimate parameters, sampling is performed from a defined prior distribution through the Markov Chain Monte Carlo algorithm to construct a Markov chain, so as to make the stationary distribution of the Markov chain the posterior distribution of the parameters.
7. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 6, characterized in that: The specific process of step S33 is: S331. For each parameter , , and , update the parameter value by acceptance-rejection sampling or conditional distribution sampling according to the current value and prior distribution; S332. Through multiple sampling, a large number of samples are generated from the prior distribution. The goal of the Markov Chain Monte Carlo algorithm is to gradually approach the posterior distribution of the target parameter through continuous iterative sampling; S333. Use diagnostic tools to check the convergence of the Markov chain; perform multiple sampling until the Markov chain converges, and use the posterior distribution to perform parameter estimation and inference; the diagnostic tools include Gelman-Rubin diagnosis and trajectory diagram.
8. A method for establishing a regional lake nutrient salt benchmark based on a Bayesian hierarchical model as claimed in claim 1, characterized in that: The specific process of step S4 is: S41. Using the posterior distribution samples in the Bayesian hierarchical model, the nutrient concentration of each lake is calculated using the Bayesian hierarchical model formula. The predicted value of S42. For the concentration distribution samples of all lakes and reservoirs, repeated calculations are performed based on the posterior distribution of the Bayesian method to generate concentration prediction values under multiple posterior samples, which are used to reflect the different performances of nutrient concentrations under different environmental conditions and constitute multi-dimensional samples of concentration distribution; S43. Integrate the concentration distributions obtained from all posterior samples to obtain the overall distribution of nutrient concentrations in regional lakes and reservoirs. The overall distribution of nutrient concentrations in regional lakes and reservoirs can be used to reflect the concentration variation range and uncertainty of each lake and reservoir in the region under different time and environmental conditions, providing a basis for the formulation of regional nutrient standards.
Citation Information
Patent Citations
Slope risk assessment method based on Bayesian hierarchical space-time model
CN116227162A
KR20210048948A