Regional lake and reservoir water quality uncertainty simulation method based on Bayesian transfer learning

Through Bayesian transfer learning and uncertainty quantification methods, the problems of missing and uncertainty in lake reservoir water quality monitoring data are solved, accurate water quality simulation and key factor identification are achieved, and the scientificity and reliability of water quality management are improved.

CN120388638AActive Publication Date: 2025-07-29XIAMEN UNIV +1

Patent Information

Application Number
CN202510874396.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with the data loss and uncertainty of lake reservoir water quality monitoring data, resulting in insufficient accuracy and reliability of water quality assessment and management.

Method used

Using a Bayesian transfer learning method, through time series decomposition, outlier recognition and missing value interpolation, combined with Bayesian additive regression tree and cross-validation strategy, a water quality simulation model across lake reservoirs was constructed, and uncertainty analysis was performed to identify key factors.

Benefits of technology

The accuracy and reliability of lake reservoir water quality simulation and prediction have been improved, the key factors affecting water quality have been identified, and the scientific basis for lake reservoir water quality management has been provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388638A_ABST
    Figure CN120388638A_ABST
Patent Text Reader

Abstract

The invention discloses a regional lake and reservoir water quality uncertainty simulation method based on Bayesian transfer learning, and the method comprises the following steps: S1, collecting water quality automatic monitoring data of a plurality of lakes and reservoirs in a region, and carrying out the preprocessing of the monitoring data; s2, numbering the plurality of lakes and reservoirs respectively, dividing the lakes and reservoirs into a target domain and a source domain, and selecting a water quality index needing to be simulated as a target index; s3, constructing a target index simulation model in the source domain by adopting a Bayesian additive regression tree, and performing model training by adopting a cross validation strategy; s4, performing parameter updating on the target index simulation model in the source domain by adopting the data of the target domain, constructing a Bayesian migration model of the target index of the target domain, and performing uncertainty simulation on the water quality target index of the target domain; and S5, according to an output result of the Bayesian migration model, identifying a key factor influencing the mean value and uncertainty of the water quality target indexes of the target domain, and analyzing a driving mechanism of the key factor to fluctuation of the water quality indexes of the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of water quality monitoring and water body management, and particularly relates to a method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning. Background Art

[0002] With the rapid development of industrialization and urbanization, the water quality problems of lakes and reservoirs have attracted increasing attention. The accuracy and integrity of water quality monitoring data are crucial for water quality assessment and management, as these data are the basis for formulating scientific and reasonable treatment measures. However, water quality monitoring data are often affected by factors such as monitoring instrument failures and environmental interferences, resulting in data missing or anomalies. In this case, transfer learning can be used as an effective solution. Through transfer learning, the existing relevant domain knowledge can be used to fill in the missing parts of the data, thereby improving the generalization ability and prediction accuracy of the model.

[0003] In addition, the simulation and prediction of lake and reservoir water quality indicators usually need to consider the comprehensive effects of multiple relevant factors, but traditional methods are often difficult to effectively handle the uncertainty and complexity of data. To address this challenge, a method capable of quantifying model uncertainty is needed to more accurately evaluate the reliability of model predictions by quantifying uncertainty and provide more scientific decision-making support for water quality management. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning. This method combines transfer learning and uncertainty quantification methods, which can not only effectively handle data missing problems, but also achieve accurate simulation and uncertainty analysis of lake and reservoir water quality indicators, and identify the key factors affecting water quality, improving the accuracy and reliability of water quality simulation and prediction, thereby providing strong support for the sustainable management of lake and reservoir water quality.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning, comprising the following steps: S1. Collect the automatic monitoring data of the water quality of multiple lakes and reservoirs in the region, and preprocess the monitoring data; S2. Number the multiple lakes and reservoirs respectively, divide the lakes and reservoirs into a target domain and a source domain, and select the water quality indicator to be simulated as the target indicator; S3. Use a Bayesian additive regression tree to construct a target indicator simulation model in the source domain, and use a cross-validation strategy for model training; S4. Select the variable parameters of the target index simulation model in the source domain, update the parameters of the target index simulation model in the source domain using the data in the target domain, construct the Bayesian transfer model of the target index in the target domain, and conduct uncertainty simulation of the water quality target index in the target domain; S5. According to the output results of the Bayesian transfer model, give the mean and confidence interval of the predicted value, identify the key factors affecting the mean and uncertainty of the water quality target index in the target domain, and analyze the driving mechanism of the key factors on the fluctuation of the water quality index in the target domain.

[0006] Preferably, the monitoring data in step S1 includes water temperature, pH, dissolved oxygen, ammonia nitrogen, total nitrogen, total phosphorus, and chlorophyll a concentration.

[0007] Preferably, the specific process of step S1 is as follows: S11. Collect the automatic water quality monitoring data of multiple lakes and reservoirs in the region, obtain the time series data of each monitoring index; then decompose the time series data using the STL method, and split the time series data into a long-term trend term, a seasonal cycle term, and a residual term; S12. For the decomposed residual term, use the generalized extreme studentized residual method to conduct outlier tests, identify the abnormal data, and mark the abnormal data as missing values; S13. Use the Kalman filter method to fill in the missing values, and gradually improve the accuracy of the state estimate by continuously fusing the measurement values and the state estimate to complete the filling of the missing values.

[0008] Preferably, in step S13, the Kalman filter method includes establishing a state space model to describe the concentration change process, setting the covariance matrices of the process noise and the observation noise, and realizing the dynamic interpolation of the missing values through a prediction-update cycle.

[0009] Preferably, the specific process of step S2 is as follows: S21. Assign a unique number from 1 to N to each lake and reservoir in the region, use some lakes and reservoirs as the target domain for individual water quality simulation; then use the remaining lakes and reservoirs as the source dataset for model training; S22. Select the water quality indicators to be simulated, and ensure that the selected water quality indicators are recorded in the monitoring data of all lakes and reservoirs.

[0010] Preferably, the specific process of step S3 is as follows: S31. Use the data of the source domain lakes and reservoirs to construct a simulation dataset of the target index; among them, the target index is used as the response variable, and other water quality indicators are used as the prediction variables; S32. Construct a target metric simulation model for predicting the target metric using Bayesian additive regression trees; among them, the Bayesian additive regression tree can capture complex non-linear relationships and provide stable prediction results by constructing Bayesian models of multiple additive regression trees; The specific process of step S32 is as follows: S321. The mathematical form of the target metric simulation model is: where, is the predicted target metric; is the index of the th regression tree; is the number of regression trees; are the features or explanatory variables used to predict the target metric; is the th prediction function of the binary regression tree; is the error term, representing the difference between the model prediction value and the actual observed value; represents the error term follows a normal distribution with a mean of 0 and a variance of ; is the standard deviation of the error term, used to measure the dispersion degree between the model prediction value and the actual observed value; S322. Each tree contains the splitting rules of internal nodes and the leaf node parameters , and each tree is made into a weak learner through regularization prior constraints: where, is to control the tree depth when, the probability that a node becomes a terminal node; is to control the tree depth; and are both hyperparameters for controlling the tree depth ; S323. The Bayesian additive regression tree uses Markov chain Monte Carlo for posterior inference, which is used to iteratively update each tree, and at the same time perform conditional updates based on other trees, and allows generating samples from the posterior distribution; the regression function at a specific value, the posterior mean is estimated by averaging all Markov chain Monte Carlo samples, and the calculation formula is: where, is the posterior mean estimate of the regression function at a specific value; is the total number of Markov chain Monte Carlo samples; is the index of the th Markov chain Monte Carlo sample; is the th Markov chain Monte Carlo sample at The value of the additive tree model evaluated at a value; S324. Regression function The confidence interval of is constructed from the quantiles of the posterior samples, and the calculation formula is: , where is the regression function at a specific value; represents an interval; is the regression function quantile function of the posterior samples of; S33. Use K-fold cross-validation to train and evaluate the Bayesian additive regression tree to improve the generalization ability of the target metric simulation model.

[0011] Preferably, in step S33, the K-fold cross-validation includes dividing the source domain data into K parts, where K - 1 parts are used for training and the remaining 1 part is used for validation. It is cycled K times, and the prediction errors of the target metric simulation model on different data subsets are calculated to improve the prediction stability and reliability of the target metric simulation model for unseen data.

[0012] Preferably, the specific process of step S4 is: S41. Analyze the target metric simulation model constructed in the source domain and determine the variable parameters; among them, the variable parameters include the depth of the decision tree and the node splitting rule; S42. Use the data in the target domain to adjust and optimize the variable parameters of the target metric simulation model in the selected source domain through the Bayesian inference method, so that the target metric simulation model in the source domain can adapt to the actual situation in the target domain; S43. After the parameter update, obtain the Bayesian transfer model applicable to the target domain; among them, the Bayesian transfer model integrates the prior knowledge of the target metric simulation model in the source domain and the data information in the target domain.

[0013] Preferably, the specific process of step S5 is: S51. Calculate the predicted mean and confidence interval of the water quality target metrics in the target domain according to the output results of the Bayesian transfer model; S52. Use the feature importance index output by the Bayesian transfer model to identify the variables that have the greatest impact on the mean and uncertainty of the water quality target metrics in the target domain, and use the variables with the greatest impact on uncertainty as the key factors; S53. For the selected key factors, quantify the response relationship between the key factors and the target metrics, and analyze the driving mechanism of the fluctuations of the water quality target metrics in the target domain.

[0014] Preferably, the specific process of step S52 is: S521. Calculate the initial root mean square error (RMSE) of the Bayesian transfer model on the original test set, and use the initial RMSE as the benchmark metric. S522. Randomly shuffle the values of each feature, re-predict with the perturbed data, and calculate the perturbed RMSE. Use the perturbed RMSE as the perturbed metric. S523. Measure the feature importance using the difference between the benchmark metric and the perturbed metric. The larger the difference, the more critical the feature is for model prediction.

[0015] After adopting the above technical solutions, the present invention has the following beneficial effects: By combining transfer learning and uncertainty quantification methods, the present invention can not only effectively handle the problem of missing data, but also achieve accurate simulation and uncertainty analysis of lake and reservoir water quality indicators, identify key factors affecting water quality, improve the accuracy and reliability of water quality simulation and prediction, and thus provide strong support for the sustainable management of lake and reservoir water quality. Specifically, the regional lake and reservoir water quality uncertainty simulation method based on Bayesian transfer learning of the present invention is directly applicable and easy to use. In the data preprocessing stage, through time series decomposition, outlier identification, and missing value imputation, it effectively solves the data quality problem, reduces the interference of abnormal data on model construction and accuracy evaluation, and ensures the integrity and reliability of the data. In terms of model construction, using Bayesian additive regression trees (BART) combined with cross-validation strategies can accurately capture the complex non-linear relationships of water quality indicators, provide stable prediction results, and significantly outperform traditional statistical models and empirical methods. The introduction of Bayesian transfer learning realizes knowledge transfer and model optimization across lakes and reservoirs, breaks through the dependence of traditional methods on a single data set, and improves the generalization ability and adaptability of the model. In addition, this method can identify key factors and analyze their driving mechanisms, providing a scientific basis for water quality management. Therefore, the present invention has obvious advantages in quantifying the commonalities and differences of regional lakes and reservoirs, can accurately and scientifically simulate water quality uncertainty, provides solid theoretical and technical support for regional lake and reservoir water environment management and regulation, and has important theoretical significance and broad application prospects. Description of the Drawings

[0016] Figure 1 It is a flow chart of the present invention. Detailed Embodiments

[0017] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] As Figure 1 shown, the regional lake and reservoir water quality uncertainty simulation method based on Bayesian transfer learning includes the following steps: S1. Collect the automatic water quality monitoring data of multiple lakes and reservoirs in the area, and preprocess the monitoring data; The monitoring data described in step S1 includes water temperature, pH, dissolved oxygen, ammonia nitrogen, total nitrogen, total phosphorus, and chlorophyll a concentration; The specific process of step S1 is as follows: S11. Collect the automatic water quality monitoring data of multiple lakes and reservoirs in the area to obtain the time series data of each monitoring index; then use the STL method to decompose the time series data, and split the time series data into a long-term trend term, a seasonal cycle term, and a residual term; S12. For the residual term obtained by decomposition, use the generalized extreme studentized residual method for outlier detection, identify the abnormal data, and mark the abnormal data as missing values; S13. Use the Kalman filter method to fill in the missing values. By continuously fusing the measured values and the state estimates, gradually improve the accuracy of the state estimates and complete the filling of the missing values; In step S13, the Kalman filter method includes establishing a state space model to describe the concentration change process, setting the covariance matrices of the process noise and the observation noise, and realizing the dynamic interpolation of the missing values through the prediction-update cycle; S2. Number the multiple lakes and reservoirs separately, divide the lakes and reservoirs into a target domain and a source domain, and select the water quality index to be simulated as the target index; The specific process of step S2 is as follows: S21. Assign unique numbers from 1 to N to each lake and reservoir in the area. Take some lakes and reservoirs as the target domain for individual water quality simulation; then take the remaining lakes and reservoirs as the source domain for the source dataset of model training; S22. Select the water quality index to be simulated and ensure that the selected water quality index is recorded in the monitoring data of all lakes and reservoirs; S3. Use the Bayesian additive regression tree to construct a target index simulation model in the source domain, and use the cross-validation strategy for model training; The specific process of step S3 is as follows: S31. Use the data of the lakes and reservoirs in the source domain to construct a simulation dataset of the target index; among them, the target index is used as the response variable, and other water quality indexes are used as the predictor variables; S32. Use the Bayesian additive regression tree to construct a target index simulation model for predicting the target index; among them, the Bayesian additive regression tree can capture complex non-linear relationships and provide stable prediction results by constructing a Bayesian model of multiple additive regression trees; The specific process of step S32 is as follows: S321. The mathematical form of the target index simulation model is: , where is the predicted target index; is the index of the th regression tree; is the number of regression trees; are the features or explanatory variables used to predict the target metric; is the th prediction function of the binary regression tree; where the specific process of constructing the prediction function is: for each tree, a binary decision tree is constructed by fitting on the training data set, specifically including selecting the feature dimension and its division threshold as the splitting rule for the internal nodes, and fitting the local target value at each leaf node; is the error term, representing the difference between the model predicted value and the actual observed value; represents the error term follows a normal distribution with mean 0 and variance ; is the standard deviation of the error term, used to measure the dispersion between the model predicted value and the actual observed value; S322. Each tree contains the splitting rules of the internal nodes and the leaf node parameters , and each tree is made into a weak learner through regularization prior constraints: , where is to control the tree depth when, the probability that a node becomes a terminal node; is to control the tree depth; and are both hyperparameters for controlling the tree depth ; S323. The Bayesian additive regression tree uses Markov chain Monte Carlo for posterior inference, which is used to iteratively update each tree, and at the same time, conditional updates are made based on other trees, and it allows generating samples from the posterior distribution; the regression function at a specific value is estimated by averaging all Markov chain Monte Carlo samples, and the calculation formula is: , where is the regression function at a specific value of the posterior mean estimate; is the total number of Markov chain Monte Carlo samples; is the th index of the Markov chain Monte Carlo sample; is the th value of the additive tree model evaluated by the Markov chain Monte Carlo sample at the value; where the specific process of obtaining the value of the additive tree model is: take Values are input into each tree of the sample, reaching the leaf nodes along the splitting path, taking the values of the leaf nodes, and summing up the output values of all the trees to obtain the overall predicted value under this iteration, which is the value of the additive tree model; S324. Regression function The confidence interval of is constructed from the quantiles of the posterior samples, and the calculation formula is: where, is the regression function at a specific value; represents the interval; is the regression function of the quantile function of the posterior samples; S33. Use K-fold cross-validation to train and evaluate the Bayesian additive regression tree to improve the generalization ability of the target index simulation model; In step S33, the K-fold cross-validation includes dividing the source domain data into K parts, where K - 1 parts are used for training and the remaining 1 part is used for validation. This is done in K rounds, and the prediction errors of the target index simulation model on different data subsets are calculated to improve the prediction stability and reliability of the target index simulation model for unseen data; S4. Select the variable parameters of the target index simulation model in the source domain, use the data in the target domain to update the parameters of the target index simulation model in the source domain, construct the Bayesian transfer model of the target domain target index, and perform uncertainty simulation of the target domain water quality target index; The specific process of step S4 is as follows: S41. Analyze the target index simulation model constructed in the source domain and determine the variable parameters; among them, the variable parameters include the depth of the decision tree and the node splitting rule; S42. Use the data in the target domain to adjust and optimize the variable parameters of the target index simulation model selected in the source domain through the Bayesian inference method, so that the target index simulation model in the source domain can adapt to the actual situation in the target domain; S43. After parameter update, obtain the Bayesian transfer model applicable to the target domain; among them, the Bayesian transfer model integrates the prior knowledge of the target index simulation model in the source domain and the data information in the target domain; S5. According to the output results of the Bayesian transfer model, give the mean and confidence interval of the predicted values, identify the key factors affecting the mean and uncertainty of the target domain water quality target index, and analyze the driving mechanism of the key factors on the fluctuations of the target domain water quality index; The specific process of step S5 is as follows: S51. According to the output results of the Bayesian transfer model, calculate the predicted mean and confidence interval of the target domain water quality target index; S52. Identify the variables that have the greatest impact on the mean and uncertainty of the water quality target indicators in the target domain by using the feature importance indicators output by the Bayesian transfer model, and use the variable with the greatest impact on uncertainty as the key factor; The specific process of step S52 is as follows: S521. Calculate the initial root mean square error of the Bayesian transfer model on the original test set, and use the initial root mean square error as the benchmark indicator; S522. Randomly shuffle the values of each feature, re-predict with the perturbed data, and calculate the perturbed root mean square error, using the perturbed root mean square error as the perturbed indicator; S523. Measure the feature importance by the difference between the benchmark indicator and the perturbed indicator. The larger the difference, the more critical the feature is to the model prediction; S53. Quantify the response relationship between the selected key factor and the target indicator, and analyze the driving mechanism of the fluctuation of the water quality target indicator in the target domain.

[0019] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning, characterized in that It includes the following steps: S1. Collect the automatic water quality monitoring data of multiple lakes and reservoirs in the area, and preprocess the monitoring data; S2. Number the multiple lakes and reservoirs respectively, divide the lakes and reservoirs into a target domain and a source domain, and select the water quality indicators to be simulated as target indicators; S3. Use Bayesian additive regression trees to construct a target indicator simulation model in the source domain, and use a cross-validation strategy for model training; S4. Select the variable parameters of the target indicator simulation model in the source domain, update the parameters of the target indicator simulation model in the source domain with the data of the target domain, construct a Bayesian transfer model for the target indicator of the target domain, and perform uncertainty simulation of the water quality target indicators in the target domain; S5. According to the output results of the Bayesian transfer model, give the mean and confidence interval of the predicted values, identify the key factors affecting the mean and uncertainty of the water quality target indicators in the target domain, and analyze the driving mechanism of the key factors on the fluctuations of the water quality indicators in the target domain.

2. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, wherein: The monitoring data described in step S1 includes water temperature, pH, dissolved oxygen, ammonia nitrogen, total nitrogen, total phosphorus, and chlorophyll a concentration.

3. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, characterized in that, The specific process of step S1 is: S11. Collect the automatic water quality monitoring data of multiple lakes and reservoirs in the area, and obtain the time series data of each monitoring indicator; then use the STL method to decompose the time series data, and split the time series data into a long-term trend term, a seasonal cycle term, and a residual term; S12. For the residual term obtained by decomposition, use the generalized extreme studentized residual method for outlier detection, identify the abnormal data, and mark the abnormal data as missing values; S13. Use the Kalman filter method to fill in the missing values, and gradually improve the accuracy of the state estimate by continuously fusing the measurement values and the state estimate to complete the filling of the missing values.

4. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 3, characterized in that: In step S13, the Kalman filter method includes establishing a state space model to describe the concentration change process, setting the covariance matrices of the process noise and the observation noise, and realizing the dynamic interpolation of the missing values through a prediction-update cycle.

5. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, wherein, The specific process of step S2 is: S21. Assign a unique number from 1 to N to each lake and reservoir in the area, use some lakes and reservoirs as the target domain for individual water quality simulation; then use the remaining lakes and reservoirs as the source domain for the source dataset of model training; S22. Select the water quality indicators to be simulated, and ensure that the selected water quality indicators are recorded in the monitoring data of all lakes and reservoirs.

6. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, characterized in that, The specific process of step S3 is: S31. Use the data of the lakes and reservoirs in the source domain to construct a simulation dataset for the target indicator; among them, the target indicator is used as the response variable, and other water quality indicators are used as the prediction variables; S32. Use Bayesian additive regression trees to construct a target indicator simulation model for predicting the target indicator; among them, the Bayesian additive regression tree can capture complex non-linear relationships and provide stable prediction results by constructing a Bayesian model of multiple additive regression trees; The specific process of step S32 is: S321. The mathematical form of the target index simulation model is as follows: , where is the predicted target index; is the index of the -th regression tree; is the number of regression trees; is the feature or explanatory variable used to predict the target index; is the -th prediction function of the binary regression tree; is the error term, representing the difference between the model prediction value and the actual observed value; indicates the error term follows a normal distribution with a mean of 0 and a variance of ; is the standard deviation of the error term, used to measure the degree of dispersion between the model prediction value and the actual observed value. S322. Each tree contains the splitting rules of internal nodes and the leaf node parameters , and each tree is made into a weak learner through regularization prior constraints: , where is the probability that a node becomes a terminal node when controlling the tree depth ; is used to control the tree depth; and are both hyperparameters for controlling the tree depth ; S323. The Bayesian Additive Regression Tree uses Markov Chain Monte Carlo for posterior inference to iteratively update each tree, conditionally update based on other trees, and allow generating samples from the posterior distribution; the regression function at a specific value is estimated by averaging all Markov Chain Monte Carlo samples, and the calculation formula is: , where is the estimated posterior mean of the regression function at a specific value; is the total number of Markov Chain Monte Carlo samples; is the index of the th Markov Chain Monte Carlo sample; is the value of the additive tree model evaluated at the th Markov Chain Monte Carlo sample at the value. S324. Regression function The confidence interval is constructed from the quantiles of the posterior samples, and the calculation formula is: where is the confidence interval of the regression function at a specific value; represents the interval; is the quantile function of the posterior samples of the regression function ; S33. Use K-fold cross-validation to train and evaluate the Bayesian additive regression tree to improve the generalization ability of the target indicator simulation model.

7. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 6, wherein: In step S33, the K-fold cross-validation includes dividing the source domain data into K parts, where K - 1 parts are used for training and the remaining 1 part is used for validation. This is carried out in K rounds, and the prediction errors of the target metric simulation model on different data subsets are calculated to improve the prediction stability and reliability of the target metric simulation model for unseen data.

8. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, characterized in that The specific process of step S4 is as follows: S41. Analyze the target metric simulation model constructed for the source domain and determine the variable parameters. Among them, the variable parameters include the depth of the decision tree and the node splitting rule. S42. Use the data of the target domain to adjust and optimize the variable parameters of the target metric simulation model in the selected source domain through the Bayesian inference method, so that the target metric simulation model in the source domain can adapt to the actual situation of the target domain. S43. After parameter update, a Bayesian transfer model applicable to the target domain is obtained. Among them, the Bayesian transfer model integrates the prior knowledge of the target metric simulation model in the source domain and the data information of the target domain.

9. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 1, characterized in that, The specific process of step S5 is as follows: S51. Calculate the predicted mean and confidence interval of the water quality target metric of the target domain according to the output result of the Bayesian transfer model. S52. Use the feature importance index output by the Bayesian transfer model to identify the variables that have the greatest impact on the mean and uncertainty of the water quality target metric of the target domain, and regard the variable with the greatest uncertainty impact as the key factor. S53. For the selected key factors, quantify the response relationship between the key factors and the target metric, and analyze the driving mechanism of the fluctuation of the water quality target metric in the target domain.

10. The method for simulating the uncertainty of regional lake and reservoir water quality based on Bayesian transfer learning according to claim 9, wherein, The specific process of step S52 is as follows: S521. Calculate the initial root mean square error of the Bayesian transfer model on the original test set, and use the initial root mean square error as the benchmark index. S522. Randomly shuffle the values of each feature, re-predict with the perturbed data and calculate the perturbed root mean square error, and use the perturbed root mean square error as the perturbed index. S523. Use the difference between the benchmark index and the perturbed index to measure the feature importance. The larger the difference, the more critical the feature is to the model prediction.

Citation Information

Patent Citations

  • Water body carbon emission long-term trend prediction method based on Bayesian additive regression tree

    CN119168152A

  • Method for inverting water quality lacking remote sensing image by combining migration and space-time deep learning

    CN119625551A

  • Regional lake and reservoir nutritive salt reference setting method based on Bayesian hierarchical model

    CN119848026A

  • Lake and reservoir nutritive salt concentration simulation method based on coupling mechanism model and deep learning

    CN119962408A

  • Method and system of sudden water pollutant source detection by forward-inverse coupling

    US20220358266A1

Cited By

  • Multi-periodic ecological environment index anomaly identification and deletion interpolation method

    CN120724360A