Water quality prediction method and device under incomplete information, equipment and medium
By integrating multi-source water quality data and using data assimilation and Bayesian inversion mechanisms to dynamically correct the model state and parameters, the problem of insufficient water quality prediction accuracy under incomplete information is solved, achieving high-precision and stable water quality prediction, and supporting water environment management and pollution prevention and control decisions.
Patent Information
- Application Number
- CN202511799386.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-17
AI Technical Summary
In water environment forecasting, incomplete information makes it difficult for existing technologies to accurately capture the patterns of water quality changes, resulting in insufficient forecast accuracy and poor stability, and failing to provide reliable decision support for water environment management.
By integrating multi-source water quality data, performing standardized preprocessing, combining the random generation and optimal selection of multiple network structures, and using data assimilation and Bayesian inversion mechanisms to dynamically correct the model state and parameters, a water quality prediction model adapted to scenarios with missing data is constructed.
It significantly improves the accuracy and stability of water quality prediction under incomplete information conditions, provides reliable data support, provides a reliable basis for water environment management and pollution prevention and control decisions, and enhances the practicality and adaptability of technology applications.
Smart Images

Figure CN121684152A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of water environment prediction technology, and in particular to a method, apparatus, equipment and medium for water quality prediction under incomplete information. Background Technology
[0002] In water environment prediction scenarios, incomplete information manifests itself in the spatiotemporal gaps of monitoring data. This includes data gaps due to malfunctions at some monitoring stations, insufficient spatial observations caused by the lack of monitoring equipment at key locations, and incomplete observations of key influencing factors such as tributary inputs and gate scheduling parameters. Many water quality prediction methods rely on modeling with complete datasets, making it difficult to accurately capture patterns of water quality changes, resulting in insufficient prediction accuracy and an inability to provide reliable decision support for water environment management. Summary of the Invention
[0003] This application provides a method, apparatus, equipment, and medium for water quality prediction under incomplete information, in order to solve the problems of low accuracy and poor stability of water quality prediction methods under incomplete information in related technologies.
[0004] The first aspect of this application provides a method for water quality prediction under incomplete information, comprising the following steps: acquiring data and monitoring data from water quality monitoring stations within a target area; preprocessing the data and monitoring data, generating a dataset based on the preprocessed multi-source data, training a pre-constructed water quality prediction model using the dataset, randomly generating multiple network structures of the preset water quality model during the training process, determining the target network structure from the multiple network structures based on the training structure, and correcting the model state and model parameters of the water quality prediction model based on data assimilation and Bayesian inversion when some monitoring data is missing; and using the trained water quality prediction model to predict the water quality of the target area.
[0005] Based on the aforementioned technical means, this application embodiment integrates multi-source water quality related data of the target area and performs standardized preprocessing. It combines the random generation and optimal screening of multiple network structures, and utilizes data assimilation and Bayesian inversion mechanisms to dynamically correct the model state and parameters under incomplete information. This constructs a water quality prediction model adapted to scenarios with missing data. This not only breaks through the dependence of fixed-structure models on complete data in related technologies, but also effectively compensates for the modeling bias caused by the spatiotemporal gaps in monitoring data and incomplete observation of key information. It significantly improves the accuracy and stability of water quality prediction under incomplete information conditions, and provides reliable data support for decision-making in water environment management and pollution prevention and control, thereby enhancing the practicality and adaptability of the technology application.
[0006] Optionally, the data includes gate size and location information, and the monitoring data includes water quality monitoring data from tributary monitoring stations and water quality monitoring data from predicted cross sections.
[0007] Based on the aforementioned technical means, this application embodiment explicitly includes information such as gate size and location that affect water flow and pollutant migration and diffusion in the data. The monitoring data focuses on indicators such as water quality data from tributary monitoring stations and water quality data from prediction sections, providing accurate and crucial input data support for the water quality prediction model. This enables the model to capture the driving factors of water quality changes more comprehensively and accurately, effectively reducing modeling bias caused by incomplete or irrelevant input data. Consequently, it improves the pertinence and accuracy of water quality prediction under incomplete information conditions, providing a more reliable basis for water environment management decisions that is more in line with actual hydrological and water quality scenarios.
[0008] Optionally, the input to the water quality prediction model is water quality monitoring data from tributary monitoring stations, and the output of the water quality prediction model is water quality monitoring data from the prediction section.
[0009] Based on the above technical means, the embodiments of this application clearly define that the input of the water quality prediction model is the water quality monitoring data of the tributary monitoring station and the output is the water quality monitoring data of the prediction section. This effectively avoids modeling redundancy caused by irrelevant data interference, improves the model's efficiency and accuracy in capturing the laws of water quality migration and transformation, and thus outputs the water quality results of the prediction section more efficiently under incomplete information conditions, providing technical support for water quality early warning and control of key sections in water environment management.
[0010] Optionally, the pre-built water quality prediction model is trained using the dataset, including: dividing the dataset into a training set and a validation set, wherein the training set and the validation set include training samples and ground truth labels, the training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels include water quality monitoring data from prediction sections; training the water quality prediction model using the training set, updating the model parameters of the trained water quality prediction model using the ground truth labels, and validating the prediction performance of the water quality prediction model using the validation set to determine whether the model is overfitting.
[0011] Based on the above technical means, the embodiments of this application divide the dataset into a training set and a validation set containing training samples and ground truth labels. The training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels are water quality monitoring data from prediction sections. While training the model using the training set and updating the parameters using the ground truth labels, the validation set is used to verify the model's prediction performance and determine whether it is overfitting. This effectively ensures the model's fitting ability and generalization ability, and improves the reliability of water quality prediction under incomplete information conditions.
[0012] Optionally, the model state and model parameters of the water quality prediction model are corrected based on data assimilation and Bayesian inversion, including: during the data assimilation process, the missing flow data is used as an extended state variable, the posterior state of the water quality prediction model is calculated based on the observed and predicted values corresponding to the extended state variables, and the model state of the water quality prediction model is corrected based on the posterior state; during the Bayesian inversion process, the missing pollution source data is used as a random variable, the prior distribution of the random variable is corrected through monitoring data, the posterior distribution and confidence interval of the random variable are determined based on the corrected prior distribution, and the model parameters of the water quality prediction model are corrected based on the posterior distribution and confidence interval.
[0013] Based on the above technical means, this application embodiment adopts differentiated imputation logic for two types of key missing data in incomplete information: for missing flow data, it is treated as an extended state variable during the data assimilation process, and the posterior state is calculated by combining the observed value and predicted value corresponding to the variable, so as to achieve accurate correction of the model state; for missing pollution source data, it is regarded as a random variable in Bayesian inversion, and its prior distribution is corrected by existing monitoring data to obtain the posterior distribution and confidence interval to calibrate the model parameters, thereby achieving targeted imputation of different types of key missing data and effectively making up for the modeling bias caused by data gaps.
[0014] Optionally, the formula for the posterior state is: The formula for the posterior distribution is:
[0015] in, For unknown parameters of the model, For observation data, For the prior distribution of parameters set based on experience or historical data, Let θ be the likelihood function of the observed data D. To update the posterior distribution of parameters by integrating prior information with observed data, This is the normalization constant.
[0016] Optionally, the target network structure includes at least one hidden layer, fully connected layer, convolutional layer, pooling layer, recurrent layer, and dropout layer, and the number of hidden layers is determined by random generation.
[0017] Based on the above technical means, the embodiments of this application incorporate multiple layer types such as fully connected layers and convolutional layers and support flexible combinations. At the same time, the number of hidden layers is determined by random generation, making the target network structure more diverse and adaptable. This enables it to better capture the complex spatiotemporal characteristics of water quality data, effectively enhance the model's generalization ability, and thus improve the accuracy of water quality prediction under incomplete information conditions.
[0018] A second aspect of this application provides a water quality prediction device under incomplete information, comprising: an acquisition module for acquiring data and monitoring data from water quality monitoring stations within a target area; a training module for preprocessing the data and monitoring data, generating a dataset based on the preprocessed multi-source data, training a pre-constructed water quality prediction model using the dataset, randomly generating multiple network structures of the preset water quality model during training, determining the target network structure from the multiple network structures based on the training structure, and correcting the model state and model parameters of the water quality prediction model based on data assimilation and Bayesian inversion when some monitoring data is missing; and a prediction module for using the trained water quality prediction model to predict the water quality of the target area.
[0019] Optionally, the data includes gate size and location information, and the monitoring data includes water quality monitoring data from tributary monitoring stations and water quality monitoring data from predicted cross sections.
[0020] Optionally, the training module is further used to divide the dataset into a training set and a validation set, wherein the training set and the validation set include training samples and ground truth labels. The training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels include water quality monitoring data from prediction sections. The water quality prediction model is trained using the training set, the model parameters of the trained water quality prediction model are updated using the ground truth labels, and the prediction performance of the water quality prediction model is verified using the validation set to determine whether the model is overfitting.
[0021] Optionally, the training module is further used in the data assimilation process to treat the missing flow data as extended state variables, calculate the posterior state of the water quality prediction model based on the observed and predicted values corresponding to the extended state variables, and correct the model state of the water quality prediction model based on the posterior state; in the Bayesian inversion process, the missing pollution source data is treated as random variables, the prior distribution of the random variables is corrected by monitoring data, the posterior distribution and confidence interval of the random variables are determined based on the corrected prior distribution, and the model parameters of the water quality prediction model are corrected based on the posterior distribution and confidence interval.
[0022] Optionally, the formula for the posterior state is: The formula for the posterior distribution is:
[0023] in, For unknown parameters of the model, For observation data, For the prior distribution of parameters set based on experience or historical data, Let θ be the likelihood function of the observed data D. To update the posterior distribution of parameters by integrating prior information with observed data, This is the normalization constant.
[0024] Optionally, the target network structure includes at least one hidden layer, fully connected layer, convolutional layer, pooling layer, recurrent layer, and dropout layer, with the number of hidden layers determined by random generation.
[0025] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to implement the water quality prediction method under incomplete information as described in the above embodiments.
[0026] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the water quality prediction method under incomplete information as described in the above embodiments.
[0027] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0028] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a water quality prediction method under incomplete information according to an embodiment of this application; Figure 2 This is a flowchart of a method for constructing a water quality prediction model under incomplete information according to an embodiment of this application; Figure 3 This is a statistical chart of the MAPE error distribution of the one-stage model training set and validation set provided in the embodiments of this application; Figure 4 This is a comparison chart of predicted and measured values of a representative model validation set provided in an embodiment of this application. Figure 5 This is a histogram showing the distribution of training duration of a one-stage candidate model according to an embodiment of this application. Figure 6 A heatmap showing the correlation between the layer types of the one-stage model and MAPE according to an embodiment of this application; Figure 7 This is a histogram showing the cumulative distribution of the number of hidden layers and the number of layers of each type in a one-stage candidate model provided according to an embodiment of this application. Figure 8 This is a comparison chart of the optimal model prediction and measured values after data assimilation according to the embodiments of this application; Figure 9 This is a graph showing the error distribution of all model training and validation sets after data assimilation according to the embodiments of this application; Figure 10 This is a comparison chart of the predicted and measured values of the optimal model after introducing Bayesian inversion according to the embodiments of this application; Figure 11 This is a comparison chart of the predicted and measured values of the two-stage final optimal model provided in the embodiments of this application; Figure 12 This is a graph showing the change in performance indicators of the two-stage final model robustness test provided in the embodiments of this application; Figure 13 This is a block diagram illustrating a water quality prediction device under incomplete information according to an embodiment of this application. Figure 14 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0029] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0030] Water quality forecasting is a crucial link in water environment management and pollution control, and is of great significance for ensuring the ecological security of rivers and lakes, optimizing scheduling strategies, and realizing the scientific utilization of water resources. However, in actual monitoring, due to limited sensor deployment density, inconsistent sampling cycles, equipment failures, or abnormal data transmission, monitoring data often suffers from problems such as missing, abnormal, or discontinuous data. This incomplete information makes it difficult for water quality modeling methods in related technologies to accurately depict the evolution of water quality, resulting in unstable prediction results and large errors, which in turn affects the scientific nature and real-time performance of management decisions. The problems of data sparsity and multi-source heterogeneity are particularly prominent in complex river networks or dynamic water systems.
[0031] The following description, with reference to the accompanying drawings, describes a method, apparatus, device, and medium for water quality prediction under incomplete information according to embodiments of this application. Addressing the problems of low prediction accuracy and poor stability in related water quality prediction methods mentioned in the background art under conditions of incomplete information such as missing, abnormal, or discontinuous data, this application provides a method for constructing a water quality prediction model under incomplete information. This method includes: collecting basic data and monitoring data from automatic water quality monitoring stations; establishing a water quality prediction model based on existing information; a first-stage prediction model accuracy and structure analysis; establishing a water quality prediction model incorporating incomplete information; and a second-stage model performance evaluation and verification. Through a two-stage modeling strategy, combined with random structure generation, data assimilation, and Bayesian inversion techniques, high-precision water quality prediction under incomplete information is achieved. This method addresses the problem of data loss caused by limited sensor deployment, discontinuous sampling, or observation errors during water quality monitoring. It comprehensively utilizes observation data and multi-source auxiliary information to construct a water quality prediction model with adaptive modeling capabilities. By introducing a data-driven prediction mechanism and a multi-objective optimization strategy, the model can be dynamically updated and adaptively adjusted under conditions of missing data and uncertainty, thereby improving prediction accuracy and stability. This application can effectively improve the robustness and generalization ability of water quality prediction models, and provide reliable decision support for river and lake water environment management and interactive regulation.
[0032] Specifically, Figure 1 This is a flowchart illustrating a water quality prediction method under incomplete information, provided as an embodiment of this application.
[0033] like Figure 1 As shown, the water quality prediction method under incomplete information includes the following steps: In step S101, data and monitoring data from water quality monitoring stations within the target area are acquired.
[0034] It is understood that by clearly defining the scope of data acquisition as the target area, the source as water quality monitoring stations, and the types of data acquired as including data and monitoring data, the basic data required for modeling are determined, laying a reliable data foundation for subsequent model training, parameter correction, and water quality prediction processes.
[0035] In this embodiment of the application, the data includes gate size and location information, and the monitoring data includes water quality monitoring data from tributary monitoring stations and water quality monitoring data from predicted cross sections.
[0036] It is understood that the embodiments of this application, by explicitly including information such as gate size and location that affect water flow and the migration and diffusion of pollutants in the data, and by focusing on indicators such as water quality data from tributary monitoring stations and water quality data from prediction sections, provide accurate and crucial input data support for the water quality prediction model. This enables the model to capture the driving factors of water quality changes more comprehensively and accurately, effectively reducing modeling bias caused by incomplete or irrelevant input data. Consequently, it improves the pertinence and accuracy of water quality prediction under incomplete information conditions, and provides a more reliable basis for water environment management decisions that is more in line with actual hydrological and water quality scenarios.
[0037] Specifically, the data acquisition process includes collecting basic hydrological, hydrodynamic, and water quality information within the target area, simultaneously acquiring real-time data from automatic water quality monitoring stations, and integrating relevant information from tributary monitoring stations, forecast sections, and major gates. Gate size and location information can serve as a reference for the initial predicted values of flow data for each tributary. Water quality indicators include ammonia nitrogen, total nitrogen, total phosphorus, dissolved oxygen, pH, permanganate index, and chlorophyll. This preliminary processing of multi-source data lays the foundation for subsequent preprocessing work.
[0038] In this embodiment of the application, the input of the water quality prediction model is the water quality monitoring data of the tributary monitoring station, and the output of the water quality prediction model is the water quality monitoring data of the prediction section.
[0039] It is understood that the embodiments of this application clearly define the input of the water quality prediction model as water quality monitoring data from tributary monitoring stations and the output as water quality monitoring data from the prediction section. This effectively avoids modeling redundancy caused by irrelevant data interference, improves the model's efficiency and accuracy in capturing the laws of water quality migration and transformation, and thus outputs the water quality results of the prediction section more efficiently under incomplete information conditions, providing technical support for water quality early warning and control of key sections in water environment management.
[0040] In step S102, the data and monitoring data are preprocessed, and a dataset is generated based on the preprocessed multi-source data. The pre-built water quality prediction model is trained using the dataset. During the training process, multiple sets of network structures of the water quality preset model are randomly generated. The target network structure is determined from the multiple sets of network structures based on the training structure. In the case of missing monitoring data, the model state and model parameters of the water quality prediction model are corrected based on data assimilation and Bayesian inversion.
[0041] It is understood that the embodiments of this application preprocess the data and monitoring data to ensure data quality, generate a dataset based on the preprocessed multi-source data to train the water quality prediction model, and at the same time, select the target structure by randomly generating multiple sets of network structures, and correct the model state and parameters by combining data assimilation and Bayesian inversion. This not only improves the model's adaptability to water quality characteristics, but also solves the problem of missing data, thereby ensuring the accuracy and reliability of water quality prediction under incomplete information.
[0042] It should be noted that the target network structure is the optimal configuration selected from multiple randomly generated network structures. Data assimilation and Bayesian inversion are adaptation techniques for scenarios with incomplete information. The former is responsible for missing state variables, such as dynamic estimation of flow, while the latter is responsible for unknown parameters, such as inference of pollution source intensity. The two fill the data gaps through differentiated logic. Data assimilation adjusts the model's state variables to fit the actual changing trend, while Bayesian inversion calibrates the model parameters to adapt to the water environment conditions.
[0043] Specifically, the data preprocessing and dataset construction process is as follows: The multi-source data obtained in step S101 is standardized to form a unified dataset, providing a reliable foundation for subsequent model training; a data-driven prediction model framework is built based on this dataset, with water quality monitoring data from tributary monitoring stations as model input and water quality monitoring data from prediction sections as model output.
[0044] In this embodiment of the application, training a pre-built water quality prediction model using a dataset includes: dividing the dataset into a training set and a validation set, wherein the training set and the validation set include training samples and ground truth labels, the training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels include water quality monitoring data from prediction sections; training the water quality prediction model using the training set; updating the model parameters of the trained water quality prediction model using the ground truth labels; and validating the prediction performance of the water quality prediction model using the validation set to determine whether the model is overfitting.
[0045] It is understood that, in this embodiment of the application, the dataset is divided into a training set and a validation set containing training samples and ground truth labels. The training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels are water quality monitoring data from prediction sections. While training the model using the training set and updating the parameters using the ground truth labels, the model's prediction performance is verified using the validation set, and overfitting is determined. This effectively ensures the model's fitting ability and generalization ability, and improves the reliability of water quality prediction under incomplete information conditions.
[0046] It should be noted that the ground truth label refers to the actual water quality monitoring data of the predicted section corresponding to the training samples. It serves as a reference standard for measuring the accuracy of the model's prediction results. During model training, the ground truth label is used to calculate the error between the model's predicted values and the actual values. Then, through the error backpropagation mechanism, the model parameters are updated, allowing the model to gradually learn the mapping relationship between tributary monitoring data and water quality at the predicted section. In the validation phase, comparing the prediction results of the ground truth label with those of the validation set can quantify the model's prediction performance and determine whether the model has overfitting problems due to insufficient generalization ability caused by overfitting the training set data. This provides a key basis for model optimization and selection.
[0047] Specifically, the dataset partitioning and model training and validation process is as follows: The unified dataset constructed in step S102 is divided into a training set and a validation set according to a target ratio, such as an 8:2 ratio. The training samples in both sets are water quality monitoring data from tributary monitoring stations, and the ground truth labels are water quality monitoring data from predicted cross-sections. The model is iteratively trained using the training set, and the model parameters are updated through backpropagation of the error between the ground truth labels and the prediction results. After training, the model's prediction performance is tested using the validation set. The deviation between the prediction results and the ground truth labels is compared to determine if overfitting exists, ensuring that the model possesses both fitting and generalization abilities. The accuracy and structure analysis process of the first-stage prediction model is as follows: Model accuracy is evaluated using various evaluation metrics, common metrics are shown in Table 1. Table 1
[0048] The first stage uses the indicators in Table 1 to compare the performance of different network structure models. For example, RMSE (Root Mean Square Error) and MAE (Mean Absolute Error) are used to measure the absolute deviation between predicted and observed values, and NSE (Nash-Sutcliffe Efficiency Coefficient) is used to evaluate the model's fit to the water quality change trend, thereby selecting the structure with better prediction accuracy. Statistical and correlation analyses are conducted on the structural configuration, number of parameters, number of hidden layers, and layer types of different models to identify key structural features affecting prediction performance. Based on the analysis results, the optimal network structure suitable for the current water quality prediction problem is integrated and determined. At the same time, sensitivity analysis and uncertainty analysis methods are combined to evaluate the model's response characteristics to input disturbances, providing a basis and data support for subsequent second-stage modeling and model optimization.
[0049] In this embodiment, the model state and model parameters of the water quality prediction model are corrected based on data assimilation and Bayesian inversion, including: during the data assimilation process, missing flow data is used as extended state variables, the posterior state of the water quality prediction model is calculated based on the observed and predicted values corresponding to the extended state variables, and the model state of the water quality prediction model is corrected based on the posterior state; during the Bayesian inversion process, missing pollution source data is used as random variables, the prior distribution of the random variables is corrected through monitoring data, the posterior distribution and confidence interval of the random variables are determined based on the corrected prior distribution, and the model parameters of the water quality prediction model are corrected based on the posterior distribution and confidence interval.
[0050] It is understood that the embodiments of this application employ differentiated imputation logic for two types of key missing data in incomplete information: for missing flow data, it is treated as an extended state variable during data assimilation, and the posterior state is calculated by combining the observed and predicted values corresponding to the variable, thereby achieving accurate correction of the model state; for missing pollution source data, it is treated as a random variable in Bayesian inversion, and its prior distribution is corrected by existing monitoring data to obtain the posterior distribution and confidence interval to calibrate the model parameters, thereby achieving targeted imputation of different types of key missing data and effectively compensating for the modeling bias caused by data gaps.
[0051] In this embodiment of the application, the formula for the posterior state is: The formula for the posterior distribution is:
[0052] in, For unknown parameters of the model, For observation data, For the prior distribution of parameters set based on experience or historical data, Let θ be the likelihood function of the observed data D. To update the posterior distribution of parameters by integrating prior information with observed data, This is the normalization constant.
[0053] Specifically, water quality prediction models often face input uncertainty and parameter unidentification issues when partial observational data is missing. To address this problem, this application employs data assimilation and Bayesian inversion methods to dynamically correct the model state and parameters and characterize uncertainties. Data assimilation combines kinetic model predictions with observational data, real-time correcting system variables within a state-space framework, with the gain used to balance the uncertainties of model predictions and observational data. In water quality modeling, missing flow data is assimilated as an extended state variable along with pollutant concentrations to achieve dynamic estimation of water quality. Bayesian inversion treats missing pollution source data and other unknown parameters as random variables, correcting their prior distributions using observational data to obtain posterior distributions and confidence intervals. Numerical implementation can employ Markov chain Monte Carlo, Bayesian optimization, or variational inference methods to provide parameter point estimates and uncertainty quantification.
[0054] In this embodiment of the application, the target network structure includes at least one hidden layer, fully connected layer, convolutional layer, pooling layer, recurrent layer and dropout layer, and the number of hidden layers is determined by random generation.
[0055] It is understood that the embodiments of this application incorporate multiple layer types such as fully connected layers and convolutional layers and support flexible combinations, while determining the number of hidden layers in a random generation manner, making the target network structure more diverse and adaptable, better capturing the complex spatiotemporal characteristics of water quality data, effectively enhancing the model's generalization ability, and thus improving the accuracy of water quality prediction under incomplete information conditions.
[0056] It should be noted that the target network structure is a water quality prediction adaptation model architecture selected from multiple randomly generated candidate architectures through performance evaluation. Its layer type includes at least one of the following: hidden layer, fully connected layer, convolutional layer, pooling layer, recurrent layer, and dropout layer. The number of hidden layers is determined by random generation and is not specifically limited to a range such as 2 to 6.
[0057] Specifically, the network structure generation and selection process is as follows: a random structure generation method is adopted, randomly selecting a number of hidden layers from 2 to 6, with layer types covering fully connected layers, convolutional layers, pooling layers, recurrent layers, and dropout layers. Tens of thousands of candidate models with different structures are trained through this mechanism; the relative percentage error of each candidate model on the training set and validation set is calculated and statistically analyzed, and the prediction accuracy and stability under different structure configurations are comprehensively evaluated. Finally, the optimal network structure is selected for subsequent water quality prediction and analysis.
[0058] In step S103, the trained water quality prediction model is used to predict the water quality of the target area.
[0059] It is understood that the embodiments of this application apply the water quality prediction model, which has undergone data preprocessing, structural optimization, and state and parameter correction, to the target area to output accurate water quality prediction results, providing a direct and reliable decision-making basis for practical needs such as water environment management and pollution early warning.
[0060] Specifically, the water quality prediction model construction process incorporating incomplete information includes: problem modeling, prior setting, introduction of observational data, parameter and state updates, and uncertainty analysis. This approach fully utilizes limited observational data to improve prediction accuracy, dynamically estimates the most probable trajectory when key inputs are missing, and provides a scientific basis for water environment risk management. The two-stage model performance evaluation and verification process is as follows: Model accuracy is assessed using various evaluation indicators, common indicators are shown in Table 1; the indicators in Table 1 are reused in the second stage to monitor changes in model accuracy under incomplete information scenarios, such as the fluctuation range of RMSE and MAE when observational data is missing, and the stability of NSE, thereby verifying the model's predictive reliability under extreme conditions; sensitivity analysis methods are mainly divided into two categories: local sensitivity analysis and global sensitivity analysis, common methods are shown in Table 2. Table 2
[0061] Monte Carlo simulation, Bayesian inference, and ensemble simulation methods are employed to analyze output uncertainty, thereby identifying sources of uncertainty and quantifying prediction reliability. Simultaneously, robustness analysis is conducted to evaluate the model's stability under input disturbances, missing observations, or anomalous events. Methods include scenario disturbance testing and sensitive parameter disturbance simulation. By observing changes in key indicators such as RMSE, MAE, and NSE, the model's ability to maintain prediction accuracy under extreme or incomplete information conditions is verified, providing reliable decision support for water environment management.
[0062] The water quality prediction method under incomplete information proposed in this application integrates multi-source water quality data of the target area and performs standardized preprocessing. It combines the random generation and optimal selection of multiple network structures, and uses data assimilation and Bayesian inversion mechanisms to dynamically correct the model state and parameters under incomplete information. This constructs a water quality prediction model adapted to scenarios with missing data. It not only breaks through the dependence of fixed-structure models on complete data in related technologies, but also effectively makes up for the modeling bias caused by the spatiotemporal gap of monitoring data and the incomplete observation of key information. It significantly improves the accuracy and stability of water quality prediction under incomplete information conditions, and provides reliable data support for decision-making in water environment management and pollution prevention and control, thereby enhancing the practicality and adaptability of the technology.
[0063] The following will illustrate the method for constructing a water quality prediction model under incomplete information provided in this application through a specific embodiment, such as... Figure 2 As shown, the specific steps are as follows: In step one, basic data and monitoring data from automatic water quality monitoring stations were collected: approximately 19,000 sets of water quality monitoring data from 17 tributary monitoring stations were collected from February 4, 2023 to December 31, 2023. Monitoring indicators included pH (hydrogen ion concentration index), turbidity, DO (dissolved oxygen), water temperature, conductivity, TP (total phosphorus), CODMn (potassium permanganate index), and ammonia nitrogen (…). - The monitoring interval was 2 hours. At the same time, 1,207 water quality monitoring data were collected from the light chemical warehouse section from April 7, 2023 to December 31, 2023. The indicators included algae density, pH, dissolved oxygen, turbidity, water temperature, conductivity, total phosphorus, ammonia nitrogen, chlorophyll a, total nitrogen and permanganate index. The monitoring interval was 2 hours. In addition, the size, location and operation status of the main gates in the river section and upstream and downstream were investigated to provide hydrodynamic information support.
[0064] In step two, a water quality prediction model based on the collected monitoring data and hydrodynamic information is established: using tributary water quality monitoring data as model input and light chemical warehouse section water quality monitoring data as output, a data-driven black box prediction model is constructed; the dataset is divided into training and validation sets in an 8:2 ratio, and model structures containing 2 to 6 hidden layers are randomly generated, with layer types covering fully connected layers, convolutional layers, pooling layers, recurrent layers, and dropout layers, training approximately 60,000 black box models in total.
[0065] In step three, an accuracy and structure analysis is performed on the one-stage prediction model: the model training results show that the mean absolute percentage error (MAPE) of the training set is distributed in the range of 15% to 25%, while the MAPE of the validation set is concentrated in the range of 25% to 35%. Figure 3 As shown, the model with a MAPE of 20.71% on the validation set and 18.73% on the training set can capture the overall fluctuations and average levels of water quality changes, but its response to sudden and drastic changes is limited. Figure 4 As shown, the training time for most models is concentrated between 1 and 6 seconds, with an average of about 2.8 seconds. Figure 5 As shown, the structure-performance correlation analysis reveals a positive correlation between MAPE on the training and validation sets, a positive correlation between the number of pooling layers and error, a negative correlation between the number of convolutional layers and error, no accuracy advantage from excessive hidden layers, and poor fitting performance from models with only one fully connected layer. The heatmap showing the correlation between layer types and MAPE in the first-stage model is shown below. Figure 6 As shown, the histogram of the cumulative distribution of the number of hidden layers and the number of each type of layer in the first-stage candidate model is as follows: Figure 7 As shown.
[0066] In step four, a water quality prediction model incorporating incomplete information is established: unknown water flow data is estimated using data assimilation, and the optimal model predictions after data assimilation are compared with the measured values, for example… Figure 8 As shown, the error distributions of all models on the training and validation sets after data assimilation are as follows: Figure 9 As shown, after introducing the method, the MAPE range of the training set is 8.62%~17.21%, and that of the validation set is 9.60%~18.21%, indicating a significant improvement in model performance. Furthermore, Bayesian inversion is introduced to estimate unknown pollution source inputs. A comparison between the optimal model predictions and actual values after introducing Bayesian inversion is shown below. Figure 10 As shown, the overall prediction accuracy of the two-stage model is improved, with predictions largely consistent with measured values for most periods. However, local errors remain significant in the peak-to-trough intervals. The final optimal two-stage model achieved a MAPE of 7.54% on the training set and 7.93% on the validation set, with an R² of 0.9101 on the training set and 0.9083 on the validation set. The model exhibits good fitting with no overfitting. A comparison of the final optimal two-stage model predictions with measured values is shown below. Figure 11 As shown.
[0067] In step five, a two-stage model performance evaluation and validation is performed: robustness testing is conducted, such as... Figure 12 As shown, when the input noise increases from 0.01 to 0.1, MAE and RMSE rise slowly, MAPE increases from about 7.76% to 13.87%, and R² decreases from 0.908 to 0.795. Although the model prediction error increases slightly, it still maintains high accuracy and interpretability overall.
[0068] In summary, the embodiments of this application have at least the following beneficial effects: (1) Effectively solves the problems of low prediction accuracy and poor model stability caused by missing data, incomplete observation and parameter uncertainty in water quality prediction in related technologies, and fills the technical gap in the scenario of incomplete information; (2) To achieve model adaptation and dynamic optimization, through multi-source data fusion, adaptive modeling, black box model training and two-stage optimization strategy, combined with data assimilation and Bayesian inversion technology, the dynamic update of model state and parameters and the quantification of uncertainty are realized, thereby improving the model's adaptability to complex water environments. (3) Ensure the reliability of predictions under extreme scenarios. Through uncertainty analysis and robustness verification, ensure that the model can maintain stable prediction performance under complex working conditions such as extreme events and input disturbances, and reduce the risk of missed or false alarms in pollution early warning and water environment management. (4) Significantly improves performance indicators. Compared with related water quality prediction methods, it greatly improves the accuracy, robustness and generalization ability of water quality prediction. The model error is comparable to the error level of the standard monitoring method, and there is no overfitting phenomenon. It has strong practical application adaptability. (5) Supports intelligent water environment regulation and decision-making, provides stable and reliable technical support for intelligent water environment regulation, pollution prevention and control and scientific management, improves the scientific nature and timeliness of decision-making, and has high practical value and promotion significance.
[0069] Next, referring to the accompanying drawings, a water quality prediction device based on embodiments of this application is described, which is based on incomplete information.
[0070] Figure 13 This is a block diagram of a water quality prediction device under incomplete information according to an embodiment of this application.
[0071] like Figure 13 As shown, the water quality prediction device 130 under incomplete information includes: an acquisition module 1301, a training module 1302, and a prediction module 1303.
[0072] The acquisition module 1301 is used to acquire data and monitoring data from water quality monitoring stations within the target area; the training module 1302 is used to preprocess the data and monitoring data, generate a dataset based on the preprocessed multi-source data, train a pre-built water quality prediction model using the dataset, randomly generate multiple sets of network structures for the water quality preset model during the training process, determine the target network structure from the multiple sets of network structures based on the training structure, and correct the model state and model parameters of the water quality prediction model based on data assimilation and Bayesian inversion when some monitoring data is missing; the prediction module 1303 is used to predict the water quality of the target area using the trained water quality prediction model.
[0073] In this embodiment of the application, the data includes gate size and location information, and the monitoring data includes water quality monitoring data from tributary monitoring stations and water quality monitoring data from predicted cross sections.
[0074] In this embodiment, the training module 1302 is further used to divide the dataset into a training set and a validation set, wherein the training set and the validation set include training samples and ground truth labels. The training samples are water quality monitoring data from tributary monitoring stations, and the ground truth labels include water quality monitoring data from prediction sections. The water quality prediction model is trained using the training set, the model parameters of the trained water quality prediction model are updated using the ground truth labels, and the prediction performance of the water quality prediction model is verified using the validation set to determine whether the model is overfitting.
[0075] In this embodiment, the training module 1302 is further configured to, during the data assimilation process, use the missing flow data as an extended state variable, calculate the posterior state of the water quality prediction model based on the observed and predicted values corresponding to the extended state variable, and correct the model state of the water quality prediction model based on the posterior state; during the Bayesian inversion process, use the missing pollution source data as a random variable, correct the prior distribution of the random variable through monitoring data, determine the posterior distribution and confidence interval of the random variable based on the corrected prior distribution, and correct the model parameters of the water quality prediction model based on the posterior distribution and confidence interval.
[0076] In this embodiment of the application, the formula for the posterior state is: The formula for the posterior distribution is:
[0077] in, For unknown parameters of the model, For observation data, For the prior distribution of parameters set based on experience or historical data, Let θ be the likelihood function of the observed data D. To update the posterior distribution of parameters by integrating prior information with observed data, This is the normalization constant.
[0078] In this embodiment of the application, the target network structure includes at least one hidden layer, fully connected layer, convolutional layer, pooling layer, recurrent layer and dropout layer, and the number of hidden layers is determined by random generation.
[0079] It should be noted that the foregoing explanation of the water quality prediction method embodiment under incomplete information also applies to the water quality prediction device under incomplete information in this embodiment, and will not be repeated here.
[0080] The water quality prediction device under incomplete information proposed in this application integrates multi-source water quality data of the target area and performs standardized preprocessing. It combines the random generation and optimal selection of multiple network structures, and uses data assimilation and Bayesian inversion mechanisms to dynamically correct the model state and parameters under incomplete information. This constructs a water quality prediction model adapted to scenarios with missing data. It not only breaks through the dependence of fixed structure models on complete data in related technologies, but also effectively makes up for the modeling bias caused by the spatiotemporal gap of monitoring data and the incomplete observation of key information. It significantly improves the accuracy and stability of water quality prediction under incomplete information conditions, and provides reliable data support for decision-making in water environment management and pollution prevention and control, thereby enhancing the practicality and adaptability of the technology.
[0081] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 1401, the processor 1402, and the computer program stored on the memory 1401 and executable on the processor 1402.
[0082] When the processor 1402 executes the program, it implements the water quality prediction method under incomplete information provided in the above embodiments.
[0083] Furthermore, electronic devices also include: Communication interface 1403 is used for communication between memory 1401 and processor 1402.
[0084] The memory 1401 is used to store computer programs that can run on the processor 1402.
[0085] The memory 1401 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0086] If the memory 1401, processor 1402, and communication interface 1403 are implemented independently, then the communication interface 1403, memory 1401, and processor 1402 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 14 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0087] Optionally, in a specific implementation, if the memory 1401, processor 1402, and communication interface 1403 are integrated on a single chip, then the memory 1401, processor 1402, and communication interface 1403 can communicate with each other through an internal interface.
[0088] The processor 1402 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0089] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for water quality prediction under incomplete information.
[0090] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0091] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0092] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0093] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0094] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0095] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A water quality prediction method under incomplete information, characterized by, The method comprises the following steps: Obtaining data and monitoring data of water quality monitoring stations in a target area; Preprocessing the data and the monitoring data, generating a data set according to the preprocessed multi-source data, training a pre-constructed water quality prediction model using the data set, randomly generating multiple sets of network structures of the water quality prediction model during the training process, determining a target network structure from the multiple sets of network structures according to the training structure, and correcting the model state and the model parameters of the water quality prediction model based on data assimilation and Bayesian inversion in the case of missing part of the monitoring data; Using the trained water quality prediction model to predict the water quality of the target area.
2. The water quality prediction method under incomplete information according to claim 1, characterized in that, The data includes gate size and position information, and the monitoring data includes water quality monitoring data of tributary monitoring stations and predicted cross-section water quality monitoring data.
3. The water quality prediction method under incomplete information according to claim 2, characterized in that, The input of the water quality prediction model is the water quality monitoring data of the tributary monitoring stations, and the output of the water quality prediction model is the predicted cross-section water quality monitoring data.
4. The water quality prediction method under incomplete information according to claim 3, characterized in that, The training of the pre-constructed water quality prediction model using the data set comprises: Dividing the data set into a training set and a validation set, wherein the training set and the validation set include training samples and true value labels, the training samples are the water quality monitoring data of the tributary monitoring stations, and the true value labels include the predicted cross-section water quality monitoring data; Training the water quality prediction model using the training set, updating the model parameters of the trained water quality prediction model using the true value labels, and verifying the prediction performance of the water quality prediction model using the validation set to determine whether the model is overfitting.
5. The water quality prediction method under incomplete information according to claim 1, characterized in that, The correction of the model state and the model parameters of the water quality prediction model based on data assimilation and Bayesian inversion comprises: In the data assimilation process, the missing flow data is taken as an extended state variable, the posterior state of the water quality prediction model is calculated according to the observation value and the prediction value corresponding to the extended state variable, and the model state of the water quality prediction model is corrected based on the posterior state; In the Bayesian inversion process, the missing pollution source data is taken as a random variable, the prior distribution of the random variable is corrected through the monitoring data, the posterior distribution and the confidence interval of the random variable are determined according to the corrected prior distribution, and the model parameters of the water quality prediction model are corrected according to the posterior distribution and the confidence interval.
6. The water quality prediction method under incomplete information according to claim 5, characterized in that, The formula of the posterior state is: ; The formula of the posterior distribution is: where, is the model unknown parameter, is the observation data, is the parameter prior distribution set based on experience or historical data, is the likelihood function of the observation data D given the parameter θ, is the updated parameter posterior distribution after fusing the prior information and the observation data, is the normalization constant.
7. The water quality prediction method under incomplete information according to claim 1, characterized in that, The target network structure includes at least one of a hidden layer, a full connection layer, a convolution layer, a pooling layer, a cycle layer and a dropout layer, and the number of the hidden layers is determined in a random generation manner.
8. A water quality prediction device under incomplete information, characterized by, Comprise: An acquisition module is configured to acquire data and monitoring data of water quality monitoring stations in a target area; The training module is configured to preprocess the material data and the monitoring data, generate a data set according to the preprocessed multi-source data, and train a pre-constructed water quality prediction model by using the data set; in the training process, a plurality of network structures of the water quality preset model are randomly generated, a target network structure is determined from the plurality of network structures according to a training structure, and in the case of missing part of the monitoring data, the model state and the model parameter of the water quality prediction model are corrected based on data assimilation and Bayesian inversion. The prediction module is configured to perform water quality prediction on the target region by using the trained water quality prediction model.
9. An electronic device, comprising: The computer program or instructions are executed to implement the water quality prediction method under incomplete information according to any one of claims 1-7. The computer program or instructions are executed to implement the water quality prediction method under incomplete information according to any one of claims 1-7.
10. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that,