Training method of water quality prediction model, water quality prediction method and water quality monitoring system
By training the water quality prediction model on the edge server side and aggregating the central cloud server, the problem of difficulty in dealing with large-scale and real-time water quality data and large-scale water quality pollution prediction in the existing technology is solved, and efficient and accurate water quality pollution prediction and data privacy protection are achieved.
Patent Information
- Application Number
- CN202510102282.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively process large-scale and real-time water quality data, and it is difficult to conduct large-scale water quality pollution prediction. Traditional methods have efficiency and feasibility problems when processing distributed data.
By training the water quality prediction model on the edge server side, local training is performed using the data collected in real time by intelligent IoT sensors, and the adjusted local learning parameters are sent to the central cloud server for aggregation, and the global water quality prediction model is updated.
It realizes the accuracy of water quality pollution on a large scale while reducing network communication costs, and each edge server does not need to interact with data, protecting the privacy of water quality detection data.
Smart Images

Figure CN120047268A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental monitoring, and in particular, to a method for training a water quality prediction model, a water quality prediction method, and a water quality monitoring system. Background Art
[0002] The quality management of urban drinking water is an important issue related to public health and a crucial link in the construction of smart cities. With the acceleration of urbanization, the problem of water resource pollution is becoming increasingly serious. Water quality detection is the key means to ensure the quality of urban water supply.
[0003] Traditional water quality detection is mainly based on laboratory research and uses sample data for analysis. This method has obvious limitations in dealing with complex and large-scale real-time data and cannot effectively reflect the real-time data sampling in the real world.
[0004] In the prior art, all sensor data is centralized to a central computing platform for processing and analysis. Due to geographical distance and network throughput limitations, this method is not suitable for processing a large amount of real-time water quality data from various locations along the river.
[0005] If an edge server is used to predict water quality pollution, for example, ML and DL algorithms are used to analyze various environmental index data from IoT (Internet of Things) sensors to deduce the statistical relationship between sensor data and water quality. These models usually have limitations and are only applicable to the prediction of water quality within a specific time and location, making it difficult to accurately predict pollution over a large area. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method for training a water quality prediction model, a water quality prediction method, and a water quality monitoring system, which can predict the water quality pollution situation over a large area while reducing network communication costs.
[0007] In a first aspect, a method for training a water quality prediction model is provided, which is applied to an edge server. The method includes:
[0008] Obtaining a training sample each time of iteration; the training sample is water quality detection data collected in real time by a variety of intelligent Internet of Things sensors;
[0009] Judging whether the number of training samples meets a preset training sample threshold;
[0010] If it meets the requirement, training the current version of the global water quality prediction model sent by the central cloud server based on the training sample and adjusting the local learning parameters of the model;
[0011] After reaching the preset number of iterations, send the finally adjusted local learning parameters to the central cloud server, so that the central cloud server aggregates the local learning parameters uploaded by each edge server to obtain global learning parameters, and updates the global water quality prediction model based on the global learning parameters;
[0012] Receive the updated global water quality prediction model sent by the central cloud server.
[0013] Optionally, obtaining a training sample each time an iteration is performed includes:
[0014] Obtain the water quality detection data collected in real time by a variety of intelligent Internet of Things sensors in the target monitoring area each time an iteration is performed;
[0015] Store the water quality detection data collected in real time into the pre-constructed data queue model;
[0016] Determine the water quality detection data collected in real time and the existing water quality detection data in the data queue model as the training sample together.
[0017] Optionally, the pre-constructed data queue model is:
[0018]
[0019] In the formula, represents the size of the water quality detection data stored by the edge server ES j within the time step t; represents the size of the water quality detection data updated by the edge server ES j at the moment t + 1; j represents the number of the edge server; is the set of edge servers ES j ; represents the intelligent Internet of Things sensor; i represents the number of the intelligent Internet of Things sensor; m represents the data collected by the intelligent Internet of Things sensor; S is the set of intelligent Internet of Things sensors; represents the intelligent Internet of Things sensor the size of the water quality data collected within the time step t; y (i,j) [t] is the scheduling vector between the edge server ES j and the intelligent Internet of Things sensor , specifically indicating whether the intelligent Internet of Things sensor is scheduled to the edge server ES j at the time step t. If it is scheduled, then y (i,j) [t] = 1, otherwise y (i,j) [t] = 0; I j is the scheduling vector of the data queue model. When I jWhen I = 1, it means that the amount of data stored in the data queue model is greater than or equal to a preset training sample threshold; when I j = 0, it means that the amount of data stored in the data queue model is less than the preset training sample threshold.
[0020] Optionally, the method further includes:
[0021] If the number of training samples meets the preset training sample threshold, clear the data in the data queue model.
[0022] Optionally, the global water quality prediction model is composed of a Pre-DBN encoder, a D-Tencoder encoder, and an L-Decoder decoder;
[0023] The Pre-DBN encoder is composed of multiple layers of Gaussian-Bernoulli restricted Boltzmann machines and a fully connected output layer, and is used to extract feature vectors from water quality detection data layer by layer based on the multiple layers of Gaussian-Bernoulli restricted Boltzmann machines, and output the feature vectors based on the fully connected output layer. After passing through the Pre-DBN encoder, the initial local learning parameters of the global water quality prediction model are obtained;
[0024] The D-Tencoder encoder is used to capture the relationships between feature vectors based on the self-attention mechanism and fine-tune the initial local learning parameters;
[0025] The L-Decoder decoder is used to predict the pollution situation of water quality detection data based on the relationships between feature vectors, and the pollution situation includes pollution type, pollution area, and pollution time.
[0026] In a second aspect, a method for training a water quality prediction model is provided, which is applied to a central cloud server. The method includes:
[0027] Send the current version of the global water quality prediction model to each edge server for local training by each edge server;
[0028] Receive the local learning parameters uploaded by each edge server;
[0029] Aggregate the local learning parameters of each edge server based on a preset parameter aggregation formula to obtain global learning parameters; the preset parameter aggregation formula is:
[0030]
[0031] where ω G is the global learning parameter; ω i is the local learning parameter; is the set of edge servers ES j ; represents an intelligent Internet of Things sensor.
[0032] Update the global water quality prediction model based on the global learning parameters;
[0033] Send the updated global water quality prediction model to each edge server.
[0034] Optionally, before aggregating the local learning parameters of each edge server, the method further includes:
[0035] Calculate the statistical values of the local learning parameters of each edge server, and the statistical values are at least one of the following: mean, median, standard deviation;
[0036] Mutually verify the credibility of the locally trained learning parameters based on the statistical values of each edge server;
[0037] Eliminate the local learning parameters with credibility lower than the preset threshold.
[0038] In a third aspect, a water quality prediction method is provided, which is applied to an edge server. The method includes:
[0039] Obtain the water quality detection data uploaded by the intelligent Internet of Things sensors that match it based on the preset matching relationship between the intelligent Internet of Things sensors and the edge server;
[0040] Input the water quality detection data into the water quality prediction model trained by any method in the first aspect for water quality prediction to obtain a water pollution prediction result.
[0041] In a fourth aspect, a water quality monitoring system is provided. The system includes distributed intelligent Internet of Things sensors, distributed edge servers, and a central cloud server;
[0042] The intelligent Internet of Things sensors are used to collect water quality detection data. The intelligent Internet of Things sensors at least include a temperature sensor, a pH sensor, a dissolved oxygen sensor, a total organic carbon sensor, a total nitrogen sensor, a chlorophyll-a sensor, and a turbidity sensor;
[0043] The edge server is used to collect the water quality detection data collected by the intelligent Internet of Things sensors, train the global water quality prediction model based on the collected water quality detection data, and predict the water pollution situation based on the trained global water quality prediction model;
[0044] The central cloud server is used to update the global learning parameters of the global water quality prediction model and send the updated global water quality prediction model to each edge server.
[0045] In a fifth aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0046] A memory for storing a computer program;
[0047] A processor for implementing the method steps described in any one of the first to third aspects when executing the program stored in the memory.
[0048] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of the first to third aspects are implemented.
[0049] A method for training a water quality prediction model, a water quality prediction method, and a water quality monitoring system provided by the present invention. By training the model on the edge server side, it is not necessary to share the sensing data with the central cloud server. It only needs to be transmitted to the local edge server, and local training can be performed on the global water quality prediction model issued by the central cloud server. This enables the system to approach the central computing performance while significantly reducing the communication cost of the entire network system. Moreover, there is no need for data interaction between edge servers, protecting the privacy of water quality detection data. And through the global water quality prediction model, it is possible to predict the water quality pollution situation in a large area.
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, so they should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0052] Figure 1 Shows the structural schematic diagram of the water quality monitoring system provided by the embodiments of the present invention;
[0053] Figure 2 Shows the flowchart of a method for training a water quality prediction model provided by the embodiments of the present invention;
[0054] Figure 3 Shows the structural schematic diagram of storing water quality detection data by the data queue model provided by the embodiments of the present invention;
[0055] Figure 4 Shows the structural schematic diagram of the water quality prediction model provided by the embodiments of the present invention;
[0056] Figure 5 Shows the structural schematic diagram of a single-layer RBM provided by the embodiments of the present invention;
[0057] Figure 6 Shows the schematic diagram of the solution process by the CD-k algorithm provided by the embodiments of the present invention;
[0058] Figure 7 Shows the schematic diagram of the structure of the D-Tencoder encoder provided by the embodiments of the present invention;
[0059] Figure 8 Shows the flowchart of a method for training a water quality prediction model provided by another embodiment of the present invention;
[0060] Figure 9 Shows the flowchart of a water quality prediction method provided by the embodiments of the present invention;
[0061] Figure 10 Shows the schematic diagram of the structure of an electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0062] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] Before the embodiments of the present invention, the existing related technologies for water quality detection are introduced first:
[0064] (1) Traditional water quality monitoring research technology: In the early stage, it was mainly laboratory-based research, using sample data for analysis. This method has obvious limitations in dealing with complex and large-scale real-time data and cannot effectively reflect the real-time data sampling of the actual environment.
[0065] (2) ML (Machine Learning) and DL (Deep Learning) technologies: Use ML and DL algorithms to analyze various environmental index data from IoT (Internet of Things) sensors and deduce the statistical relationship between the sensor data and water quality. These models usually have limitations and are only applicable to the prediction of water quality within a specific time and location, and it is difficult to perform accurate large-scale pollution prediction.
[0066] (3) Drinking water quality prediction model technology: Liu et al. proposed using long short-term memory networks for time series prediction. This model is constructed after collecting all data and is not suitable for real-time computing. In addition, it mainly predicts the measurement change trends of various water quality indicators rather than directly predicting pollution results.
[0067] (4) Centralized computing method / technology: All sensor data is centralized to a central computing platform for processing and analysis. Due to geographical distance and network throughput limitations, this method is not suitable for processing a large amount of real-time water quality data distributed across rivers.
[0068] (5) Smart Internet of Things sensor design technology: Cloete et al. and Vijayakumar et al. proposed the design of smart Internet of Things sensors for real-time water quality monitoring. These studies mainly focus on sensor hardware design and do not address the issues of large-scale data processing and analysis.
[0069] (6) Unmanned surface vehicle technology: Madeo et al. proposed a low-cost unmanned surface vehicle for extensive water quality monitoring. Although this method improves the flexibility of data collection, it cannot meet the requirements of real-time data transmission.
[0070] (7) Wireless sensor network technology: Olatinwo and Joubert studied the throughput and fairness maximization of a wireless sensor network system for water quality monitoring. This research mainly focuses on network performance optimization rather than data analysis and the development of prediction models.
[0071] Although these existing technologies have made progress in their respective fields, they all have some limitations. The main problems include:
[0072] 1. Difficulty in processing large-scale and real-time water quality data.
[0073] 2. Lack of the ability to effectively integrate and analyze geographically distributed data.
[0074] 3. Prediction models are usually limited to specific times and locations and lack universality.
[0075] 4. Centralized computing methods face efficiency and feasibility issues when dealing with distributed data.
[0076] Based on this, the embodiments of the present invention provide a water quality monitoring system based on edge computing and federated learning, as Figure 1 shown. The system includes distributed smart Internet of Things sensors, distributed edge servers, and a central cloud server.
[0077] In a feasible implementation, intelligent Internet of Things sensors are deployed on the bank of the water intake area of urban water use to collect water quality detection data. When the quality of the collected data is insufficient or poor, the prediction model cannot accurately reflect the actual situation of water quality. Therefore, in the embodiments of the present invention, multiple types of intelligent Internet of Things sensors are provided, and the intelligent Internet of Things sensors at least include a temperature sensor, a pH value sensor, a dissolved oxygen sensor, a total organic carbon sensor, a total nitrogen sensor, a chlorophyll-a sensor, and a turbidity sensor. This improves the diversity and quality of the data.
[0078] The edge server is used to collect the water quality detection data collected by the intelligent Internet of Things sensors, train a global water quality prediction model based on the collected water quality detection data, and predict the water quality pollution situation based on the trained global water quality prediction model.
[0079] In a feasible implementation, narrowband Internet of Things NB-IoT communication can be used between the intelligent Internet of Things sensors and the edge server to establish a large number of connections between the intelligent Internet of Things sensors and the edge server. NB-IoT allows more than 100,000 devices to be connected in each small area, and the connection number can be increased by using multiple carriers. The maximum throughput rates are 200 kbps and 20 kbps in the downlink and uplink respectively, and the payload size of each piece of data is 1600 bytes. In addition, the communication range of NB-IoT is about 10 kilometers. In addition, using NB-IoT is relatively power-saving, which can extend the service life of intelligent Internet of Things sensors with limited battery life.
[0080] In another feasible implementation, drones can also be deployed around the intelligent Internet of Things sensors to receive the data collected by the intelligent Internet of Things sensors through irregular cruising, and then send the data to the corresponding edge server. The intelligent Internet of Things sensors and the drones transmit data through a wireless network. This not only reduces the on-site wiring of the intelligent Internet of Things sensors, but also improves the real-time performance of data transmission.
[0081] In the embodiments of the present invention, in order to ensure the data balance of each edge server, a scheduling rule is set between the edge server and the intelligent Internet of Things sensors: an edge server can be connected to multiple intelligent Internet of Things sensors, and an intelligent Internet of Things sensor can only be scheduled by one edge server at the same time.
[0082] In a feasible implementation, the edge server and the intelligent Internet of Things sensors are constrained by a predefined model. The predefined model is as follows:
[0083]
[0084] In the formula, ES jis an edge server, and j is the number of the edge server; is a set of edge servers; represents an intelligent Internet of Things sensor; i represents the number of the intelligent Internet of Things sensor; m represents the data collected by the intelligent Internet of Things sensor; S represents the set of intelligent Internet of Things sensors; α j represents the threshold of the number of intelligent Internet of Things sensors that each edge server can schedule simultaneously; y (i,j) [t] is at the time step t and is the scheduling vector between; if it is scheduled, then y (i,j) [t]=1, otherwise y (i,j) [t]=0.
[0085] Among them, Equation 1 restricts that an intelligent Internet of Things sensor can only be scheduled by one edge server;
[0086] Equation 2 restricts that the number of intelligent Internet of Things sensors scheduled by one edge server simultaneously does not exceed α j .
[0087] Through this scheduling rule, it is possible to prevent significant differences in the data volume between edge servers, ensure data balance, and this effect is particularly obvious in environments where the distribution degree of intelligent Internet of Things sensors is different or some edge servers may monopolize the data of intelligent Internet of Things sensors.
[0088] The central cloud server is used to update the global learning parameters of the global water quality prediction model and send the updated global water quality prediction model to each edge server.
[0089] Through the water quality monitoring system constructed by the implementation of the present invention, it is not necessary to share the sensing data with the central cloud server, and it only needs to be transmitted to the local edge server, and the water quality prediction is carried out by the global water quality prediction model sent by the central cloud server, so that the system can significantly reduce the communication cost of the entire network system while approaching the central computing performance; and each edge server does not need to perform data interaction, protecting the privacy of the water quality detection data. And through the global water quality prediction model, the prediction of the water quality pollution situation in a large range is satisfied.
[0090] Based on the constructed water quality monitoring system, an embodiment of the present invention provides a training method for a water quality prediction model. The execution subject of this method is an edge server, such as Figure 2 shown, including the following steps:
[0091] Step S201: Obtain a training sample each time of iteration.
[0092] In a feasible implementation, obtaining a training sample each time of iteration includes:
[0093] Step S201A: Obtain the water quality detection data collected in real time by a variety of intelligent Internet of Things sensors in the target monitoring area during each iteration. The water quality detection data includes parameters such as temperature, pH value, dissolved oxygen, total organic carbon, total nitrogen, chlorophyll-a, and turbidity.
[0094] In the embodiment of the present invention, a time step is set to obtain the water quality detection data collected by the intelligent Internet of Things sensors. Every time the time step t elapses, the water quality detection data is obtained once. That is, the moment when the time step t is reached is the iteration moment.
[0095] Step S201B: Store the water quality detection data collected in real time into a pre-constructed data queue model.
[0096] Among them, the pre-constructed data queue model is:
[0097]
[0098] In the formula, represents the size of the water quality detection data stored by the edge server ES j within the time step t; represents the size of the water quality detection data updated by the edge server ES j at the moment of t + 1; is the set of the edge server ES j ; represents the intelligent Internet of Things sensor; S is the set of intelligent Internet of Things sensors; represents the intelligent Internet of Things sensor the size of the water quality data collected within the time step t; y (i,j) [t] is the scheduling vector between the edge server ES j and the intelligent Internet of Things sensor , specifically indicating whether the intelligent Internet of Things sensor is scheduled to the edge server ES j during the time step t. If it is scheduled, then y (i,j) [t] = 1, otherwise y (i,j) [t] = 0; I j is the scheduling vector of the data queue model. When I j = 1, it means that the data volume stored in this data queue model is greater than or equal to the preset training sample threshold; when I j = 0, it means that the data volume stored in this data queue model is less than the preset training sample threshold.
[0099] According to this data queue model, when , I j = 0; it means that the edge server ES jThe amount of data received is insufficient to start training, and the queue continues to accumulate data until
[0100] When I j = 1, it means that the edge server ES j starts training, uses the data that meets the training start threshold D j for local training, and at the same time the queue will be emptied, and after emptying, the subsequent collected data will be accumulated again.
[0101] Through this loop accumulation and judgment mechanism, it is ensured that no additional training process will be triggered before the data meets the training requirements, thereby optimizing the use of computing resources and network bandwidth.
[0102] And it should be noted that the storage capacity of the data queue model is not limited by the training start threshold D j and is designed to have an infinite storage space. This setting allows the edge server to receive more data than D j in a single iteration and store all the data in the queue.
[0103] Step S201C: Determine the real-time collected water quality detection data and the existing water quality detection data in the data queue model as training samples together.
[0104] Step S202: Judge whether the number of training samples meets the preset training sample threshold.
[0105] In this step, the amount of data collected by the intelligent Internet of Things sensors is constrained by a predefined model; the predefined model is:
[0106]
[0107] In the formula, represents the sensing data collected by the intelligent Internet of Things sensor at iteration time t; represents the amount of data existing in the data queue model of the edge server at iteration time t; other parameters are the same as formula 3 above.
[0108] As Figure 3 shown, is the amount of new data that needs to be collected to start local training. According to formula 4, when local training needs to be started, then the amount of data uploaded by the intelligent Internet of Things sensors scheduled by the edge server ES j should be greater than or equal to When local training does not need to be started, It only needs to be greater than 0, which relaxes the range of the amount of data collected by the intelligent Internet of Things sensor at this time.
[0109] By setting a preset training sample threshold, it is possible to ensure that sufficient training samples are provided, avoiding the situation of overfitting or underfitting in model training due to insufficient data, thereby improving the performance of the model.
[0110] Step S203: If satisfied, train the current version of the global water quality prediction model sent by the central cloud server based on the training samples, and adjust the local learning parameters of the model.
[0111] In the embodiment of the present invention, if the edge server conducts local training for the first time, the central cloud server sends the initially constructed global water quality prediction model, and the parameters of this initially global water quality prediction model can be randomly generated. When the first training is actually deployed and applied, for example, the model is retrained every month to update the model parameters, then the central cloud server sends the global water quality prediction model after the previous training.
[0112] Step S204: After reaching the preset number of iterations, send the finally adjusted local learning parameters to the central cloud server, so that the central cloud server aggregates the local learning parameters uploaded by each edge server to obtain global learning parameters, and updates the global water quality prediction model based on the global learning parameters.
[0113] This algorithm ensures that by fairly aggregating the local learning parameters of each edge server on the central cloud server, it can effectively utilize distributed data resources to construct a more accurate and robust global model.
[0114] Step S205: Receive the updated global water quality prediction model sent by the central cloud server.
[0115] In the embodiment of the present invention, if the updated global water quality prediction model has met the requirements, for example, the loss of the model reaches the minimum, it can be directly deployed and applied. If it has not met the requirements, repeat steps S201 - S205 until the training requirements are met.
[0116] Based on the above embodiments, the method further includes:
[0117] If the number of training samples meets the preset training sample threshold, clear the data in the data queue model.
[0118] According to the constructed data queue model, when the preset training sample threshold is met, the data queue model will automatically clear the data and then accumulate data again, which can increase the storage space of this data queue model and achieve infinite loop storage.
[0119] Based on the above embodiments, as Figure 4 shown, the global water quality prediction model consists of a Pre-DBN (Deep Belief Network) encoder, a D-T encoder (DBN-Transformer Encoder, a Transformer encoder fine-tuned based on DBN), and an L-Decoder decoder.
[0120] The global water quality prediction model of the embodiment of the present invention adopts a dual encoder: the encoding layers based on DBN and Transformer respectively: Pre-DBN and DBN-Transformer Encoder, and these two layers of encoders are the core layers of representation learning in the model.
[0121] The Pre-DBN encoder consists of multiple layers of Gaussian-Bernoulli restricted Boltzmann machines and a fully connected output layer, and is used to extract feature vectors in the water quality detection data layer by layer based on the multiple layers of Gaussian-Bernoulli restricted Boltzmann machines, and output the feature vectors based on the fully connected output layer. After passing through the Pre-DBN encoder, the initial local learning parameters of the global water quality prediction model are obtained.
[0122] As Figure 4 shown, the Pre-DBN encoder consists of three layers of GG-RBM (Gated Gaussian Restricted Boltzmann Machine) and a fully connected output layer. By using three or more layers of RBMs to form the encoding part of Pre-DBN, as the depth of the network increases, the features extracted by each layer will be more obvious, that is, the coupling and correlation features of the data between each dimension are stronger.
[0123] As Figure 5 shown, it is the network structure of a single layer of RBM. The input layer of Pre-DBN is an observable random vector i.e., the training sample vector The first hidden random vector of the RBM υ (1) = h (0) ; the second hidden random vector is υ (2) = h (1) ; the third hidden random vector is The final output layer is
[0124] In each RBM, the dimension of the observable variable is D v , and the dimension of the hidden variable is D h , that is, Bias of the observable variables in the visible layer Bias of the hidden layer variables Weight matrix between the two layers Corresponding a i Is υ i Bias of, b j Is h j Bias of, W ij Is υ i To h j Weight of the edge between.
[0125] The energy function plays a central role in the RBM. It is used to define the probability distribution of the configuration between the visible input layer and the hidden layer in the model. Specifically, in the Gaussian-Bernoulli Restricted Boltzmann Machine GG-RBM, the energy function E(v, h) is used to calculate the system energy given the visible layer state v and the hidden layer state h. Given the state v of the visible layer, the energy function can be used for inference, such as predicting or reconstructing the state h of the hidden layer, or vice versa, predicting v given h. This helps with feature extraction, dimensionality reduction, and generating new samples.
[0126] In one example, the energy function E(v, h) of the GG-RBM can be defined as:
[0127]
[0128] Among them, the observable variable υ i And the hidden variable h j The calculation formulas are defined as:
[0129]
[0130] σ in the above formula i And σ j Respectively represent the standard deviations of the noise of the observable variable υ i And the hidden variable h j That satisfy the Gaussian distribution. In application, through the normalization process of the data, σ i And σ j Are both equal to 1. When solving the training, σ i Can be obtained from the original input data as σ i = cov(v) i That is, σ j Then according to Figure 5 The two-layer relationship of, its calculation formula can be defined as: Among them, the meanings of the parameters in formula 5-8 are the same.
[0131] Such as Figure 6As shown, it is the process of solving the parameters of each layer of GG-RBM in the Pre-DBN encoding layer through the CD-k algorithm. The specific process is as follows: For the original data v of the given observable variables (0) By sampling the conditional probability p(h|v (0) ) to obtain the hidden vector h (0) ; Then, using h (0) to calculate the conditional probability p(v|h (0) ) and sampling to obtain v (1) ; Then repeat the above process to obtain p(h|v (1) ) → h (1) , p(v|h (1) ) → v (2) .
[0132] According to the above process and formulas (5)-(7), and the calculation formulas of σ i and σ j , the update formulas for the parameters W, bias a, and bias b of GG-RBM can be obtained as follows:
[0133]
[0134]
[0135] Among them, α > 0 is the learning rate of the model, and at the same time <·> model represents the expected value generated by the model in the l-th sampling process; <·> data represents the expected value calculated based on the training data; v i represents the visible unit; h j represents the hidden unit.
[0136] In the same way, the parameters θ ggr-j = {W ggr_k , a ggr_k , b ggr_k} of the L layers of GG-RBM in the Pre-DBN can be obtained, where k = 1, 2,..., L.
[0137] As a pre-training model, the last layer of the Pre-DBN is a fully connected hidden layer as the output layer, and it is trained through the backpropagation algorithm. Finally, the weights and biases of each layer of the neural network are obtained. Its training process is as follows:
[0138] The first step: Use the CD-k algorithm to perform layer-by-layer pre-training on each layer of GG-RBM from front to back according to formulas (5)-(11) to obtain the parameters θ ggr = {W ggr , a ggr , b ggr}. The iterative calculation formula for training is:
[0139]
[0140] Among them, represents the parameter set of the k-th layer of GG-RBM at the e-th iteration; α is the learning rate, e is the number of iterations of each layer of GG-RBM, k is the layer number where GG-RBM is located, and Δ respectively includes: the update amount of the weights of the model the update amount of the bias of the visible unit v i and the update amount of the bias of the hidden unit h j
[0141] Step 2: After completing the pre-training of GG-RBM, use the fully connected hidden layer of the last layer of Pre-DBN as the output, convert the network structure of GG-RBM into the corresponding neural network structure, and train and refine the network parameters again with a deep neural network.
[0142] The initial parameters of the corresponding network layer are: θ pre-DBN ={{W ggr , b ggr} l}, l = 1, 2, … L, z p is the output of the hidden variable layer corresponding to the last GG-RBM.
[0143] The output of Pre-DBN is The output of the model and the loss function are defined as:
[0144]
[0145]
[0146] Among them, is the predicted value at time step t; z p is the output of the hidden variable layer corresponding to the last GG-RBM; W p is the weight matrix of the model; b p is the bias matrix of the hidden layer; |P| is the number of samples in the pre-training set; is the predicted value of the i-th feature of the model at time point t; x(t, i) ∈ X represents the true value of the input sample, and M is the dimension of the data sequence.
[0147] Step 3: Use the backpropagation algorithm and Adam as the optimizer for backpropagation training to obtain the network parameters of Pre-DBN These are also the initial network parameters of the corresponding L-layer neural network in the encoder D-TEncoder.
[0148] The D-T encoder is used to capture the relationships between feature vectors based on the self-attention mechanism and fine-tune the initial local learning parameters.
[0149] In an embodiment of the present invention, the Transformer multi-layer multi-head attention module is applied to encode the feature vectors output by the Pre-DBN, and the interaction information on each dimension is captured in multiple different projection spaces.
[0150] In a feasible embodiment, the output feature vectors of the Pre-DBN are That is, the input sequence of the D-T Encoder.
[0151] Let the initial input sequence of the D-T Encoder be H (0) , add the positional encoding, and the learnable parameter W pos : Then
[0152]
[0153] Among them,
[0154] Such as Figure 7 As shown, it is a schematic structural diagram of an encoder of the multi-head self-attention. When encoding, through multiple learnable matrices That is, the self-attention model is applied in N projection spaces respectively. The calculation formula of each attention head is as follows:
[0155] MultiHead(H) = W 0 [head 1 ; …, head N (14);
[0156]
[0157]
[0158] Among them, is the output projection matrix; H is the input matrix; D is the dimension of the input feature matrix H; Q n is the query matrix; K n is the key matrix; V n is the value matrix; are the weight matrices corresponding to each matrix respectively; d k is the dimension of the column vectors of the matrices Q n and V n ; d v is the dimension of the column vectors of V n ; N is the number of attention heads.
[0159] From Figure 4 the global water quality prediction model, a total of M Transformer encoders are set, and each encoding layer is calculated through a multi-head self-attention module and a non-linear FNN(·) position-wise feed-forward neural network.
[0160] Figure 7 In, in the layer calculation of each encoder, the residuals (such as Figure 7 the data sequences z′ and v′ in
[0161] FNN(·) is a two-layer fully connected layer connected together with ReLu as the activation function, and its definition is:
[0162] FNN(v′) = W 2FNN ReLu(W 1FNN v′ + b 1FNN ) + b 2FNN (17)
[0163] where v′ ∈ Z (l) is the feature vector at each position in the input sequence of the previous layer, W 1FNN and W 2FNN are the weight matrices of the two-layer neural network, b 1FNN and b 2FNN are the biases corresponding to these two layers of the network, and they are all learnable network parameters. The connection weights of each layer in the encoder D-TEncoder are obtained in the dynamic calculation of the self-attention mechanism.
[0164] The L-Decoder decoder is used to predict the pollution situation of water quality detection data based on the relationship between feature vectors. The pollution situation includes pollution type, pollution area, and pollution time.
[0165] Through the double encoding of Pre-DBN and D-TEncoder, the input water quality detection data sequence is reconstructed through representation learning, and the reconstructed encoding is used as the input of the last fully connected layer for training. Let the vector of the encoding reconstructed by D-TEncoder at time t be The preset time window width is Then all the encodings are concatenated into a vector as a linear layer with network parameters , and τ is the time window width of the output prediction. The output of L-Decoder is defined as:
[0166]
[0167] The meaning of its parameters refers to the meaning of the same symbols above.
[0168] According to the predicted water quality pollution situation in the future τ time to predict the demand, the loss function of the model is defined as:
[0169]
[0170] where y is the actual water quality pollution situation; is the predicted water quality pollution situation.
[0171] Based on the same inventive concept, a training method for a water quality prediction model is provided, which is applied to a central cloud server, as Figure 8 shown, and the method includes the following steps:
[0172] Step S801: Send the current version of the global water quality prediction model to each edge server for local training by each edge server.
[0173] Step S802: Receive the local learning parameters uploaded by each edge server.
[0174] In a feasible implementation, the number of nodes of the edge servers connected to the central cloud server is constrained by a predefined model; the predefined model is as follows:
[0175]
[0176] where the meanings of the parameters in the formula are the same as those in formula 3.
[0177] The maximum number of edge servers connected to the central cloud server at the same time is achieved through this formula 17.
[0178] Step S803: Aggregate the local learning parameters of each edge server based on a preset parameter aggregation formula to obtain global learning parameters; the preset parameter aggregation formula is:
[0179]
[0180] where ω G is the global learning parameter; ω i is the local learning parameter; is the set of edge servers ES j ; represents an intelligent Internet of Things sensor.
[0181] Step S804: Update the global water quality prediction model based on the global learning parameters.
[0182] Step S805: Send the updated global water quality prediction model to each edge server.
[0183] The performance of the local models participating in the aggregation affects the performance of the global model generated by the central cloud. Therefore, to ensure that only the local learning parameters of edge servers with high credibility can participate in the aggregation, based on the above embodiments, before aggregating the local learning parameters of each edge server, the method further includes:
[0184] Calculating the statistical values of the local learning parameters of each edge server, where the statistical values are at least one of the following: mean, median, standard deviation;
[0185] Based on the statistical values of each edge server, mutually verify the credibility of the local learning parameters trained by each other.
[0186] Eliminate the local learning parameters with credibility lower than the preset threshold.
[0187] In one example, if the standard deviation calculated from the local learning parameters of a certain edge sensor exceeds 3 standard deviations compared to other edge servers, it indicates that there is a significant difference between this edge server and other edge servers, and the credibility is low.
[0188] Through mutual verification, ensure that the local learning parameters trained by edge servers with high credibility participate in the aggregation, thereby improving the accuracy of model training.
[0189] Based on the same inventive concept, a water quality prediction method is provided, which is applied to an edge server. As Figure 9 shown, the method includes:
[0190] Step S901: Obtain the water quality detection data uploaded by the intelligent Internet of Things sensors that match it based on the preset matching relationship between the intelligent Internet of Things sensors and the edge server.
[0191] Step S902: Input the water quality detection data into the trained water quality prediction model for water quality prediction to obtain a water pollution prediction result.
[0192] The water quality prediction model is trained by the method of the above embodiments.
[0193] Based on the same technical concept, an embodiment of the present invention further provides an electronic device. As Figure 10 shown, it includes a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004.
[0194] The memory 1003 is used to store a computer program;
[0195] The processor 1001 is configured to implement the steps of the water quality prediction model training and the water quality prediction method when executing the program stored in the memory 1003.
[0196] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0197] The communication interface is used for communication between the above electronic device and other devices.
[0198] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0199] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0200] The computer program product for water quality prediction model training and water quality prediction provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated herein.
[0201] The device for training a water quality prediction model and predicting water quality provided by the embodiments of the present invention may be specific hardware on a device, or software or firmware installed on the device, etc. The device provided by the embodiments of the present invention has the same implementation principle and technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing described systems, devices, and units can all refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0202] In the embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0203] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, the functional units in the embodiments provided by the present invention can be integrated into a processing unit, or each unit exists physically alone, or two or more units can be integrated into one unit.
[0205] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0206] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.
[0207] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All of them should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A training method for a water quality prediction model, characterized in that: Applied to an edge server, the method comprises: A training sample is obtained in each iteration; the training sample is water quality detection data collected in real time by a variety of intelligent Internet of Things sensors; Determining whether the number of the training samples meets a preset training sample threshold; If satisfied, then training the current version of the global water quality prediction model issued by the central cloud server based on the training samples, and adjusting the local learning parameters of the model; After reaching a preset number of iterations, the finally adjusted local learning parameters are sent to the central cloud server, so that the central cloud server aggregates the local learning parameters uploaded by each edge server to obtain the global learning parameters, and updates the global water quality prediction model based on the global learning parameters; Receive the updated global water quality prediction model sent by the central cloud server.
2. The method according to claim 1, characterized in that: Acquiring a training sample once in each iteration includes: At each iteration, water quality test data collected in real time by a variety of smart IoT sensors in the target monitoring area is obtained; The real-time collected water quality test data is stored in a pre-built data queue model; The water quality detection data collected in real time and the existing water quality detection data in the data queue model are jointly determined as training samples.
3. The method according to claim 2, characterized in that The pre-built data queue model is: In the formula, Indicates edge server ES j The size of the water quality test data stored in time step t; Indicates edge server ES j The size of the water quality detection data updated in the next time step; j represents the number of the edge server; It is the edge server ES j A collection of; It stands for Smart IoT Sensor; i represents the number of the smart IoT sensor; m represents the data collected by the smart IoT sensor; S is the set of smart IoT sensors; Represents smart IoT sensors The size of water quality data collected within time step t; y (i,j) [t] is the edge server ES j and smart IoT sensors The scheduling vector between them specifically indicates that within the time step t, the smart IoT sensor Whether it is dispatched to the edge server ES j , if scheduled, then y (i,j) [t]=1, otherwise y (i,j) [t] = 0; I h is the scheduling vector of the data queue model, when I j =1, it means that the amount of data stored in the data queue model is greater than or equal to the preset training sample threshold; when I j =0, it means that the amount of data stored in the data queue model is less than the preset training sample threshold.
4. The method according to claim 2, characterized in that: The method further comprises: If the number of training samples meets the preset training sample threshold, the data in the data queue model is cleared.
5. The method according to claim 1, characterized in that The global water quality prediction model consists of a Pre-DBN encoder, a D-Tencoder encoder and an L-Decoder decoder; The Pre-DBN encoder is composed of a multi-layer Gauss-Bernoulli restricted Boltzmann machine and a fully connected output layer, and is used to extract feature vectors in water quality detection data layer by layer based on the multi-layer Gauss-Bernoulli restricted Boltzmann machine, and output the feature vector based on the fully connected output layer, and obtain the initial local learning parameters of the global water quality prediction model through the Pre-DBN encoder; The D-Tencoder encoder is used to capture the relationship between feature vectors based on a self-attention mechanism and fine-tune the initial local learning parameters; The L-Decoder is used to predict the pollution situation of water quality detection data based on the relationship between feature vectors, and the pollution situation includes pollution type, pollution area and pollution time.
6. A training method for a water quality prediction model, characterized in that: Applied to a central cloud server, the method comprises: Send the current version of the global water quality prediction model to each edge server for local training; Receive local learning parameters uploaded by each edge server; The local learning parameters of each edge server are aggregated based on a preset parameter aggregation formula to obtain a global learning parameter; the preset parameter aggregation formula is: Among them, ω G is the global learning parameter; ω i for local learning parameters; It is the edge server ES j A collection of; It stands for Smart IoT Sensor; Updating the global water quality prediction model based on the global learning parameters; The updated global water quality prediction model is sent to each edge server.
7. The method according to claim 6, characterized in that Before aggregating the local learning parameters of each edge server, the method further includes: Calculate the statistical value of the local learning parameter of each edge server, where the statistical value is at least one of the following: mean, median, and standard deviation; Based on the statistical values of each edge server, the credibility of the local learning parameters trained by each server is verified; Eliminate local learning parameters whose credibility is lower than a preset threshold.
8. A water quality prediction method, characterized in that: Applied to an edge server, the method comprises: Based on the preset matching relationship between the smart IoT sensor and the edge server, the water quality detection data uploaded by the smart IoT sensor matching the smart IoT sensor is obtained; The water quality detection data is input into the water quality prediction model trained by any one of the methods described in claims 1-5 to perform water quality prediction and obtain a water pollution prediction result.
9. A water quality monitoring system, characterized in that: The system includes distributed intelligent IoT sensors, distributed edge servers and a central cloud server; The smart IoT sensor is used to collect water quality detection data, and the smart IoT sensor includes at least a temperature sensor, a pH sensor, a dissolved oxygen sensor, a total organic carbon sensor, a total nitrogen sensor, a chlorophyll-a sensor, and a turbidity sensor; The edge server is used to collect water quality detection data collected by the intelligent Internet of Things sensor, train a global water quality prediction model based on the collected water quality detection data, and predict water pollution based on the trained global water quality prediction model; The central cloud server is used to update the global learning parameters of the global water quality prediction model and send the updated global water quality prediction model to each edge server.
10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.