Water quality prediction method, device and equipment and readable storage medium

Through variational modal decomposition and bidirectional long and short-term memory network combined with differential autoregressive moving average model, parameters are optimized to solve the problems of low efficiency and accuracy of water quality prediction, and more efficient and accurate water quality prediction is achieved.

CN120277991APending Publication Date: 2025-07-08INTELLIGENT TECH CO LTD OF CHINESE CONSTR THIRD ENG BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242946.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing water quality prediction algorithm model is inefficient and has low accuracy, and cannot effectively predict the trend of water quality change.

Method used

The water quality data is decomposed by a variational modal decomposition algorithm, combined with a two-way long and short-term memory network and a differential autoregressive moving average model for prediction, and the parameters are optimized through the NGO algorithm, and non-essential features are eliminated and weighted fusion is performed.

Benefits of technology

It improves the accuracy and efficiency of water quality prediction, can better identify the interactions of water quality parameters in complex scenarios, reduce noise impact, and improve model performance and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277991A_ABST
    Figure CN120277991A_ABST
Patent Text Reader

Abstract

The invention discloses a water quality prediction method, device and equipment and a readable storage medium. The water quality prediction method comprises the following steps: preprocessing a time sequence formed by multiple pieces of water quality data to obtain a first input sequence; decomposing the first input sequence by using a variational mode decomposition algorithm to obtain a plurality of mode components, wherein an optimal parameter of the variational mode decomposition algorithm is determined through an NGO algorithm during training; and inputting the plurality of modal components into a bidirectional long-short-term memory network to obtain first water quality prediction data. According to the method, the input data is decomposed through the variational mode decomposition algorithm, on one hand, the dimension of the data processed by the bidirectional long-short-term memory network can be reduced, the performance of the model can be improved on the whole, and on the other hand, noise in the input data can be reduced; the bidirectional long-short-term memory network considers past and future context information at each time step at the same time, so that compared with a traditional unidirectional long-short-term memory network, the prediction of the time sequence has higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of water quality monitoring, and particularly to a water quality prediction method, device, equipment, and readable storage medium. Background Art

[0002] Smart communities are the basic and important components of smart cities. They are new community forms established by using modern technologies such as digitization, informatization, and intelligence. More and more devices upload the collected data to the community Internet of Things platform through technical communication methods, such as water quality sensors. By using an algorithm model to predict the changing trend of water quality based on the water quality data collected by the water quality sensors, measures can be taken in advance to avoid problems caused by water quality deterioration.

[0003] However, the existing algorithm models for water quality prediction have problems of low efficiency and low accuracy. Summary of the Invention

[0004] This application provides a water quality prediction method, device, equipment, and readable storage medium, aiming to solve the technical problems of low efficiency and low accuracy existing in the current algorithm models for water quality prediction.

[0005] In a first aspect, an embodiment of this application provides a water quality prediction method, and the water quality prediction method includes:

[0006] Preprocess a time series composed of multiple water quality data to obtain a first input sequence;

[0007] Use the variational mode decomposition algorithm to decompose the first input sequence into multiple mode components, where the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training;

[0008] Input the multiple mode components into a bidirectional long short-term memory network to obtain first water quality prediction data.

[0009] Optionally, after the step of inputting the multiple mode components into a bidirectional long short-term memory network to obtain first water quality prediction data, it includes:

[0010] Input the first input sequence into an autoregressive integrated moving average model to obtain second water quality prediction data;

[0011] Sum the first water quality prediction data and the second water quality prediction data after multiplying them by a first weight coefficient and a second weight coefficient respectively to obtain final water quality prediction data.

[0012] Optionally, before the step of summing the first water quality prediction data and the second water quality prediction data after multiplying them by a first weight coefficient and a second weight coefficient respectively to obtain final water quality prediction data, it includes:

[0013] Preprocess the time series composed of multiple historical water quality data and construct a water quality data training set, where the water quality data training set includes multiple groups of second input sequences, and each group of second input sequences includes multiple preprocessed historical water quality data and corresponding true output water quality data;

[0014] Use each group of second input sequences to train a bidirectional long short-term memory network to obtain the first training prediction value of each group of second input sequences;

[0015] Use each group of second input sequences to train an autoregressive integrated moving average model to obtain the second training prediction value of each group of second input sequences;

[0016] Multiply the first training prediction value and the second training prediction value of each group of second input sequences by different third weight coefficients and fourth weight coefficients respectively and sum them to obtain the final training prediction value of each group of second input sequences;

[0017] Calculate the mean square error between the final training prediction values of all groups of second input sequences and the corresponding true output water quality data;

[0018] Determine the third weight coefficient and the fourth weight coefficient corresponding to the minimum mean square error as the first weight coefficient and the second weight coefficient.

[0019] Optionally, the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training, including:

[0020] Generate an initial population, where the initial population includes multiple individuals, and each individual includes the number of mode components and the penalty parameter, and the number of mode components and the penalty parameter are random values within a preset range;

[0021] Use each individual as the parameter of the variational mode decomposition algorithm, decompose the second input sequence using the variational mode decomposition algorithm to obtain multiple mode components, and input the multiple mode components into the bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0022] Based on the second training prediction value of each individual and the corresponding true output water quality data, calculate the fitness of the population through formula one;

[0023] If the number of adjustment times is less than or equal to the preset number of times, adjust the individual through formula two and return to execute the step of using each individual as the parameter of the variational mode decomposition algorithm, decomposing the second input sequence using the variational mode decomposition algorithm to obtain multiple mode components, and inputting the multiple mode components into the bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0024] If the number of adjustments is greater than the preset number, the number of modal components and the numerical value of the penalty parameter of the individual with the smallest fitness loss function in the population are used as the optimal parameters of the variational mode decomposition algorithm;

[0025] The first formula is as follows:

[0026] The second formula is as follows:

[0027] where MSE is the fitness loss function of the population, N is the number of individuals, y i is the true output water quality data corresponding to the i-th individual, is the second training prediction value of the i-th individual, is the position of the i-th individual in the (t + 1)-th adjustment, is the position of the i-th individual in the t-th adjustment, ω is the inertia coefficient, used to control the influence of the previous adjustment speed, is the historical best position of the i-th individual, c1 and c2 are acceleration constants, used to control the speed of the individual moving towards the individual optimum and the global optimum, r1 and r2 are random numbers within the range of [0, 1], is the global best position among all individuals.

[0028] Optionally, each water quality data includes multiple features and corresponding feature data. Before using each group of second input sequences to train the bidirectional long short-term memory network to obtain the first training prediction value of each group of second input sequences, it includes:

[0029] If the number of features is greater than the preset number, the importance score of each feature is calculated using the water quality data training set through a linear regression model;

[0030] The feature with the lowest importance score and the corresponding feature data are removed;

[0031] If the number of features after removal is greater than the preset number, the water quality data training set after removal is used as the new water quality data training set, and the step of calculating the importance score of each feature using the water quality data training set through the linear regression model is returned.

[0032] Optionally, the preprocessing of the time series composed of multiple water quality data to obtain the first input sequence includes:

[0033] Detect outliers and missing values in the time series composed of multiple water quality data;

[0034] Replace outliers and missing values with the mean or median of multiple water quality data;

[0035] Normalize each piece of substituted water quality data to obtain a first input sequence.

[0036] Optionally, detecting outliers in the time series composed of multiple pieces of water quality data includes:

[0037] For each piece of water quality data, use the water quality data, the mean and standard deviation of multiple pieces of water quality data to calculate the standard score of the water quality data through Formula 3. If the absolute value of the standard score of the water quality data is greater than the preset score, then determine the water quality data as an outlier. The formula 3 is:

[0038]

[0039] Where Z is the standard score of the water quality data, x is the water quality data, μ is the mean of multiple pieces of water quality data, and σ is the standard deviation of multiple pieces of water quality data.

[0040] In a second aspect, an embodiment of the present application provides a water quality prediction device, and the water quality prediction device includes:

[0041] A preprocessing module for preprocessing the time series composed of multiple pieces of water quality data to obtain a first input sequence;

[0042] A decomposition module for decomposing the first input sequence using the variational mode decomposition algorithm to obtain multiple modal components, where the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training;

[0043] A prediction module for inputting the multiple modal components into a bidirectional long short-term memory network to obtain first water quality prediction data.

[0044] In a third aspect, an embodiment of the present application provides a water quality prediction device, and the water quality prediction device includes a processor, a memory, and a water quality prediction program stored on the memory and executable by the processor. When the water quality prediction program is executed by the processor, the steps of the water quality prediction method as described above are implemented.

[0045] In a fourth aspect, an embodiment of the present application provides a readable storage medium, and a water quality prediction program is stored on the readable storage medium. When the water quality prediction program is executed by a processor, the steps of the water quality prediction method as described above are implemented.

[0046] The beneficial effects brought by the technical solutions provided by the embodiments of the present application include:

[0047] In the embodiments of the present application, a time series composed of multiple water quality data is preprocessed to obtain a first input sequence; the variational mode decomposition algorithm is used to decompose the first input sequence to obtain multiple mode components, where the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training; the multiple mode components are input into a bidirectional long short-term memory network to obtain first water quality prediction data. Through the embodiments of the present application, the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training, which has a high convergence accuracy and good stability. By decomposing the input data through the variational mode decomposition algorithm, on the one hand, it can reduce the dimension of the data processed by the bidirectional long short-term memory network, and can improve the performance of the model as a whole. On the other hand, it can reduce the noise in the input data and reduce the impact of unnecessary water quality characteristics on water quality prediction, and is more suitable for scenarios where it is necessary to accurately identify the complex interactions between water quality parameters. The bidirectional long short-term memory network has a higher accuracy in predicting time series compared to the traditional unidirectional long short-term memory network because it considers the context information of the past and future at each time step. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is the first process schematic diagram of an embodiment of the water quality prediction method of the present application;

[0049] Figure 2 It is the second process schematic diagram of an embodiment of the water quality prediction method of the present application;

[0050] Figure 3 For the present application Figure 1 It is the refined process schematic diagram of step S10;

[0051] Figure 4 It is the third process schematic diagram of an embodiment of the water quality prediction method of the present application;

[0052] Figure 5 It is the fourth process schematic diagram of an embodiment of the water quality prediction method of the present application;

[0053] Figure 6 It is the functional module schematic diagram of an embodiment of the water quality prediction device of the present application;

[0054] Figure 7 It is the hardware structure schematic diagram of the water quality prediction device involved in the solution of the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0056] To make the purpose, technical solutions and advantages of this application clearer, the embodiments of this application will be further described in detail below in conjunction with the accompanying drawings.

[0057] In a first aspect, an embodiment of this application provides a water quality prediction method.

[0058] In one embodiment, referring to Figure 1 , Figure 1 is the first process schematic diagram of an embodiment of the water quality prediction method of this application. As Figure 1 shows, the water quality prediction method includes:

[0059] Step S10: Preprocess the time series composed of multiple water quality data to obtain a first input sequence.

[0060] In this embodiment, the community water quality data at multiple moments can be collected through a water quality sensor, and communicated and uploaded to the community Internet of Things platform based on Modbus (a serial communication protocol). Further, the water quality data is preprocessed and the water quality is predicted. The community water quality data at multiple collected moments is the time series composed of multiple water quality data that needs to be preprocessed. Each piece of water quality data may include, for example, pH value, conductivity, dissolved oxygen, turbidity, temperature, residual chlorine, suspended solids, and collection timestamp, etc.

[0061] Step S20: Use the variational mode decomposition algorithm to decompose the first input sequence to obtain multiple modal components, where the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training.

[0062] In this embodiment, VMD (Variational Mode Decomposition) is an adaptive and completely non-recursive method for modal variational and signal processing. Its adaptability is manifested in determining the number of modal decompositions of a given sequence according to the actual situation. During the subsequent search and solution process, it can adaptively match the optimal center frequency and finite bandwidth of each mode, and can effectively separate IMFs (Intrinsic Mode Functions, i.e., modal components), divide the frequency domain of the signal, and then obtain the effective components of the given signal, and finally obtain the optimal solution of the variational problem. The parameters of the variational mode decomposition algorithm include the number of modal components and the penalty parameter. The optimal parameters of the variational mode decomposition algorithm are determined by the NGO (Nonlinear Granger Optimizer) algorithm during training, which has high convergence accuracy and good stability. Decomposing the input data through the variational mode decomposition algorithm can, on the one hand, reduce the dimension of the data processed by the bidirectional long short-term memory network, improve the performance of the model as a whole, and on the other hand, reduce the noise in the input data, reduce the impact of unnecessary water quality characteristics on water quality prediction, and is more suitable for scenarios that require precise identification of the complex interactions between water quality parameters. Among them, the goal of VMD is to decompose the first input sequence f(t) into K modal components uk(t), and the center frequency of each modal component is w k , and its calculation formula is The first term among them represents the bandwidth constraint of each modal component uk(t), δ(t) is the Dirac delta function, is the kernel of the Hilbert transform, and * represents the convolution operation, represents the time derivative, w k is the center frequency of the k-th modal component, and the second term λ represents the reconstruction error, ensuring that the sum of the decomposed modal components is equal to the original signal f(t). λ represents the regularization parameter, which is used to balance the broadband constraint and the reconstruction error. The third term means that the sum of all modal components must be equal to the original signal f(t).

[0063] Step S30: Input multiple modal components into the bidirectional long short-term memory network to obtain the first water quality prediction data.

[0064] In this embodiment, the BiLSTM (Bidirectional Long Short-Term Memory) is composed of LSTM units in two directions. Among them, the forward LSTM unit processes the input sequence from front to back, and the backward LSTM unit processes the input sequence from back to front. Since the bidirectional long short-term memory network considers the context information of the past and the future simultaneously at each time step, it has a higher accuracy rate for predicting time series compared to the traditional unidirectional long short-term memory network.

[0065] In this embodiment, the community water quality data at multiple moments can be collected by a water quality sensor, communicated and uploaded to the community Internet of Things platform based on Modbus, and further preprocessed and water quality prediction is carried out. The community water quality data at multiple collected moments is a time series composed of multiple water quality data to be preprocessed. The optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training, which has a high convergence accuracy and good stability. Decomposing the input data by the variational mode decomposition algorithm can, on the one hand, reduce the dimension of the data processed by the bidirectional long short-term memory network, improve the performance of the model as a whole, and on the other hand, reduce the noise in the input data and reduce the impact of unnecessary water quality features on water quality prediction. It is more suitable for scenarios that require accurate identification of the complex interactions between water quality parameters. The bidirectional long short-term memory network is composed of LSTM units in two directions. Among them, the forward LSTM unit processes the input sequence from front to back, and the backward LSTM unit processes the input sequence from back to front. Since the bidirectional long short-term memory network considers the context information of the past and the future simultaneously at each time step, it has a higher accuracy rate for predicting time series compared to the traditional unidirectional long short-term memory network.

[0066] Further, in one embodiment, referring to Figure 2 , Figure 2 is the second process schematic diagram of an embodiment of the water quality prediction method of this application. As shown in Figure 2 , after step S30, it includes:

[0067] Step S40, input the first input sequence into the autoregressive integrated moving average model to obtain the second water quality prediction data;

[0068] Step S50, multiply the first water quality prediction data and the second water quality prediction data by the first weight coefficient and the second weight coefficient respectively and sum them to obtain the final water quality prediction data.

[0069] In this embodiment, ARIMA (Autoregressive Integrated Moving Average model) consists of three parts: autoregressive term (AR), differencing term (I), and moving average term (MA), corresponding to the influence of historical values, the influence of random error terms, and the influence of past random error terms respectively. The ARIMA model can effectively handle trends and periodic changes in time series, and can well cope with the linear trends and seasonal components usually contained in water quality data. Further, the water quality prediction data of the BiLSTM and ARIMA models are weighted and fused to obtain the final water quality prediction data, so that the overall model can not only capture complex nonlinear relationships, but also handle linear trends and seasonal patterns, improving the generalization ability of the model.

[0070] Further, in one embodiment, before step S50, it includes:

[0071] Preprocess the time series composed of multiple historical water quality data and construct a water quality data training set, where the water quality data training set includes multiple groups of second input sequences, and each group of second input sequences includes multiple preprocessed historical water quality data and corresponding true output water quality data;

[0072] Use each group of second input sequences to train the bidirectional long short-term memory network to obtain the first training prediction value of each group of second input sequences;

[0073] Use each group of second input sequences to train the autoregressive integrated moving average model to obtain the second training prediction value of each group of second input sequences;

[0074] Multiply the first training prediction value and the second training prediction value of each group of second input sequences by different third weight coefficients and fourth weight coefficients respectively and sum them to obtain the final training prediction value of each group of second input sequences;

[0075] Calculate the mean square error between the final training prediction values of all groups of second input sequences and the corresponding true output water quality data;

[0076] Determine the third weight coefficient and the fourth weight coefficient corresponding to the minimum mean square error as the first weight coefficient and the second weight coefficient.

[0077] In this embodiment, the time series composed of multiple historical water quality data is, for example, the community water quality data at multiple moments (such as every hour) within a period of time (such as one month) collected by a water quality sensor. A set of second input sequences in the constructed water quality data training set includes, for example, 5 preprocessed historical water quality data (such as water quality data for 5 consecutive hours) and the corresponding true output water quality data (such as the water quality data for the next hour that the BiLSTM model needs to predict). It is easy to understand that the input samples and prediction targets can be specifically set according to the prediction requirements by using the sliding window technique. By continuously moving the sliding window, a series of input samples and prediction targets are generated according to the prediction requirements.

[0078] In this embodiment, when training the bidirectional long short-term memory network using each set of second input sequences, let the input sequence be X = {x1, x2,... x T}, and the formula calculation of LSTM is i t = σ(W xi x t + W hi h t-1 + b i ), f t = σ(W xf x t + W hf h t-1 + b f ), c t = f t ° c t-1 + i t ° tanh(W xc x t + W hc h t-1 + b c ), o t = σ(W xo x t + W ho h t-1 + b o ), h t = o t ° tanh(c t ), where i t , f t , c t , o t , h t respectively represent the input gate, forget gate, cell state, output gate, and hidden state, σ is the sigmoid function, ° represents element-wise multiplication, W and b are the weight matrix and bias vector, and tanh represents the hyperbolic tangent function. The output of the forward LSTM at time t is The input sequence of the forward LSTM is processed in chronological order 1, 2, 3…,T, and the input sequence of the backward LSTM is processed in reverse chronological order T, T-1,…1. Finally, the output h of the BiLSTM at each time step t i is the concatenation of the outputs of the forward LSTM and the backward LSTM wherein, denotes the concatenation with .

[0079] In this embodiment, when using each group of second input sequences to train the autoregressive integrated moving average model, the ARIMA algorithm model consists of three parts: autoregressive term (AR), differencing term (I), and moving average term (MA). The calculation formula of the autoregressive term (AR) is wherein, X t is the value of the time series at time t, c is the constant item, is the autoregressive coefficient, ∈ t is the white noise error term. The calculation formula of the differencing term (I) is Y t =(1 - B) d X t , wherein, B is the lag operator, BX t =X t-1 , (1 - B)d represents the d-th order difference. The calculation formula of the moving average term (MA) is X t =μ + ∈ t +θ1∈ t-1 +θ2∈ t-2 +…+θ q ∈ t-q , wherein, μ is the constant term, θ q is the moving average coefficient, ∈ t is the white noise error term.

[0080] In this embodiment, the first training prediction value and the second training prediction value of each group of second input sequences are multiplied by different third weight coefficients and fourth weight coefficients respectively and then summed to obtain the final training prediction value of each group of second input sequences. Its weighted fusion calculation formula can be expressed as wherein, ω3 and ω4 represent the third weight coefficient and the fourth weight coefficient respectively, satisfying ω3 + ω4 = 1, is the final training prediction value, is the first training prediction value obtained by processing through the VMD, NGO and BiLSTM models, is the second training prediction value obtained by the ARIMA model. The optimal weights ω3 and ω4 are found through the genetic algorithm to minimize the prediction error and used as the first weight coefficient and the second weight coefficient for the final prediction. The specific method is Among them, is the real output water quality data, and MSE is the mean square error. Further, in one embodiment, the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training, including:

[0081] Generate an initial population, the initial population includes multiple individuals, each individual includes the number of modal components and the penalty parameter, where the number of modal components and the penalty parameter are random values within a preset range;

[0082] Using each individual as the parameters of the variational mode decomposition algorithm, decompose the second input sequence using the variational mode decomposition algorithm to obtain multiple modal components, and input the multiple modal components into the bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0083] Based on the second training prediction value of each individual and the corresponding real output water quality data, calculate the fitness of the population through Formula 1;

[0084] If the number of adjustment times is less than or equal to the preset number of times, adjust the individual through Formula 2 and return to execute the step of using each individual as the parameters of the variational mode decomposition algorithm, decomposing the second input sequence using the variational mode decomposition algorithm to obtain multiple modal components, and inputting the multiple modal components into the bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0085] If the number of adjustment times is greater than the preset number of times, then use the numerical values of the number of modal components and the penalty parameter of the individual with the minimum fitness loss function in the population as the optimal parameters of the variational mode decomposition algorithm;

[0086] The Formula 1 is:

[0087] The Formula 2 is:

[0088] Among them, MSE is the fitness loss function of the population, N is the number of individuals, y i is the real output water quality data corresponding to the i-th individual, is the second training prediction value of the i-th individual, is the position of the i-th individual in the (t + 1)-th adjustment, is the position of the i-th individual in the t-th adjustment, ω is the inertia coefficient, used to control the influence of the previous adjustment speed, is the historical best position of the i-th individual, c1 and c2 are acceleration constants, used to control the speed of the individual moving towards the individual optimum and the global optimum, r1 and r2 are random numbers within the range of [0, 1], is the global best position among all individuals.

[0089] In this embodiment, the NGO (Nonlinear Granger Optimizer) algorithm is used to determine the optimal values of the number of modal components and the penalty parameter of the variational mode decomposition algorithm. Specifically, an initial population is first generated. For example, the initial population includes 50 individuals, and the number of modal components and the penalty parameter of each individual are random values within a preset range. The values of the number of modal components and the penalty parameter in the 50 individuals are respectively used as the parameters of the variational mode decomposition algorithm. The variational mode decomposition algorithm is used to decompose the second input sequence to obtain multiple modal components. The multiple modal components are input into a bidirectional long short-term memory network for training to obtain the second training prediction values of each individual, a total of 50 training prediction values of individuals. Based on the training set, the true output water quality data corresponding to the training prediction values of these 50 individuals can be obtained. Furthermore, the fitness of the population is calculated through Formula 1 based on the true output water quality data corresponding to the training prediction values of these 50 individuals. The fitness of the population can be reflected by the mean square error in Formula 1. If the number of adjustment times is less than or equal to the preset number of times, the individuals are adjusted through Formula 2, and the adjusted individuals are used as the parameters of the variational mode decomposition algorithm. The second input sequence is continued to be decomposed to obtain multiple modal components, and the adjusted second training prediction values are obtained through the bidirectional long short-term memory network. Furthermore, the fitness of the population is recalculated until the number of adjustment times is greater than the preset number of times. Then, the values of the number of modal components and the penalty parameter of the individual with the minimum fitness loss function in the population are used as the optimal parameters of the variational mode decomposition algorithm. Thus, it can be seen that through the NGO, the optimal values can be determined from a large number of values of the number of modal components and the penalty parameter, and the optimal parameters can be determined for the variational mode decomposition algorithm, with high convergence accuracy and good stability.

[0090] Further, in one embodiment, each water quality data includes multiple features and corresponding feature data. Before using each group of second input sequences to train the bidirectional long short-term memory network to obtain the first training prediction value of each group of second input sequences, it includes:

[0091] If the number of features is greater than the preset number, the importance score of each feature is calculated through a linear regression model using the water quality data training set;

[0092] The feature with the lowest importance score and the corresponding feature data are removed;

[0093] If the number of features after removal is greater than the preset number, the water quality data training set after removal is used as the new water quality data training set, and the step of calculating the importance score of each feature through the linear regression model using the water quality data training set is returned and executed.

[0094] In this embodiment, as described above, the water quality data includes, for example, pH value, conductivity, dissolved oxygen, turbidity, temperature, residual chlorine, and suspended solids. Among them, the pH value, conductivity, dissolved oxygen, turbidity, temperature, residual chlorine, and suspended solids are respectively water quality characteristics. There are a total of 7 water quality characteristics. If the number of characteristics is greater than the preset number (the preset number is 5 for example), the linear regression model is used to eliminate the characteristics of the water quality data, and the relatively unimportant water quality characteristics and the corresponding characteristic data in the water quality data are eliminated to improve the processing performance of the variational mode decomposition algorithm, the bidirectional long short-term memory network, and the autoregressive integrated moving average model. Specifically, set the initial feature set F = {f1, f2, …, fn}, and use all features to train the linear regression model. Its calculation formula is y = β0 + β1f1 + β2f2 + … + β n f n + ∈, where β0 is the intercept term, β i is the coefficient of the feature f n , and ∈ is the error term. Calculate the absolute value of the coefficient of each feature as the importance score, Importance Score(fi) = ∣βi∣, and select the feature with the lowest importance score for elimination until the number of remaining features is not greater than the preset number 5.

[0095] Further, in one embodiment, referring to Figure 3 , Figure 3 is the detailed flowchart of step S10 in this application. As Figure 1 shown, step S10 includes: Figure 3 shown, step S10 includes:

[0096] Step S101, detecting outliers and missing values in the time series composed of multiple water quality data;

[0097] Step S102, replacing the outliers and missing values with the mean or median of multiple water quality data;

[0098] Step S103, normalizing each piece of water quality data after replacement to obtain the first input sequence.

[0099] In this embodiment, normalizing each piece of water quality data after replacement can be performed using the formula where, is the value of each piece of initial water quality data after normalization, x represents each piece of initial water quality data, x min represents the minimum value among multiple water quality data, and x max represents the maximum value among multiple water quality data.

[0100] Further, in one embodiment, the detection of outliers in the time series composed of multiple water quality data includes:

[0101] For each piece of water quality data, the standard score of the water quality data is calculated using the water quality data, the mean and standard deviation of multiple pieces of water quality data through Formula 3. If the absolute value of the standard score of the water quality data is greater than the preset score, the water quality data is determined to be an outlier. Formula 3 is as follows:

[0102]

[0103] Among them, Z is the standard score of the water quality data, x is the water quality data, μ is the mean of multiple pieces of water quality data, and σ is the standard deviation of multiple pieces of water quality data.

[0104] In this embodiment, the standard score, that is, the Z-score, is a dimensionless value used to describe the deviation degree of a data point relative to the average value of the data set. Its value may be positive, negative or zero. Therefore, a preset score for an acceptable deviation degree can be set in advance. If the absolute value of the calculated standard score of the water quality data is greater than this preset score, this piece of water quality data is determined to be an outlier.

[0105] In this embodiment, referring to Figure 4 , Figure 4 is the schematic diagram of the third process of an embodiment of the water quality prediction method of the present application. As shown in Figure 4 , multiple water quality sensors can be arranged to collect water quality data at multiple locations, and an overall model including VMD-NGO-BiLSTM-ARIMA is established for final water quality prediction. Referring to Figure 5 , Figure 5 is the schematic diagram of the fourth process of an embodiment of the water quality prediction method of the present application. As shown in Figure 5 , it is the specific processing process of the VMD-NGO-BiLSTM-ARIMA overall model, thus realizing full-process coverage: from the collection of the water quality data set to the preprocessing of the data set, using VMD to decompose the data set, determining the optimal parameters of VMD (through the NGO algorithm), then using the optimal parameters of VMD to decompose the data set, and then to deep learning modeling (BiLSTM) and statistical model (ARIMA). The whole process is seamlessly connected to form a complete prediction chain; automation: once the overall model framework is established and trained, it can automatically receive new water quality data, and after a series of processing steps, directly output the prediction result without manual intervention; integration of multiple technologies: by combining signal processing, deep learning and statistical models, this framework can integrate the advantages of different methods to improve the accuracy and robustness of prediction; flexibility: although it is an end-to-end framework, each component is modular and can be adjusted and optimized according to specific application requirements.

[0106] In the second aspect, the embodiments of the present application also provide a water quality prediction device.

[0107] In one embodiment, referring toFigure 6 , Figure 6 is a schematic diagram of the functional modules of an embodiment of the water quality prediction device of the present application. As Figure 6 shown, the water quality prediction device includes:

[0108] A preprocessing module 10 for preprocessing a time series composed of multiple water quality data to obtain a first input sequence;

[0109] A decomposition module 20 for decomposing the first input sequence using the variational mode decomposition algorithm to obtain multiple modal components, wherein the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training;

[0110] A prediction module 30 for inputting multiple modal components into a bidirectional long short-term memory network to obtain first water quality prediction data.

[0111] Further, in one embodiment, the water quality prediction device further includes a final prediction module for:

[0112] Inputting the first input sequence into an autoregressive integrated moving average model to obtain second water quality prediction data;

[0113] Multiplying the first water quality prediction data and the second water quality prediction data by a first weight coefficient and a second weight coefficient respectively and summing them to obtain final water quality prediction data.

[0114] Further, in one embodiment, the water quality prediction device further includes a weight determination module for:

[0115] Preprocessing a time series composed of multiple historical water quality data and constructing a water quality data training set, wherein the water quality data training set includes multiple groups of second input sequences, and each group of second input sequences includes multiple preprocessed historical water quality data and corresponding real output water quality data;

[0116] Training a bidirectional long short-term memory network using each group of second input sequences to obtain a first training prediction value for each group of second input sequences;

[0117] Training an autoregressive integrated moving average model using each group of second input sequences to obtain a second training prediction value for each group of second input sequences;

[0118] Multiplying the first training prediction value and the second training prediction value of each group of second input sequences by different third weight coefficients and fourth weight coefficients respectively and summing them to obtain a final training prediction value for each group of second input sequences;

[0119] Calculating the mean square error between the final training prediction values of all groups of second input sequences and the corresponding real output water quality data;

[0120] Determine the third weight coefficient and the fourth weight coefficient corresponding to the minimum mean square error as the first weight coefficient and the second weight coefficient.

[0121] Further, in one embodiment, the water quality prediction device further includes an optimal parameter determination module, configured to:

[0122] Generate an initial population, where the initial population includes a plurality of individuals, and each individual includes the number of modal components and a penalty parameter, and the number of modal components and the penalty parameter are random values within a preset range;

[0123] Using each individual as the parameter of the variational mode decomposition algorithm, decompose the second input sequence using the variational mode decomposition algorithm to obtain a plurality of modal components, and input the plurality of modal components into a bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0124] Based on the second training prediction value of each individual and the corresponding true output water quality data, calculate the fitness of the population through Formula 1;

[0125] If the number of adjustment times is less than or equal to the preset number of times, adjust the individual through Formula 2, and return to execute the step of using each individual as the parameter of the variational mode decomposition algorithm, decomposing the second input sequence using the variational mode decomposition algorithm to obtain a plurality of modal components, and inputting the plurality of modal components into a bidirectional long short-term memory network for training to obtain the second training prediction value of each individual;

[0126] If the number of adjustment times is greater than the preset number of times, use the values of the number of modal components and the penalty parameter of the individual with the minimum fitness loss function in the population as the optimal parameters of the variational mode decomposition algorithm;

[0127] The Formula 1 is:

[0128] The Formula 2 is:

[0129] Where, MSE is the fitness loss function of the population, N is the number of individuals, y i is the true output water quality data corresponding to the i-th individual, is the second training prediction value of the i-th individual, is the position of the i-th individual in the (t + 1)-th adjustment, is the position of the i-th individual in the t-th adjustment, ω is the inertia coefficient, used to control the influence of the previous adjustment speed, is the historical best position of the i-th individual, c1 and c2 are acceleration constants, used to control the speed of the individual moving towards the individual optimum and the global optimum, r1 and r2 are random numbers within the range of [0, 1], is the global best position among all individuals.

[0130] Further, in one embodiment, the water quality prediction device further includes a feature elimination module, which is used for:

[0131] If the number of features is greater than the preset number, use the water quality data training set to calculate the importance score of each feature through a linear regression model;

[0132] Eliminate the feature with the lowest importance score and the corresponding feature data;

[0133] If the number of features after elimination is greater than the preset number, use the water quality data training set after elimination as the new water quality data training set, and return to execute the step of calculating the importance score of each feature through the linear regression model using the water quality data training set.

[0134] Further, in one embodiment, the preprocessing module 10 includes:

[0135] A detection unit, which is used to detect outliers and missing values in the time series composed of multiple water quality data;

[0136] A substitution unit, which is used to substitute outliers and missing values with the mean or median of multiple water quality data;

[0137] A normalization unit, which is used to normalize each water quality data after substitution to obtain a first input sequence.

[0138] Further, in one embodiment, the detection unit is further used for:

[0139] For each water quality data, use the water quality data, the mean and standard deviation of multiple water quality data to calculate the standard score of the water quality data through Formula 3. If the absolute value of the standard score of the water quality data is greater than the preset score, determine the water quality data as an outlier. Formula 3 is:

[0140]

[0141] Among them, Z is the standard score of the water quality data, x is the water quality data, μ is the mean of multiple water quality data, and σ is the standard deviation of multiple water quality data.

[0142] Among them, the function implementation of each module in the above water quality prediction device corresponds to each step in the above water quality prediction method embodiment, and its function and implementation process will not be elaborated here one by one.

[0143] In the third aspect, an embodiment of the present application provides a water quality prediction device.

[0144] Refer to Figure 7 , Figure 7This is a schematic diagram of the hardware structure of the water quality prediction device involved in the solution of the embodiment of the present application. In the embodiment of the present application, the water quality prediction device may include a processor, a memory, a communication interface, and a communication bus.

[0145] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0146] The communication interface includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of internal devices of the water quality prediction device, as well as interfaces for implementing the interconnection between the water quality prediction device and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display screen (Display), a keyboard (Keyboard), etc.

[0147] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0148] The processor can be a general-purpose processor, and the general-purpose processor can call the water quality prediction program stored in the memory and execute the water quality prediction method provided by the embodiment of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the water quality prediction program is called can refer to the various embodiments of the water quality prediction method of the present application and will not be elaborated here.

[0149] Those skilled in the art can understand that Figure 7 the hardware structure shown in does not constitute a limitation to the present application, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0150] Fourthly, the embodiment of the present application also provides a readable storage medium.

[0151] The water quality prediction program is stored on the readable storage medium of the present application. When the water quality prediction program is executed by the processor, the steps of the water quality prediction method as described above are implemented.

[0152] Among them, the method implemented when the water quality prediction program is executed can refer to the various embodiments of the water quality prediction method of this application, which will not be elaborated here.

[0153] It should be noted that the serial numbers of the above embodiments of this application are only for description and do not represent the superiority or inferiority of the embodiments.

[0154] The terms "including" and "having" and any variations thereof in the specification, claims and above-mentioned drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. The descriptions of "first", "second", "third", etc. are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are different types.

[0155] In the description of the embodiments of this application, "exemplary", "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for instance" is intended to present relevant concepts in a specific manner.

[0156] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B can represent A or B; "and / or" in the text is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0157] In some processes described in the embodiments of this application, a plurality of operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0159] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A water quality prediction method, characterized in that, The water quality prediction method includes: Preprocessing a time series composed of multiple water quality data to obtain a first input sequence; Using the variational mode decomposition algorithm to decompose the first input sequence to obtain multiple modal components, where the optimal parameters of the variational mode decomposition algorithm are determined by the NGO algorithm during training; Inputting the multiple modal components into a bidirectional long short-term memory network to obtain first water quality prediction data.

2. The water quality prediction method according to claim 1, characterized in that, After the step of inputting the multiple modal components into the bidirectional long short-term memory network to obtain the first water quality prediction data, it includes: Inputting the first input sequence into an autoregressive integrated moving average model to obtain second water quality prediction data; Multiplying the first water quality prediction data and the second water quality prediction data by a first weight coefficient and a second weight coefficient respectively and summing them to obtain final water quality prediction data.

3. The water quality prediction method according to claim 2, characterized in that Before the step of multiplying the first water quality prediction data and the second water quality prediction data by a first weight coefficient and a second weight coefficient respectively and summing them to obtain the final water quality prediction data, it includes: Preprocessing a time series composed of multiple historical water quality data and constructing a water quality data training set, where the water quality data training set includes multiple groups of second input sequences, and each group of second input sequences includes multiple preprocessed historical water quality data and corresponding true output water quality data; Using each group of second input sequences to train the bidirectional long short-term memory network to obtain a first training prediction value for each group of second input sequences; Using each group of second input sequences to train the autoregressive integrated moving average model to obtain a second training prediction value for each group of second input sequences; Multiplying the first training prediction value and the second training prediction value of each group of second input sequences by different third weight coefficients and fourth weight coefficients respectively and summing them to obtain a final training prediction value for each group of second input sequences; Calculating the mean square error between the final training prediction values of all groups of second input sequences and the corresponding true output water quality data; Determining the third weight coefficient and the fourth weight coefficient corresponding to the minimum mean square error as the first weight coefficient and the second weight coefficient.

4. The water quality prediction method according to claim 3, characterized in that, The determination of the optimal parameters of the variational mode decomposition algorithm by the NGO algorithm during training includes: Generating an initial population, where the initial population includes multiple individuals, and each individual includes the number of modal components and a penalty parameter, where the number of modal components and the penalty parameter are random values within a preset range; Taking each individual as the parameter of the variational mode decomposition algorithm, using the variational mode decomposition algorithm to decompose the second input sequence to obtain multiple modal components, inputting the multiple modal components into the bidirectional long short-term memory network for training, and obtaining a second training prediction value for each individual; Calculating the fitness of the population based on the second training prediction value of each individual and the corresponding true output water quality data through Formula 1; If the number of adjustment times is less than or equal to the preset number of times, adjusting the individual through Formula 2 and returning to execute the step of taking each individual as the parameter of the variational mode decomposition algorithm, using the variational mode decomposition algorithm to decompose the second input sequence to obtain multiple modal components, inputting the multiple modal components into the bidirectional long short-term memory network for training, and obtaining a second training prediction value for each individual; If the number of adjustment times is greater than the preset number of times, the number of modal components and the numerical value of the penalty parameter of the individual with the smallest fitness loss function in the population are used as the optimal parameters of the variational mode decomposition algorithm; The first formula is as follows: The formula two is as follows: where MSE is the fitness loss function of the population, N is the number of individuals, y i is the true output water quality data corresponding to the i-th individual, is the second training prediction value of the i-th individual, is the position of the i-th individual in the (t + 1)-th adjustment, is the position of the i-th individual in the t-th adjustment, ω is the inertia coefficient used to control the influence of the previous adjustment speed, is the historical best position of the i-th individual, c1 and c2 are acceleration constants used to control the speed of the individual moving towards the individual optimum and the global optimum, r1 and r2 are random numbers in the range [0, 1], is the global best position among all individuals.

5. The water quality pre-treatment method according to claim 3, characterized in that, Each piece of water quality data includes multiple features and corresponding feature data. Before training the bidirectional long short-term memory network using each group of second input sequences to obtain the first training prediction value of each group of second input sequences, it includes: If the number of features is greater than the preset number, the importance score of each feature is calculated using the water quality data training set through a linear regression model; The feature with the lowest importance score and the corresponding feature data are removed; If the number of features after removal is greater than the preset number, the water quality data training set after removal is used as the new water quality data training set, and the step of calculating the importance score of each feature using the water quality data training set through the linear regression model is returned; 6. The water quality prediction method according to claim 1, wherein The preprocessing of the time series composed of multiple pieces of water quality data to obtain the first input sequence includes: Detecting outliers and missing values in the time series composed of multiple pieces of water quality data; Replacing the outliers and missing values with the mean or median of multiple pieces of water quality data; Normalizing each piece of water quality data after replacement to obtain the first input sequence.

7. The water quality prediction method according to claim 6, characterized in that The detection of outliers in the time series composed of multiple pieces of water quality data includes: For each piece of water quality data, the standard score of the water quality data is calculated using the water quality data, the mean and standard deviation of multiple pieces of water quality data through formula three. If the absolute value of the standard score of the water quality data is greater than the preset score, the water quality data is determined as an outlier. The formula three is: Where Z is the standard score of the water quality data, x is the water quality data, μ is the mean of multiple pieces of water quality data, and σ is the standard deviation of multiple pieces of water quality data.

8. A water quality prediction device, characterized in that, The water quality prediction device includes: A preprocessing module for preprocessing the time series composed of multiple pieces of water quality data to obtain the first input sequence; A decomposition module for decomposing the first input sequence using the variational mode decomposition algorithm to obtain multiple modal components, where the optimal parameters of the variational mode decomposition algorithm are determined through the NGO algorithm during training; A prediction module for inputting the multiple modal components into the bidirectional long short-term memory network to obtain the first water quality prediction data.

9. A water quality prediction device, characterized in that, The water quality prediction device includes a processor, a memory, and a water quality prediction program stored on the memory and executable by the processor. When the water quality prediction program is executed by the processor, the steps of the water quality prediction method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that, A water quality prediction program is stored on the readable storage medium. When the water quality prediction program is executed by the processor, the steps of the water quality prediction method according to any one of claims 1 to 7 are implemented.