Sewage treatment multi-water-quality parallel intelligent prediction method based on combined gradient descent

By applying a multi-task parallel intelligent prediction method based on joint gradient descent in sewage treatment, the problem that the existing technology is difficult to predict multiple water quality indicators at the same time is solved, and high accuracy and high efficiency water quality prediction is achieved, reducing operating costs.

CN120218306APending Publication Date: 2025-06-27BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510226806.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

It is difficult for existing sewage treatment technologies to accurately predict multiple water quality indicators at the same time, resulting in low water quality in the sewage treatment process meeting the standards.

Method used

Using a multi-task parallel intelligent prediction method based on joint gradient descent, by building a multi-task parallel model, sharing structures and parameters, using the joint gradient descent algorithm to update shared parameters, dynamically adjust the learning rate, and achieve accurate prediction of multiple water quality parameters at the same time.

Benefits of technology

It improves the accuracy and operating efficiency of sewage treatment water quality prediction, reduces operating costs, and achieves accurate prediction of multiple water quality indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218306A_ABST
    Figure CN120218306A_ABST
Patent Text Reader

Abstract

The invention provides a sewage treatment multi-water-quality parallel intelligent prediction method based on combined gradient descent, and aims to solve the problem that water quality in a sewage treatment process is difficult to accurately predict due to the fact that water quality indexes are more, pollutants influence each other and a single-task prediction model cannot fully utilize related information among the indexes in the sewage treatment process. Firstly, a multi-water-quality parallel prediction model based on a fuzzy neural network is constructed; secondly, a joint gradient descent algorithm is designed to update model sharing parameters online, and sharing utilization of related information between tasks is achieved; finally, a self-adaptive learning rate strategy is provided, the learning rate of a single model is adjusted by training error change, the global learning rate is adjusted by training gradient change and correlation, and sewage treatment multi-water-quality accurate prediction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention designs a multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent, realizing the simultaneous and accurate prediction of multiple water-quality indicators during the sewage treatment process. The result of the effluent water quality of sewage treatment is an important basis for evaluating the effectiveness of the sewage treatment process. Predicting the effluent water quality of sewage treatment can improve the accuracy of achieving the standard of sewage treatment water quality. Applying the multi-task parallel model based on joint gradient descent to the parallel prediction of multiple water qualities in sewage treatment not only improves the prediction accuracy and the operation efficiency of the sewage treatment process, but also can effectively reduce the operation cost. Background Art

[0002] In order to ensure the stable operation of the sewage treatment system, predicting water quality changes has become an urgent need. By accurately predicting water-quality indicators, the trend of water quality changes can be discovered in advance and the treatment process can be adjusted in time to optimize the sewage treatment process and improve the water-quality compliance rate. The water-quality prediction method based on machine learning can process a large amount of non-linear complex data to obtain prediction results. However, the water-quality indicators in sewage treatment affect each other. These single-task water-quality prediction models often regard each water-quality indicator as a separate task and cannot make full use of all the information of related water qualities. How to consider multiple water-quality indicators simultaneously and realize the simultaneous prediction of the water quality in the sewage treatment process has become a key topic in water resource management. In the multi-task learning model, knowledge can be shared from multiple related tasks, which is crucial for improving the prediction accuracy of the multi-water-quality prediction model in sewage treatment. Therefore, it is of great significance to study a prediction method that can predict multiple water qualities and realize the simultaneous and accurate prediction of multiple water qualities in sewage treatment.

[0003] The present invention designs a multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent. This method analyzes multiple water-quality indicators of sewage treatment, selects corresponding characteristic variables, establishes a multi-water-quality parallel intelligent prediction model based on joint gradient descent, utilizes task-related information through model structure and parameter sharing, updates the shared parameters using the joint gradient descent algorithm, and dynamically adjusts the learning rate of the shared parameters using the task gradient similarity to achieve the simultaneous and accurate prediction of multiple water-quality parameters in sewage treatment. Summary of the Invention

[0004] The present invention obtains a multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent. This method selects characteristic variables related to the effluent water quality in the field of sewage treatment to determine the input of the model and predicts multiple effluent water qualities. A multi-task parallel model based on a fuzzy neural network (FNN) is established, and the shared structure utilizes relevant information. A joint gradient descent algorithm is designed to modify the gradient direction in the multi-task training process and jointly update the shared parameters. The task-specific learning rate and the global learning rate are dynamically adjusted respectively using the error change and the gradient similarity, enabling multiple water quality prediction tasks in the sewage treatment process to be carried out simultaneously, and improving the accuracy of sewage treatment water quality prediction.

[0005] The present invention adopts the following technical solutions and implementation steps:

[0006] A multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent, characterized by: determining the input and output variables of the model, constructing a multi-task parallel model, designing a model parameter training method, adjusting the training speed of each task of the model, and predicting multiple water quality indicators in the sewage treatment process, including the following steps:

[0007] (1) Determine the input and output variables of the multi-task parallel model: Taking the sewage treatment process as the research object, predict the total nitrogen in the effluent and the total phosphorus in the effluent. The first task is to predict the total nitrogen in the effluent, and the corresponding input variables are six variables: ammonia nitrogen in the influent, oxidation-reduction potential in the anaerobic tank, oxidation-reduction potential in the anoxic tank, nitrate nitrogen in the anoxic tank, mixed liquor suspended solid concentration in the anoxic tank, and sludge volume. The second task is to predict the total phosphorus in the effluent, and the corresponding input variables are six variables: hydrogen ion concentration in the influent, chemical oxygen demand in the influent, dissolved oxygen in the aerobic tank, orthophosphate in the effluent of the secondary sedimentation tank, external reflux flow rate, and temperature. Obtain the ammonia nitrogen data in the influent at time t Oxidation-reduction potential data in the anaerobic tank Oxidation-reduction potential data in the anoxic tank Nitrate nitrogen data in the anoxic tank Mixed liquor suspended solid concentration data in the anoxic tank Sludge volume data Construct the first task sample matrix Obtain the hydrogen ion concentration data in the influent at time t Chemical oxygen demand data in the influent Dissolved oxygen data in the aerobic tank Orthophosphate data in the effluent of the secondary sedimentation tank External reflux flow rate data Temperature data Construct the second task sample matrix U is the total number of samples, and T represents the transpose of the matrix;

[0008] (2) Construct a multi-task parallel intelligent water quality prediction model

[0009] The multi-task parallel intelligent water quality prediction model consists of two fuzzy neural networks, including an input layer, an RBF layer, a normalization layer, and an output layer. There are 12 neurons in the input layer, N1 + N2 + M neurons in the RBF layer, N1 + N2 + M neurons in the normalization layer, and 2 neurons in the output layer. M neurons are shared in the hidden layer between tasks. N1 is the number of neurons specific to the first task, and N2 is the number of neurons specific to the second task. Among them, the shared neurons are the hidden layer neurons to which the inputs of all tasks are connected, and the task-specific neurons are the hidden layer neurons to which the input of the q-th specific task is connected, which are neurons only for that specific task. When q = 1, the specific task represents predicting the total nitrogen in the effluent; when q = 2, the specific task represents predicting the total phosphorus in the effluent. The outputs of each layer of the q-th task are as follows:

[0010] Input layer: This layer is the input variable for water quality prediction. The input of the q-th task is

[0011]

[0012] RBF layer: This layer includes task-specific neurons and shared neurons, and the output is

[0013]

[0014] where e is the natural logarithm, is the output of the n-th specific neuron in the RBF layer of the q-th task at time t, q = 1, 2, n = 1, 2,..., N, is the center of the n-th specific neuron corresponding to the i-th input of the q-th task at time t, i = 1, 2,..., 6. The specific center parameter matrix of the neurons in this layer is is the width of the n-th specific neuron corresponding to the i-th input of the q-th task at time t. The specific width parameter matrix of the neurons in this layer is

[0015]

[0016] where is the output of the m-th shared neuron in the RBF layer of the q-th task at time t, m = 1, 2,..., M, is the center of the m-th shared neuron corresponding to the i-th input of each task at time t. The shared center parameter matrix of the neurons in this layer is is the width of the m-th shared neuron corresponding to the i-th input of each task at time t. The shared width parameter matrix of the neurons in this layer is

[0017] Normalization layer: The number of neurons of each type in this layer is the same as that in the RBF layer, and the output is

[0018]

[0019] where is the output of the nth specific neuron in the normalization layer of the qth task at time t, is the output of the mth shared neuron in the normalization layer of the qth task at time t;

[0020] Output layer: The number of neurons in this layer is the same as the number of tasks

[0021]

[0022] where is the output of the qth task at time t, and w q (t) is the specific weight parameter matrix of the qth task at time t, is the shared weight parameter matrix at time t,

[0023] (3) Updating of multi-task parallel model parameters

[0024] ① Initialize the multi-task parallel water quality intelligent prediction model. Let the current time t = 1, the current iteration number k = 1. The initial central values of the neurons in the RBF layer of the multi-task parallel model are random values between 0 and 1, that is and All internal elements are 1, and the initial width value is 1, that is and All internal elements of are 1. The initial weights of the normalization layer are random values between 0 and 1, that is w q (1) and w s (1) All internal elements are 1, and the maximum number of iterations is K, where K is an integer greater than or equal to 100;

[0025] ② Use formulas (2)-(6) to obtain the output of the qth task sample at the kth iteration and time t. The loss function of the qth task is Define the joint loss function of the multi-task parallel model as The update formula for the specific parameters of the multi-task parallel model at the kth iteration using the gradient descent method is:

[0026]

[0027] where is the specific learning rate of the qth task at the kth iteration, and the initial setting is

[0028] ③ Update the shared parameters using the joint gradient descent algorithm, and merge each shared parameter gradient into a gradient matrix

[0029] (t)], calculate the cosine value of the included angle between the two gradient vectors at the t-th moment of the k-th iteration, and determine whether gradient modification is required

[0030]

[0031] Among them, the symbol · represents the dot product of vectors, and ||·||2 represents the second norm, that is, the square root of the sum of the squares of all internal elements;

[0032] When cosφ k (t)>0, perform gradient modification, and calculate the gradient interference direction vector Calculate the gradient cosine value of the two gradients in the interference direction as

[0033]

[0034] Then calculate the gradient difference between the two gradients in the gradient interference direction Calculate the modified gradient as:

[0035]

[0036] Denote Represents the difference between the i-th input and output of the q-th task at the t-th moment, with the input data selected as the reference sequence The output data is selected as the comparison sequence y q (t), calculate the grey correlation coefficient as:

[0037]

[0038] Among them, min t min i Is the minimum value of the difference between all input sequences And the corresponding output sequences and at all moments t, max t max i Is the maximum value of the difference between all input sequences And the corresponding output sequences and at all moments t;

[0039] Then calculate the average value of the correlation coefficients of each reference sequence as:

[0040]

[0041] Calculate the task gradient weight α using the result of formula (13) q As:

[0042]

[0043] Then, the update formula for the network shared parameters in the k-th iteration of the multi-task parallel water quality intelligent prediction model is as follows:

[0044]

[0045] where η k is the global learning rate for updating the shared parameters, and the initial setting is

[0046] ④ If the time t < U, then t is incremented by 1, and it turns to step ②; if t = U, it turns to the steps in section (4);

[0047] (4) Adjust the training speed of the multi-task parallel model

[0048] Automatically adjust the task-specific learning rate according to the error change in each iteration as follows:

[0049]

[0050] where sign(.) is the sign function. When When When

[0051] Adjust the global learning rate according to the gradient similarity and the cumulative gradient in each iteration as follows:

[0052]

[0053] where S q is the similarity after modifying the task gradient, which is calculated using formula (8) and The cosine value is used to obtain S q , is the cumulative gradient of all tasks with respect to the shared parameters. The specific calculation method is as follows:

[0054]

[0055] where is the cumulative gradient of all shared parameters in the shared structure l with respect to the task q in the k-th iteration. The ratio of C qlk in equation (20) is the contribution degree of task q with respect to the shared parameters among all tasks in the k-th iteration;

[0056] Determine whether the maximum number of training rounds has been reached. If the iteration count K is reached, the training ends and proceed to the steps in Section (5); if the iteration count K has not been reached, then increment k by 1 and return to step ② in Section (3) to continue training;

[0057] Determine whether the network structure has reached the optimal state, adjust the number of shared neurons, and use the RMSE k value for judgment. The calculation method is as described in the explanation of formula (16). If the training results of both tasks are worse than the test results, increase the number of shared neurons; otherwise, decrease it. The number of shared neurons is greater than or equal to 1. In each specific task, adjust the number of specific neurons according to the training result of that task. If the training value decreases, increase the number of shared neurons; otherwise, decrease it. The number of specific neurons is greater than or equal to 1, and then return to Section (2) to retrain;

[0058] (5) Implement parallel intelligent prediction of multiple water qualities for sewage treatment based on joint gradient descent

[0059] Implement parallel intelligent prediction of multiple water qualities in the sewage treatment process, collect the input and output corresponding to the two task samples of the total nitrogen and total phosphorus in the effluent, and obtain the influent ammonia nitrogen data at time t Redox potential data in the anaerobic tank Redox potential data in the anoxic tank Nitrate nitrogen data in the anoxic tank Mixed liquor suspended solid concentration data in the anoxic tank Sludge volume data Construct the first task sample matrix Obtain the influent hydrogen ion concentration data at time t Influent chemical oxygen demand data Dissolved oxygen data in the aerobic tank Orthophosphate data at the effluent end of the secondary sedimentation tank External reflux flow data Temperature data Construct the second task sample matrix For t = 1, 2, …, U, where U is the total number of samples; perform normalization on the data, input the input vector into the input layer of the multi-task parallel water quality intelligent prediction model, and obtain the output value of the task parallel water quality intelligent prediction model after passing through the RBF layer, normalization layer, and output layer of the multi-task parallel water quality intelligent prediction model These are the normalized predicted values of the total nitrogen and total phosphorus in the effluent; perform denormalization on the normalized predicted values of the total nitrogen and total phosphorus in the effluent to obtain the predicted values of the total nitrogen and total phosphorus in the effluent.

[0060] The creativity of the present invention is mainly reflected in:

[0061] In view of the fact that there are many indicators required in the sewage treatment process, and there are often complex mutual influences among pollutants, a single water quality prediction model often has difficulty in processing these multi-dimensional data simultaneously. The single-task prediction model ignores the relevant information between various water quality prediction tasks, resulting in relatively low accuracy of the prediction results. A multi-task parallel intelligent prediction model based on joint gradient descent is established. This model obtains relevant information through a shared structure, proposes a joint gradient descent algorithm to solve the gradient conflict between tasks, updates the shared parameters, and dynamically adjusts the global learning rate by using the gradient similarity between tasks and the contribution of the cumulative gradient to balance the task learning speed and improve the accuracy of water quality prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 FIG. is a diagram of the multi-task parallel water quality prediction model of the present invention. The red part is the neuron part shared by all tasks, and the inputs of all tasks need to be connected. Each specific task has specific neurons of the same color, which are only connected to the inputs and outputs of the corresponding specific tasks. Taking the present invention as an example, the first specific task is to predict the total nitrogen in the effluent. The yellow part is the specific neuron part of this task, which is initially set to half of the number of input neurons, i.e., 3. When the training error RMSE1 of the total nitrogen in the effluent is less than 1, the specific number is adjusted to N1. The second specific task is to predict the total phosphorus in the effluent. The blue part is the specific neuron part of this task, which is initially set to half of the number of input neurons, i.e., 3. When the training error RMSE2 of the total phosphorus in the effluent is less than 0.003, the specific number is adjusted to N2. The shared neurons are initially set to one, and when the training errors of both tasks are standard, the specific number is adjusted to M;

[0063] Figure 2 FIG. is a diagram of the prediction results of the model of the present invention for the training samples of the total nitrogen in the sewage treatment effluent;

[0064] Figure 3 FIG. is a diagram of the prediction results of the model of the present invention for the training samples of the total phosphorus in the sewage treatment effluent;

[0065] Figure 4 FIG. is a diagram of the prediction results of the model of the present invention for the test samples of the total nitrogen in the sewage treatment effluent;

[0066] Figure 5 FIG. is a diagram of the prediction results of the model of the present invention for the test samples of the total phosphorus in the sewage treatment effluent. DETAILED DESCRIPTION OF THE INVENTION

[0067] The present invention designs a multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent. Taking two water-quality prediction tasks as examples, the selected input variables are six variables: influent ammonia nitrogen, oxidation-reduction potential in the anaerobic tank, oxidation-reduction potential in the anoxic tank, nitrate nitrogen in the anoxic tank, mixed liquor suspended solid concentration in the anoxic tank, and sludge volume, and six variables: influent hydrogen ion concentration index, influent chemical oxygen demand, dissolved oxygen in the aerobic tank, orthophosphate in the effluent of the secondary sedimentation tank, and temperature. The total nitrogen and total phosphorus in the effluent are used as the output variables of the multi-task parallel water-quality prediction model.

[0068] The experimental data comes from the actual sampling data of the operating variables of a sewage treatment plant in 2021, with a sampling interval of one minute. After screening and processing, 300 data samples are selected as training samples, and the latter 100 samples are used as test samples.

[0069] The present invention adopts the following technical solutions and implementation steps:

[0070] A multi-water-quality parallel intelligent prediction method for sewage treatment based on joint gradient descent, characterized in that: determining the input and output variables of the model, constructing a multi-task parallel model, designing a model parameter training method, adjusting the training speed of each task of the model, and predicting multiple water-quality indicators in the sewage treatment process, including the following steps:

[0071] (1) Determining the input and output variables of the multi-task parallel model: Taking the sewage treatment process as the research object, predicting the total nitrogen and total phosphorus in the effluent. The first task is to predict the total nitrogen in the effluent, and the corresponding input variables are six variables: influent ammonia nitrogen, oxidation-reduction potential in the anaerobic tank, oxidation-reduction potential in the anoxic tank, nitrate nitrogen in the anoxic tank, mixed liquor suspended solid concentration in the anoxic tank, and sludge volume; the second task is to predict the total phosphorus in the effluent, and the corresponding input variables are six variables: influent hydrogen ion concentration, influent chemical oxygen demand, dissolved oxygen in the aerobic tank, orthophosphate in the effluent of the secondary sedimentation tank, external reflux flow rate, and temperature; obtaining the influent ammonia nitrogen data at time t Oxidation-reduction potential data in the anaerobic tank Oxidation-reduction potential data in the anoxic tank Nitrate nitrogen data in the anoxic tank Mixed liquor suspended solid concentration data in the anoxic tank Sludge volume data Constructing the first task sample matrix Obtaining the influent hydrogen ion concentration data at time t Influent chemical oxygen demand data Dissolved oxygen data in the aerobic tank Orthophosphate data in the effluent of the secondary sedimentation tank External reflux flow rate data Temperature data Constructing the second task sample matrix U is the total number of samples, and T represents the transpose of the matrix;

[0072] (2) Construct a multi-task parallel water quality intelligent prediction model

[0073] The multi-task parallel water quality intelligent prediction model consists of two fuzzy neural networks, including an input layer, an RBF layer, a normalization layer, and an output layer. There are 12 neurons in the input layer, N1 + N2 + M neurons in the RBF layer, N1 + N2 + M neurons in the normalization layer, and 2 neurons in the output layer; M neurons in the hidden layer are shared between tasks. N1 is the number of neurons specific to the first task, and N2 is the number of neurons specific to the second task. Among them, the shared neurons are the hidden layer neurons to which all task inputs are connected, and the task-specific neurons refer to the hidden layer neurons to which the inputs of the q-th specific task are connected and are neurons specific to that particular task. When q = 1, the specific task represents predicting the total nitrogen in the effluent; when q = 2, the specific task represents predicting the total phosphorus in the effluent. The outputs of each layer of the q-th task are as follows:

[0074] Input layer: This layer is the input variable for water quality prediction. The input of the q-th task is

[0075]

[0076] RBF layer: This layer includes task-specific neurons and shared neurons, and the output is

[0077]

[0078] where e is the natural logarithm, is the output of the n-th specific neuron in the RBF layer of the q-th task at time t, q = 1, 2, n = 1, 2, …, N, is the center of the n-th specific neuron corresponding to the i-th input of the q-th task at time t, i = 1, 2, …, 6. The specific center parameter matrix of the neurons in this layer is is the width of the n-th specific neuron corresponding to the i-th input of the q-th task at time t. The specific width parameter matrix of the neurons in this layer is

[0079]

[0080] where is the output of the m-th shared neuron in the RBF layer of the q-th task at time t, m = 1, 2, …, M, is the center of the m-th shared neuron corresponding to the i-th input of each task at time t. The shared center parameter matrix of the neurons in this layer is is the width of the m-th shared neuron of the i-th input of each task at time t. The width parameter matrix of the neurons in this layer is

[0081] Normalization layer: The number of neurons of each type in this layer is the same as that in the RBF layer, and the output is

[0082]

[0083] where is the output of the n-th specific neuron in the normalization layer of the q-th task at time t, is the output of the m-th shared neuron in the normalization layer of the q-th task at time t;

[0084] Output layer: The number of neurons in this layer is the same as the number of tasks

[0085]

[0086] where is the output of the q-th task at time t, and w q (t) is the specific weight parameter matrix of the q-th task at time t, w s (t) is the shared weight parameter matrix at time t,

[0087] (3) Multi-task parallel model parameter update

[0088] ① Initialize the multi-task parallel water quality intelligent prediction model. Let the current time t = 1, the current iteration number k = 1. The initial central value of the neurons in the RBF layer of the multi-task parallel model is a random value between 0 and 1, that is and All internal elements are 1, and the initial width value is 1, that is and All internal elements are 1. The initial weight of the normalization layer is a random value between 0 and 1, that is w q (1) and w s (1) All internal elements are 1. The maximum number of iterations is K, and K is an integer greater than or equal to 100;

[0089] ② Use formulas (2)-(6) to obtain the output of the q-th task sample at time t of the k-th iteration. The loss function of the q-th task is Define the joint loss function of the multi-task parallel model as The update formula for the specific parameters of the multi-task parallel model at the k-th iteration using the gradient descent method is:

[0090]

[0091] where is the q-th task-specific learning rate at the k-th iteration, initialized to

[0092] ③ Update the shared parameters using the joint gradient descent algorithm, and merge the gradients of each shared parameter into a gradient matrix

[0093] (t)], calculate the cosine value of the angle between the two gradient vectors at time t of the k-th iteration, and determine whether gradient modification is required

[0094]

[0095] where the symbol · represents the dot product of vectors, and ||·||2 represents the two-norm, that is, the square root of the sum of the squares of all internal elements;

[0096] When cosφ k (t)>0, perform gradient modification and calculate the gradient interference direction vector Calculate the cosine value of the gradients of the two gradients in the interference direction as

[0097]

[0098] Then calculate the gradient difference between the two gradients in the gradient interference direction Calculate the modified gradient as:

[0099]

[0100] Denote as the difference between the i-th input and output of the q-th task at time t, with the input data selected as the reference sequence The output data is selected as the comparison sequence y q (t), and calculate the grey correlation coefficient as:

[0101]

[0102] where min t min i is the minimum value of the differences between all input sequences and the corresponding output sequences at all times t, and max t max i is the maximum value of the differences between all input sequences and the corresponding output sequences at all times t;

[0103] Then calculate the average value of the correlation coefficients of each reference sequence as:

[0104]

[0105] Calculate the task gradient weight α using the result of formula (13). q It is:

[0106]

[0107] Then, the update formula for the network shared parameters in the k-th iteration of the multi-task parallel water quality intelligent prediction model is:

[0108]

[0109] where η k is the global learning rate for updating the shared parameters, and the initial setting is

[0110] ④ If the time t < U, then t is incremented by 1, and go to step ②; if t = U, go to the steps in section (4).

[0111] (4) Adjust the training speed of the multi-task parallel model

[0112] Automatically adjust the task-specific learning rate according to the error change in each iteration as:

[0113]

[0114] where sign(.) is the sign function. When When When

[0115] Automatically adjust the global learning rate according to the gradient similarity and the cumulative gradient in each iteration as:

[0116]

[0117] where S q is the similarity after modifying the task gradient, calculated using formula (8) and The cosine value is used to obtain S q , is the cumulative gradient of all tasks with respect to the shared parameters. The specific calculation method is:

[0118]

[0119]

[0120] where is the cumulative gradient of all shared parameters in the shared structure l with respect to the task q in the k-th iteration. In formula (20), C qlkThe ratio is the contribution degree of the k-th iteration task q to the shared parameters among all tasks;

[0121] Judge whether the maximum number of training epochs is reached. If the iteration number K is reached, the training ends and go to the steps in Section (5); if the iteration number K is not reached, then k is incremented by 1, and return to the second step in Section (3) to continue training;

[0122] Judge whether the network structure reaches the optimum, adjust the number of shared neurons, and use the RMSE k value for judgment. The calculation method is as shown in the description of formula (36). If the training results of both tasks are worse than the test results, increase the number of shared neurons; otherwise, decrease it. The number of shared neurons is greater than or equal to 1. In each specific task, adjust the number of specific neurons according to the training result of this task. If the training RMSE q k value decreases, increase the number of shared neurons; otherwise, decrease it. The number of specific neurons is greater than or equal to 1, and return to Section (2) to retrain;

[0123] (5) Implement multi-water-quality parallel intelligent prediction for sewage treatment based on joint gradient descent

[0124] Implement multi-water-quality parallel intelligent prediction for the sewage treatment process, collect the input and output corresponding to the two task samples of the total nitrogen and total phosphorus in the effluent, and obtain the influent ammonia nitrogen data at time t Redox potential data in the anaerobic tank Redox potential data in the anoxic tank Nitrate nitrogen data in the anoxic tank Mixed liquor suspended solid concentration data in the anoxic tank Sludge volume data Construct the first task sample matrix Obtain the influent hydrogen ion concentration data at time t Influent chemical oxygen demand data Dissolved oxygen data in the aerobic tank Orthophosphate data at the effluent end of the secondary sedimentation tank External reflux flow data Temperature data Construct the second task sample matrix t = 1, 2, …, U, where U is the total number of samples; perform normalization on the data, input the input vector into the input layer of the multi-task parallel water-quality intelligent prediction model, and pass it through the RBF layer, normalization layer, and output layer of the multi-task parallel water-quality intelligent prediction model to obtain the output value of the task parallel water-quality intelligent prediction model It is the predicted value of the total nitrogen and total phosphorus in the effluent after normalization; the predicted value of the total nitrogen and total phosphorus in the effluent after normalization Perform inverse normalization to obtain the predicted values of the total nitrogen and total phosphorus in the effluent.

Claims

1. A parallel intelligent prediction method for multiple water qualities in sewage treatment based on joint gradient descent, characterized in that: It includes the following steps: (1) Determine the input and output variables of the multi-task parallel model: Take the sewage treatment process as the research object and predict the total nitrogen and total phosphorus in the effluent; the first task is to predict the total nitrogen in the effluent, and the corresponding input variables are six variables: influent ammonia nitrogen, redox potential in the anaerobic tank, redox potential in the anoxic tank, nitrate nitrogen in the anoxic tank, suspended solids concentration in the mixed liquor of the anoxic tank, and sludge volume; the second task is to predict the total phosphorus in the effluent, and the corresponding input variables are six variables: influent hydrogen ion concentration, influent chemical oxygen demand, dissolved oxygen in the aerobic tank, positive phosphorus at the effluent end of the secondary sedimentation tank, external reflow flow, and temperature; obtain the influent ammonia nitrogen data at time t Oxidation-reduction potential data in anaerobic tanks Anoxic pool redox potential data Anoxic pool nitrate data Suspended solids concentration data of mixed liquor in anoxic tank Sludge volume data Construct the first task sample matrix Get the inlet water hydrogen ion concentration data at time t Influent Chemical Oxygen Demand Data Aerobic pool dissolved oxygen data Phosphorus data of the secondary sedimentation tank effluent External return flow data Temperature data Construct the second task sample matrix U is the total number of samples, and T represents the transpose of the matrix; (2) Construct a multi-task parallel water quality intelligent prediction model The multi-task parallel water quality intelligent prediction model consists of two fuzzy neural networks, including an input layer, an RBF layer, a normalization layer, and an output layer. There are 12 neurons in the input layer, N1+N2+M neurons in the RBF layer, N1+N2+M neurons in the normalization layer, and 2 neurons in the output layer; M neurons are shared in the hidden layer between tasks. N1 is the number of neurons specific to the first task, and N2 is the number of neurons specific to the second task. Among them, the shared neurons are the hidden layer neurons to which the inputs of all tasks are connected, and the task-specific neurons refer to the hidden layer neurons to which the input of the q-th specific task is connected, which are neurons only for that specific task. When q = 1, the specific task represents predicting the total nitrogen in the effluent. When q = 2, the specific task represents predicting the total phosphorus in the effluent. The outputs of each layer of the q-th task are as follows: Input layer: This layer is the input variable for water quality prediction. The input of the q-th task is RBF layer: This layer includes task-specific neurons and shared neurons, and the output is where e is the natural logarithm, is the output of the nth specific neuron in the qth task RBF layer at time t, q = 1, 2, n = 1, 2, ..., N q , is the center of the nth specific neuron of the ith input of the qth task at time t, i = 1, 2, ..., 6, and the specific center parameter matrix of the neurons in this layer is is the width of the nth specific neuron of the ith input of the qth task at time t. The specific width parameter matrix of the neurons in this layer is in is the output of the mth shared neuron in the RBF layer of the qth task at time t, m = 1, 2, ..., M, is the center of the mth shared neuron of the ith input of each task at time t. The parameter matrix of the shared center of neurons in this layer is is the width of the mth shared neuron of the ith input of each task at time t. The parameter matrix of the shared width of neurons in this layer is Normalization layer: The number of neurons of each type in this layer is the same as that in the RBF layer, and the output is in is the output of the nth specific neuron in the qth task normalization layer at time t, is the output of the mth shared neuron in the qth task normalization layer at time t; Output layer: The number of neurons in this layer is the same as the number of tasks in is the output of the qth task at time t, w q (t) is the qth task-specific weight parameter matrix at time t, w s (t) is the shared weight parameter matrix at time t, (3) Update the parameters of the multi-task parallel model ① Initialize the multi-task parallel water quality intelligent prediction model, set the current time t = 1, the current iteration number k = 1, and the initial center value of the RBF layer neuron of the multi-task parallel model is a random value between 0 and 1, that is and The internal elements are all 1, and the initial width value is 1, that is, and The internal elements of are all 1, and the initial weight of the normalized layer is a random value between 0 and 1, that is, w q (1) and w s The internal elements of (1) are all 1, the maximum number of iterations is K, and K is an integer greater than or equal to 100; ②Use formulas (2)-(6) to obtain the sample output of the qth task at the kth iteration t. The loss function of the qth task is: Define the joint loss function of the multi-task parallel model as The specific parameters of the kth iteration of the multi-task parallel model are updated using the gradient descent method as follows: in is the qth task-specific learning rate at the kth iteration, initialized as ③ Use the joint gradient descent algorithm to update the shared parameters and merge the gradients of each shared parameter into a gradient matrix (t)], calculate the cosine value of the included angle between the two gradient vectors at the t-th moment of the k-th iteration, and judge whether gradient modification is needed Among them, the symbol · represents the dot product of vectors, and ||·||2 represents the two-norm, that is, the square root of the sum of the squares of all internal elements; When cosφ k When (t)>0, the gradient is modified and the gradient interference direction vector is calculated Calculate the gradient cosine of the two gradients in the interference direction as Then calculate the gradient difference between the two gradients in the gradient interference direction The modified gradient is calculated as: remember Represents the difference between the i-th input and output of the q-th task at time t, with the input data selected as the reference sequence Output data selected as comparison sequence y q (t), calculate the grey correlation coefficient as: Where min t min i is the sequence of all inputs at all times t The minimum value of the difference between the corresponding output sequence and, max t max i is the sequence of all inputs at all times t The maximum value of the difference between the corresponding output sequences and ; Then calculate the average value of the correlation coefficients of each reference sequence as: The task gradient weight α is calculated using the result of formula (13): q for: Then the update formula for the network shared parameters of the multi-task parallel water quality intelligent prediction model at the k-th iteration is: where η k is the global learning rate used to update the shared parameters, which is initialized to ④ If the moment t < U, then t is incremented by 1, and turn to step ②; if t = U, turn to the steps in section (4); (4) Adjust the training speed of the multi-task parallel model Automatically adjust the task-specific learning rate according to the error change at each iteration as: in sign(.) is the sign function. when when Automatically adjust the global learning rate according to the gradient similarity and the cumulative gradient at each iteration as: Where S q is the similarity after task gradient modification, calculated using formula (8) and The cosine value is S q , It is the cumulative gradient of all tasks to the shared parameters. The specific calculation method is: in is the cumulative gradient of all shared parameters in the shared structure l for task q at the kth iteration, where The ratio of is the contribution of the k-th iteration task q to all tasks relative to the shared parameters; Judge whether the maximum number of training rounds is reached. If the iteration number K is reached, the training ends, and turn to the steps in section (5); if the iteration number K is not reached, then k is incremented by 1, and return to step ② in section (3) to continue training; Determine whether the network structure is optimal and adjust the number of shared neurons to obtain the optimal value based on RMSE. k The calculation method is as shown in the description of formula (16). If the training results of the two tasks are worse than the test results, increase the number of shared neurons, otherwise reduce them. The number of shared neurons is greater than or equal to 1. In each specific task, the number of specific neurons is adjusted according to the training results of the task. If the training result of the task is If the value decreases, the shared neurons are increased, otherwise it decreases. If the number of specific neurons is greater than or equal to 1, return to section (2) and retrain; (5) Implement multi-water-quality parallel intelligent prediction for sewage treatment based on joint gradient descent Realize parallel intelligent prediction of multiple water qualities in the sewage treatment process, collect the input and output corresponding to the two task samples of effluent total nitrogen and effluent total phosphorus, and obtain the influent ammonia nitrogen data at time t Oxidation-reduction potential data in anaerobic tanks Anoxic pool redox potential data Anoxic pool nitrate data Suspended solids concentration data of mixed liquor in anoxic tank Sludge volume data Construct the first task sample matrix Get the inlet water hydrogen ion concentration data at time t Influent Chemical Oxygen Demand Data Aerobic pool dissolved oxygen data Phosphorus data of the secondary sedimentation tank effluent External return flow data Temperature data Construct the second task sample matrix t=1,2,…,U, where U is the total number of samples; normalize the data, input the input vector into the input layer of the multi-task parallel water quality intelligent prediction model, and pass it through the RBF layer, normalization layer and output layer of the multi-task parallel water quality intelligent prediction model to obtain the output value of the task parallel water quality intelligent prediction model It is the normalized predicted value of effluent total nitrogen and effluent total phosphorus; the normalized predicted value of effluent total nitrogen and effluent total phosphorus After reverse normalization, the predicted values ​​of effluent total nitrogen and effluent total phosphorus were obtained.

Citation Information

Cited By

  • Method and system for dynamically predicting dehydration efficiency of river and lake bottom mud and automatically adjusting parameters

    CN121292782A