Soft measurement method, device and medium for effluent COD concentration based on double-error optimization neural network

CN117649891BActive Publication Date: 2026-09-04CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410015367.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2026-09-04
Estimated Expiration
2044-01-05

AI Technical Summary

Technical Problem

然而,由于污水处理工艺机理复杂、生化反应多、工况交变快,导致基于神经网络的软测量模型在使用过程中经常面临样本数据维度高、模型计算强度大、预测性能不稳定等问题

Benefits of technology

[0049]本发明的基于双误差优化神经网络的出水COD浓度软测量方法,首先计算过程变量与出水COD浓度之间的Pearson相关系数和核密度互信息值,依据帕累托原则将两者融合,选择最优的过程变量作为辅助变量;然后建立一个单隐藏层的前馈神经网络,同时以点预测误差最小化和分布误差最小化作为寻优目标,调节网络参数,提高模型的训练效率和检测结果的精度;最后将测试数据输入到最终的神经网络中,输出结果即为出水COD浓度的检测值。与现有技术相比,本发明具有以下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117649891B_ABST
    Figure CN117649891B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on double-error optimization neural network effluent COD concentration soft measurement method, equipment and medium, method: the sample data of target sewage treatment plant process variable is collected and preprocessed, including effluent COD concentration and other multiple process variables;The Pearson correlation coefficient and kernel density mutual information value between effluent COD concentration and each other process variable are calculated, and the weighted sum value is calculated and the value is preferably selected as auxiliary variable according to the process variable;Establish neural network, with auxiliary variable as input, effluent COD concentration as output, with point prediction error minimization and distribution error minimization as optimization target, adjust neural network parameter, obtain effluent COD concentration soft measurement model;Actual measurement, collect auxiliary variable value, input into effluent COD concentration soft measurement model, obtain effluent COD concentration.The application can quickly and accurately detect effluent COD concentration on-line, and ensure the safe and stable operation of sewage treatment plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent detection technology for wastewater treatment processes, specifically relating to a soft measurement method, equipment, and medium for effluent COD concentration based on a dual-error optimized neural network. Background Technology

[0002] To reduce water pollution caused by the large-scale discharge of domestic sewage and industrial wastewater and promote the recycling of water resources, numerous sewage treatment plants have been built and put into operation in recent years. Online monitoring of the key water quality parameter, effluent COD concentration, can provide data support for optimized control and operation management, thereby ensuring the stable operation of sewage treatment plants and improving effluent quality. Although effluent COD concentration can be measured using advanced instruments or obtained through offline sampling and analysis, the high cost of equipment maintenance and limited technical conditions often result in errors and delays in the detection results. This severely hinders the transmission of feedback signals in process control, optimization, and monitoring.

[0003] With the increasing storage and computing power of computers, neural network-based soft sensor models have been widely used to solve variable detection problems in wastewater treatment processes. However, due to the complexity of wastewater treatment processes, the numerous biochemical reactions, and the rapid changes in operating conditions, neural network-based soft sensor models often face problems such as high dimensionality of sample data, high computational intensity, and unstable predictive performance. Therefore, by selecting the optimal process variable as an auxiliary variable and establishing a neural network model with high computational efficiency and predictive performance, it is possible to achieve rapid and accurate online detection of effluent COD concentration, thereby ensuring the safe and stable operation of wastewater treatment plants. Summary of the Invention

[0004] This invention provides a soft measurement method, device, and medium for effluent COD concentration based on a dual-error optimized neural network, which can achieve online detection of effluent COD concentration more quickly and accurately, ensuring the safe and stable operation of wastewater treatment plants.

[0005] To achieve the above technical objectives, the present invention adopts the following technical solution:

[0006] A soft measurement method for effluent COD concentration based on a dual-error optimized neural network includes:

[0007] Step 1: Collect sample data of process variables from the target wastewater treatment plant; the process variables include effluent COD concentration and other process variables.

[0008] Step 2: Preprocess the collected sample data;

[0009] Step 3: Calculate the Pearson correlation coefficient and kernel density mutual information value between the water COD concentration and each of the other process variables;

[0010] Step 4: According to Pareto's law, the Pearson correlation coefficient and the kernel density mutual information value are weighted and summed, and some process variables are selected as auxiliary variables based on the weighted sum value;

[0011] Step 5: Establish a neural network with auxiliary variables from the sample data as input and effluent COD concentration as output. The optimization objectives are to minimize point prediction error and distribution error. Adjust the neural network parameters. The neural network with the parameters finally adjusted is recorded as the soft measurement model of effluent COD concentration.

[0012] Step 6: When it is necessary to measure the actual effluent COD concentration, the process variable in the auxiliary variables is collected and input into the soft measurement model of effluent COD concentration. The output result is the effluent COD concentration.

[0013] Furthermore, the other process variables include, but are not limited to: water flow rate, pH value, ammonia nitrogen concentration, dissolved oxygen concentration, total suspended solids, nitrate concentration, nitrite concentration, autotrophic bacteria bioactivity, and heterotrophic bacteria bioactivity.

[0014] Further, the sample data is preprocessed, including removing duplicate data, adding missing data, and normalizing the data.

[0015] Furthermore, the Pearson correlation coefficient between the effluent COD concentration and each of the other process variables is calculated as follows:

[0016]

[0017] Where ρ(·,·) represents the Pearson correlation coefficient, Cov(·,·) represents the covariance, Var(·) represents the standard deviation, X represents the sample data set of a certain process variable, and Y represents the sample data set of effluent COD concentration.

[0018] Furthermore, the kernel density mutual information value between the effluent COD concentration and each of the other process variables is calculated as follows:

[0019]

[0020] Where I(·,·) represents the kernel density mutual information value, p(·,·) represents the joint probability density function, and p(·) represents the marginal probability density function; X represents the sample data set of a certain process variable, and x is a sample in the set X; Y represents the sample data set of effluent COD concentration, and y is a sample in the set Y.

[0021] Furthermore, the specific steps for selecting certain process variables as auxiliary variables are as follows:

[0022] Step 4.1: According to Pareto's law, the Pearson correlation coefficient and the kernel density mutual information value are weighted and summed;

[0023] S(X,Y)=α×ρ(X,Y)+β×I(X,Y)

[0024] Where S(·,·) represents the weighted sum, α and β are weighting coefficients; X represents the sample data set of a certain process variable, Y represents the sample data set of effluent COD concentration; ρ(X,Y) represents the Pearson correlation coefficient between X and Y, and I(X,Y) represents the kernel density mutual information value between X and Y;

[0025] Step 4.2: Given a threshold λ for variable selection, if S(·,·) is greater than the threshold λ, then the process variable is closely related to the effluent COD concentration, and the process variable is selected as the preferred auxiliary variable. If S(·,·) is less than the threshold λ, then the process variable is not strongly related to the effluent COD concentration, and the process variable is removed.

[0026] Furthermore, the established neural network is a feedforward neural network with a single hidden layer. The auxiliary variables are transmitted from the input layer through the hidden layer to the output layer, represented as follows:

[0027]

[0028]

[0029] Where, x i h represents the input auxiliary variable data. j and The hidden layer and the output layer represent the outputs, f(·) and g(·) represent the non-linear activation functions, and ω and ω represent the non-linear activation functions. Let b represent the weight matrix and This represents the bias vector, where m and l represent the number of neurons in the input layer and hidden layer, respectively, with m set according to the number of auxiliary variables.

[0030] Furthermore, the process of adjusting the neural network parameters with the optimization objectives of minimizing point prediction error and minimizing distribution error is as follows:

[0031] Step 5.1: Determine the gradient steepest descent direction parameters of the neural network when minimizing the point prediction error;

[0032]

[0033]

[0034]

[0035] Where E represents the point prediction error between the actual and predicted values ​​of the effluent COD concentration, and y i and These represent the actual and predicted values ​​of COD concentration in the effluent, respectively, where n represents the number of samples. This represents the gradient steepest descent direction parameter when the point prediction error is minimized;

[0036] Step 5.2: Determine the gradient steepest descent direction parameters of the neural network when minimizing the distribution error;

[0037]

[0038]

[0039]

[0040]

[0041] Where J represents the distribution prediction error between the actual and predicted values ​​of the effluent COD concentration, and T tar and T out Let φ(x) represent the kernel density estimates of the actual and predicted COD concentrations in the effluent, respectively, where φ(x) represents the Gaussian kernel function and h represents the window width. This represents the gradient steepest descent direction parameter when the distribution error is minimized;

[0042] Step 5.3: Update the network parameters using the steepest descent direction of the gradient when minimizing the point prediction error and the distribution error, until the detection accuracy requirements are met;

[0043]

[0044]

[0045] Where η represents the learning step size of the neural network.

[0046] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor enables the processor to implement the soft measurement method for effluent COD concentration as described above.

[0047] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the soft measurement method for effluent COD concentration as described above.

[0048] Beneficial effects

[0049] The present invention provides a soft measurement method for effluent COD concentration based on a dual-error optimized neural network. First, the Pearson correlation coefficient and kernel density mutual information value between the process variable and the effluent COD concentration are calculated. Based on the Pareto principle, these two values ​​are fused, and the optimal process variable is selected as the auxiliary variable. Then, a single-hidden-layer feedforward neural network is established, with the optimization objectives of minimizing point prediction error and distribution error. The network parameters are adjusted to improve the model's training efficiency and the accuracy of the detection results. Finally, the test data is input into the final neural network, and the output is the detected value of the effluent COD concentration. Compared with existing technologies, the present invention has the following advantages:

[0050] 1. By using the Pareto principle to weighted sum the Pearson correlation coefficient and kernel density mutual information value as the basis for variable selection, the optimal process variable can be selected as the auxiliary variable, thereby reducing the dimensionality of sample data and the complexity of the model.

[0051] 2. Simultaneously using the minimization of point prediction error and the minimization of distribution error as optimization objectives, adjusting network parameters can improve the training efficiency of neural networks and avoid getting trapped in local optima. Attached Figure Description

[0052] Figure 1 This is a flowchart of the algorithm of the method of the present invention;

[0053] Figure 2 Here is a simplified structural diagram of the BSM1 wastewater treatment simulation platform;

[0054] Figure 3 Histogram of variable selection results for the method of this invention;

[0055] Figure 4 This is a graph showing the COD concentration detection results of the effluent from the method of the present invention;

[0056] Figure 5 This is an error graph showing the COD concentration detection results of the effluent from the method of the present invention. Detailed Implementation

[0057] Figure 1 This is an algorithm flowchart of the method of the present invention. The present invention will be further described below with reference to specific embodiments.

[0058] This embodiment uses the BSM1 wastewater treatment simulation platform as an example. Figure 2This is a simplified structural diagram of the BSM1 wastewater treatment simulation platform, comprising five biological reaction tanks and one secondary sedimentation tank. The platform treats an average of 20,000 m³ of wastewater per day, with an average biodegradable COD concentration of 300 mg / L. The first two reaction tanks are sealed anoxic tanks, while the latter three are aerobic tanks. Organic matter removal involves nitrification and denitrification reactions. Sampling points for process variables are located at the inlet, the five biological reaction tanks, the secondary sedimentation tank, and the recirculation stage. Each sampling point includes 15 process variables, with the effluent COD concentration as the output variable. A 14-day wastewater treatment process was simulated under sunny conditions, collecting 1344 sets of sample data, with a sampling period of 15 minutes for each process variable. Data from the first seven days were used as training data, and data from the last seven days were used as test data.

[0059] The specific implementation steps are as follows:

[0060] Step 1: Collect process variables from the BSM1 wastewater treatment simulation platform for 14 days, including effluent COD concentration and 120 other easily measurable variables (water flow rate, pH value, ammonia nitrogen concentration, dissolved oxygen concentration, total suspended solids, nitrate concentration, nitrite concentration, autotrophic bacteria activity, heterotrophic bacteria activity, etc.).

[0061] Step 2: Duplicate data were removed and missing data were added from the 1344 collected sample data sets, resulting in a sample data set of 1344 × 10⁹ (including 108 easily measurable process variables and 1 effluent COD concentration). Then, to eliminate calculation errors caused by different dimensions among the dependent variables, normalization was performed as follows:

[0062]

[0063] Where x represents the original sample data, x' represents the normalized sample data, and x min and x max These are the minimum and maximum values ​​in the original sample dataset, respectively.

[0064] Step 3: Calculate the Pearson correlation coefficient and kernel density mutual information value between the effluent COD concentration and the other 108 process variables. The specific steps are as follows:

[0065] Step 3.1: Calculate the Pearson correlation coefficient;

[0066]

[0067] Where ρ(·,·) represents the Pearson correlation coefficient, Cov(·,·) represents the covariance, Var(·) represents the standard deviation, X represents the sample data set of a certain process variable, and Y represents the sample data set of effluent COD concentration.

[0068] Step 3.2: Calculate the kernel density mutual information value;

[0069]

[0070] Where I(·,·) represents the kernel density mutual information value, p(·,·) represents the joint probability density function, and p(·) represents the marginal probability density function.

[0071] Step 4: Based on Pareto's law, the Pearson correlation coefficient and kernel density mutual information value are weighted and summed to select the optimal process variable as the auxiliary variable. The specific steps are as follows:

[0072] Step 4.1: Calculate the weighted sum of the Pearson correlation coefficient and the kernel density mutual information value according to Pareto's law;

[0073] S(X,Y)=α×ρ(X,Y)+β×I(X,Y)

[0074] Where S(·,·) denotes the weighted sum, according to Pareto's law α = 0.8 and β = 0.2;

[0075] Step 4.2: Given a threshold λ = 0.9 for the selected variable, when S(·,·) is greater than the threshold λ, the process variable is closely related to the effluent COD concentration, and the process variable is selected as the auxiliary variable for modeling. When S(·,·) is less than the threshold λ, the process variable is not strongly related to the effluent COD concentration, and the process variable is removed.

[0076] Variable selection results are as follows Figure 3 As shown, a total of 28 process variables closely related to the effluent COD concentration were selected as auxiliary variables.

[0077] Step 5: Establish a single-hidden-layer feedforward neural network, using auxiliary variables from the sample data as input and effluent COD concentration as output. Simultaneously, optimize the network parameters by minimizing point prediction error and distribution error, improving the model's training efficiency and the accuracy of the detection results. Specific steps are as follows:

[0078] Step 5.1: Establish a feedforward neural network with a single hidden layer. The network structure is: input layer 28 - hidden layer 10 - output layer 1. Data information is transmitted from the input layer to the output layer through the hidden layer.

[0079]

[0080]

[0081] Where, x i h represents the input auxiliary variable data. j and The hidden layer and the output layer represent the outputs, f(·) and g(·) represent the non-linear activation functions, and ω and ω represent the non-linear activation functions. Let b represent the weight matrix and The paranoia vector is represented by m=28 and l=10, which represent the number of neurons in the input layer and the hidden layer, respectively.

[0082] Step 5.2: Determine the gradient steepest descent direction parameters of the neural network when the point prediction error is minimized;

[0083]

[0084]

[0085]

[0086] Where E represents the point prediction error between the actual and predicted values ​​of the effluent COD concentration, and y i and These represent the actual and predicted values ​​of COD concentration in the effluent, respectively, where n represents the number of samples. This represents the gradient steepest descent direction parameter when the point prediction error is minimized;

[0087] Step 5.3: Determine the gradient steepest descent direction parameters of the neural network when minimizing the distribution error;

[0088]

[0089]

[0090]

[0091] Where J represents the distribution prediction error between the actual and predicted values ​​of the effluent COD concentration, and T... tar and T out φ(x) represents the kernel density estimates of the actual and predicted COD concentrations in the effluent, respectively, where φ(x) represents the Gaussian kernel function, n represents the number of samples, and h represents the window width. This represents the gradient steepest descent direction parameter when the distribution error is minimized;

[0092] Step 5.4: Update the network parameters using the steepest descent direction of the gradient when minimizing the point prediction error and the distribution error, until the detection accuracy requirements are met;

[0093]

[0094]

[0095] Where η = 10 -5 This represents the learning step size of the neural network;

[0096] Step 6: Input 672 sets of test data into the final neural network, and the output result is the detected value of COD concentration in the effluent.

[0097] In this embodiment, the detection results of the effluent COD concentration are as follows: Figure 4 As shown, the X-axis represents the number of test data samples, and the Y-axis represents the true and detected values ​​of effluent COD concentration, in mg / L. The solid line represents the true value of effluent COD concentration, and the dashed line represents the detected value of effluent COD concentration. The detection error of effluent COD concentration is as follows: Figure 5 As shown, the X-axis represents the number of test data samples, and the Y-axis represents the detection error value, with the unit being mg / L.

[0098] Since the above embodiments are for verification of the BSM1 wastewater treatment simulation platform, they do not limit the scope of implementation of this invention. This invention is also applicable to the detection of COD concentration in the effluent of other actual wastewater treatment plants.

[0099] The above embodiments are preferred embodiments of this application. Those skilled in the art can make various changes or improvements based on them. Without departing from the overall concept of this application, these changes or improvements should fall within the scope of protection claimed in this application.

Claims

1. A soft measurement method for effluent COD concentration based on a dual-error optimized neural network, characterized in that, include: Step 1: Collect sample data of process variables from the target wastewater treatment plant; The process variables include effluent COD concentration and other process variables; Step 2: Preprocess the collected sample data; Step 3: Calculate the Pearson correlation coefficient and kernel density mutual information value between the water COD concentration and each of the other process variables; Step 4: According to Pareto's law, the Pearson correlation coefficient and the kernel density mutual information value are weighted and summed, and some process variables are selected as auxiliary variables based on the weighted sum value; Step 5: Establish a neural network with auxiliary variables from the sample data as input and effluent COD concentration as output. The optimization objectives are to minimize point prediction error and distribution error. Adjust the neural network parameters. The neural network with the parameters finally adjusted is recorded as the soft measurement model of effluent COD concentration. The established neural network is a feedforward neural network with a single hidden layer. The auxiliary variables are transmitted from the input layer through the hidden layer to the output layer, as shown below: ; ; in, This represents the input auxiliary variable data. and This represents the output of the hidden layer and the output layer. and Represents a non-linear activation function. and Represents the weight matrix. and Represents the paranoia vector. and This indicates the number of neurons in the input layer and the hidden layer. Set according to the number of auxiliary variables; The process of adjusting neural network parameters with the optimization objectives of minimizing point prediction error and minimizing distribution error is as follows: Step 5.1: Determine the gradient steepest descent direction parameters of the neural network when minimizing the point prediction error; ; ; ; in, This represents the point prediction error between the actual and predicted values ​​of the effluent COD concentration. and These represent the actual and predicted values ​​of COD concentration in the effluent, respectively. Indicates the number of samples. This represents the gradient steepest descent direction parameter when the point prediction error is minimized; Step 5.2: Determine the gradient steepest descent direction parameters of the neural network when minimizing the distribution error; ; ; ; ; in, This represents the distribution prediction error between the actual and predicted values ​​of effluent COD concentration. and These represent the kernel density estimates of the actual and predicted COD concentrations in the effluent, respectively. Represents the Gaussian kernel function. Indicates the window width. This represents the gradient steepest descent direction parameter when the distribution error is minimized; Step 5.3: Update the network parameters using the steepest descent direction of the gradient when minimizing the point prediction error and the distribution error, until the detection accuracy requirements are met; ; ; in, This represents the learning step size of the neural network; Step 6: When it is necessary to measure the actual effluent COD concentration, the process variable in the auxiliary variables is collected and input into the soft measurement model of effluent COD concentration. The output result is the effluent COD concentration.

2. The soft measurement method for effluent COD concentration according to claim 1, characterized in that, The other process variables include, but are not limited to: water flow rate, pH value, ammonia nitrogen concentration, dissolved oxygen concentration, total suspended solids, nitrate concentration, nitrite concentration, autotrophic bacteria bioactivity, and heterotrophic bacteria bioactivity.

3. The soft measurement method for effluent COD concentration according to claim 1, characterized in that, Sample data preprocessing includes removing duplicate data, adding missing data, and data normalization.

4. The soft measurement method for effluent COD concentration according to claim 1, characterized in that, The Pearson correlation coefficient between effluent COD concentration and each of the other process variables is calculated as follows: ; in, This represents the Pearson correlation coefficient. Describing covariance, Indicates standard deviation, A sample set representing a certain process variable. A set of sample data representing the COD concentration in effluent.

5. The soft measurement method for effluent COD concentration according to claim 1, characterized in that, The kernel density mutual information value between the effluent COD concentration and each of the other process variables is calculated using the following formula: ; in, Represents the kernel density mutual information value. Denotes the joint probability density function. Represents the marginal probability density function; A sample set representing a certain process variable. For set A sample in; A set of sample data representing the COD concentration in effluent. For set A sample in the sample.

6. The soft measurement method for effluent COD concentration according to claim 1, characterized in that, The specific steps for selecting certain process variables as auxiliary variables are as follows: Step 4.1: According to Pareto's law, the Pearson correlation coefficient and the kernel density mutual information value are weighted and summed; ; in, Indicates a weighted sum. and These are weighting coefficients; A sample set representing a certain process variable. A set of sample data representing the COD concentration in effluent; express and The Pearson correlation coefficient between them express and The kernel density mutual information value between them; Step 4.2, Threshold for variable selection ,when Greater than the threshold When this process variable is closely related to the effluent COD concentration, it is selected as the preferred auxiliary variable. Less than the threshold If the process variable is not strongly correlated with the effluent COD concentration, then the process variable should be removed.

7. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, the processor causes the processor to implement the method as described in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Deep learning-based dual-supervision image defogging method, system, medium and equipment

    CN110097519A

  • Knowledge-based robust effluent ammonia nitrogen soft measurement method

    CN110542748A