Soft measurement method for outlet ammonia nitrogen concentration based on modular random configuration network

CN116796145BActive Publication Date: 2026-08-28BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310564457.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2026-08-28
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

然而,当前基于随机配置网络开发的出水氨氮浓度软测量模型的研究还很少

Benefits of technology

[0055]This invention addresses the challenge of real-time and accurate detection of effluent ammonia nitrogen concentration in urban wastewater treatment processes due to the complex influence of multiple operating conditions and uncertainties. It proposes a soft-sensor modeling method based on a modular stochastic configuration network. This method simulates the human brain's "divide and conquer" approach to handle complex, multi-condition nonlinear modeling tasks, thereby achieving real-time and accurate detection of effluent ammonia nitrogen concentration. Specifically, this soft-sensor method first extracts variables with significant impact on effluent ammonia nitrogen concentration through grey relational analysis to improve the model's measurement performance; then, it reduces the modeling complexity of a single model by decomposing the complex task; further, it learns from sub-tasks using an improved stochastic configuration network to improve the overall learning efficiency and approximation accuracy of the modular network. This soft-sensor method has the following characteristics:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116796145B_ABST
    Figure CN116796145B_ABST
Patent Text Reader

Abstract

The soft measurement method for effluent ammonia nitrogen concentration based on modular stochastic configuration network relates to the field of intelligent modeling of sewage treatment process. The method solves the problem that the effluent ammonia nitrogen concentration in sewage treatment is difficult to detect accurately in real time, based on the respective advantages of brain-like modular network and stochastic configuration network in processing complex modeling tasks. The method first extracts the main variables affecting the effluent ammonia nitrogen concentration by using grey correlation analysis; then, task decomposition is carried out based on fuzzy clustering method to alleviate the complexity of each modeling task; then, the "divide and conquer" idea is used to construct the corresponding stochastic configuration network sub-model for each sub-task after decomposition, and the results output by each sub-model are integrated to improve the measurement effect of the effluent ammonia nitrogen concentration, so as to realize the accurate and effective measurement of the effluent ammonia nitrogen concentration in the sewage treatment process. The effectiveness and superiority of the method in detecting the effluent ammonia nitrogen concentration are verified based on the relevant water quality data collected from an actual sewage treatment plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial artificial intelligence and is directly applied to the field of intelligent modeling of wastewater treatment processes. Background Technology

[0002] Wastewater treatment is a crucial foundation for water resource management and prevention. Effective treatment of wastewater from domestic and industrial processes allows for a degree of water resource reuse. During wastewater treatment, effluent quality monitoring is essential for management and control; monitoring key water quality parameters can effectively prevent substandard discharges due to inadequate treatment. Among these parameters, effluent ammonia nitrogen concentration is a critical indicator for assessing effluent quality. Excessively high ammonia nitrogen concentrations can severely endanger aquatic life such as fish and shrimp, negatively impacting the surrounding ecosystem. Real-time and effective monitoring of effluent ammonia nitrogen concentration not only allows for the timely detection of deficiencies in the wastewater treatment process but also provides a basis for relevant control measures, ensuring that effluent ammonia nitrogen concentrations meet discharge standards. Therefore, real-time and effective monitoring of effluent ammonia nitrogen concentration is of paramount importance.

[0003] Currently, the methods commonly used in wastewater treatment plants both domestically and internationally to measure ammonia nitrogen concentration in effluent include electrochemical analysis, instrumental analysis, and spectrophotometry. While these methods offer high accuracy in measuring effluent ammonia nitrogen concentration, their long measurement cycles, complex processes, high costs, high consumption, and susceptibility to human factors make it difficult to guarantee precise and efficient measurement. In recent years, with the continuous development of industrial information technology and the construction of smart factories based on artificial intelligence, wastewater treatment processes have become an important future development direction. Notably, intelligent modeling methods, represented by data-driven modeling soft measurement technology, are widely used in key water quality detection areas of wastewater treatment processes due to their advantages such as simple operation, ease of execution, and rapid measurement. Among these, stochastic configuration networks, as a classic model of randomized learning, are widely used in data-driven modeling of industrial processes due to their advantages such as fast learning speed, strong approximation performance, and computational simplicity. However, actual wastewater treatment processes are characterized by multiple operating conditions and complex uncertainties due to a series of biochemical reactions and various complex uncertainties. Single network models, due to their inherent limitations in memory retention and general modeling, have certain limitations in handling such complex multi-condition modeling tasks. Furthermore, most current modular neural networks use traditional feedforward neural networks as their base model. These models suffer from difficulties in quickly learning network parameters and structures, resulting in compromised modeling accuracy and efficiency. In contrast, stochastic configuration networks offer advantages such as fast learning speed and strong approximation performance. However, research on soft measurement models for effluent ammonia nitrogen concentration based on stochastic configuration networks is still limited. To further improve the approximation performance of stochastic configuration networks and enhance the modeling performance for effluent ammonia nitrogen concentration, this invention designs a modular stochastic configuration network based on the brain-like modular approach and the respective advantages of stochastic configuration networks. This network uses a divide-and-conquer strategy to modularize complex modeling tasks and employs an improved stochastic configuration network to learn from sub-tasks, thereby improving the overall modeling performance of the network and achieving real-time and accurate detection of effluent ammonia nitrogen concentration in wastewater treatment processes. Summary of the Invention

[0004] This invention designs a modular stochastic configuration network soft measurement model for effluent ammonia nitrogen concentration with divide-and-conquer and self-learning characteristics. Inspired by the human brain's "divide-and-conquer" and regional collaborative characteristics in information analysis and processing, as well as the brain's modular partitioning structure, this method simulates the structure and functional characteristics of the human cerebral cortex to establish corresponding stochastic configuration network models for different learning tasks, thereby achieving real-time and accurate detection of effluent ammonia nitrogen concentration. Specifically, the model first analyzes the correlation between various water quality parameters and the measured effluent ammonia nitrogen concentration based on grey relational analysis theory and expert knowledge to obtain auxiliary variables with strong correlation to effluent ammonia nitrogen, reducing model complexity. Then, it decomposes the data generated by the system into tasks using fuzzy mean clustering; subsequently, based on each sub-task, it establishes corresponding sub-models using the improved stochastic configuration network. Because the randomly configured network in the sub-model can autonomously construct the corresponding network structure and parameters according to the distribution characteristics of the data in the sub-tasks to be processed, this soft measurement method can ensure the network's general approximation capability while possessing autonomous learning and data-dependent characteristics, thereby achieving accurate and efficient measurement of ammonia nitrogen concentration in effluent. To achieve the above objectives, this invention adopts the following technical solution:

[0005] Step 1: Data collection and pretreatment of relevant water quality during wastewater treatment; Collect relevant actual water quality data of a wastewater treatment plant during wastewater treatment, including: total phosphorus TP1 in the influent, oxidation-reduction potential (OPR1) at the end of the anaerobic process, dissolved oxygen concentration (DO1) at the beginning of the aerobic process, temperature, dissolved oxygen concentration (DO2) at the end of the aerobic process, suspended solids concentration (TSS), pH value, oxidation-reduction potential (OPR2) in the effluent, nitrate nitrogen concentration (NO3-N) in the effluent, total phosphorus TP2 in the effluent, and ammonia nitrogen concentration in the effluent, totaling 11 water quality parameters;

[0006] Step 2: Select auxiliary variables; use grey relational analysis to measure the correlation between the relevant water quality parameters collected during the above wastewater treatment process and the effluent ammonia nitrogen concentration;

[0007] Assume there are d = 10 comparison sequences and 1 reference sequence, with each sequence containing N samples. Let X0 = {X0(1), X0(2), ..., X0(N)} represent the reference sequence selected in the grey relational analysis, and the comparison sequence be X... i ={X i (1),X i (2),...,X i (N)}, i=1,2,3,…,d; To more accurately analyze the correlation coefficient between the reference sequence and each comparison sequence, here we take each comparison sequence X i The data in the reference sequence X0 were normalized as follows to mitigate the impact of different units on the analysis results.

[0008]

[0009]

[0010] By analyzing the reference sequence X0 and the comparison sequence X i After normalization, a new set of reference sequence x0 and comparison sequence x is obtained. i Wherein, the grey relational coefficient χ between the i-th comparison sequence and the reference sequence at time n. i (n) can be calculated using the following formula:

[0011]

[0012] Where x0(n) is the sample value of the reference sequence at time n, x i (n) represents the sample value of the comparison sequence at time n; This represents the value that minimizes the difference between all comparison sequences and the reference sequence; Let Δ represent the value that represents the largest difference between all comparison sequences and the reference sequence; denoted as Δ. i (n)=|x0(n)-x i If (n)|, then the grey relational coefficient between the i-th comparison sequence and the reference sequence (i.e., the ammonia nitrogen concentration in the effluent) is:

[0013]

[0014] Where η = 0.5 is the resolution coefficient. Therefore, the average correlation coefficients at different times are used to compare the overall correlation between the sequences, obtaining the correlation between the comparison sequence and the reference sequence, calculated as follows:

[0015]

[0016] To facilitate the analysis and comparison of the correlation between various water quality parameters and effluent ammonia nitrogen concentration, the correlation between the parameters obtained from formula (5) and ammonia nitrogen concentration is statistically analyzed. The larger the correlation coefficient, the higher the correlation between the parameter variable and the effluent ammonia nitrogen concentration. The correlation between each water quality parameter and effluent ammonia nitrogen concentration is calculated using formula (5), and the results are sorted. Six variables with a high correlation with effluent ammonia nitrogen concentration are selected as auxiliary variables for further study: aerobic front-end DO1, effluent pH, anaerobic terminal ORP1, influent TP, effluent NO3-N, and temperature T.

[0017] Step 3: Design a soft measurement model for effluent ammonia nitrogen concentration based on a modular stochastic configuration network;

[0018] The construction process of this soft measurement model includes five modules: input layer, task decomposition layer, task allocation layer, sub-network module, and output integration layer. The main construction process and specific functional implementation are detailed below:

[0019] Step 3.1: The input layer is used to transmit data; here, the data is transmitted to the task decomposition module through the input layer, where the number of nodes in the input layer is 6, which is the same as the number of input auxiliary variables.

[0020] Step 3.2: The task decomposition layer processes the input data, aiming to break down a complex learning task into several relatively simple subtasks. This module uses fuzzy mean clustering to cluster the acquired data; using C... k x represents the cluster centers of each category. n Let μ represent the nth input sample. nk Indicates sample x n For the k-th cluster center C k The membership degree. Then, the samples can be clustered by minimizing the following loss function J:

[0021]

[0022] Where N is the number of samples, K is the number of clusters, and α is a constant, usually set to α = 2. ||*|| represents the distance metric for data similarity; here, the most common Euclidean distance is used. Membership degree μ nk This represents the degree to which the nth sample belongs to the kth class, and it is calculated using the following formula:

[0023]

[0024] Where C j Let C represent the j-th cluster center, j = 1, 2, ..., K; and the corresponding cluster center C. k It can be calculated using the following formula:

[0025]

[0026] Therefore, after the sample is decomposed by the task decomposition module, multiple clusters and cluster centers are obtained.

[0027] Step 3.3: In the task allocation layer, the membership values ​​of each sample data to different clusters obtained by the fuzzy mean clustering method used by the task decomposition module are first regarded as the initial weight coefficients for the sample to belong to each cluster; then, an allocation threshold v is set in the task allocation layer. p=0.4, and its value range is [0,1], representing the lowest membership value of the sample to each cluster; if the weight coefficient of a sample is greater than this threshold, the sample is assigned to the corresponding subtask; at the same time, the task allocation layer will count the number of times the sample is assigned to each task and record the corresponding weight coefficient, so as to facilitate the weighted integration of the output layer.

[0028] Step 3.4: The sub-network module mainly models each sub-task. In modular neural networks, the modeling performance of the sub-networks is crucial to the overall modeling performance of the network. Randomized networks have advantages such as fast learning speed and strong approximation ability when handling large-scale complex data modeling tasks, and can autonomously generate corresponding network structures according to modeling requirements, significantly reducing the workload of parameter tuning. Therefore, randomized networks are used as sub-networks in the modular neural network to construct corresponding sub-models for each sub-task obtained after task decomposition.

[0029] Assume the task allocation module assigns the acquired samples {X,Y} to the corresponding subtasks t, t=1,2,…K, where K is the number of subtasks and the number of clusters. The data for each task is represented as {X}. t ,Y t};in, Let be the input data of the network in the t-th subtask, and let be the corresponding output data. Nt represents the number of samples assigned to task t. Assume that a randomly configured network with L-1 hidden nodes has been constructed for the t-th subtask, given the objective function... The network output can then be represented as:

[0030]

[0031] Among them, X t Let β represent the input data corresponding to the t-th subtask. j =[β j1 ,β j2 ,...,β jM [] represents the output weight of the j-th hidden node corresponding to this sub-network; h j Let be the output vector of the j-th hidden node, which is represented as:

[0032]

[0033] g(·) is the activation function of the hidden layer neurons. In a randomly configured network, the sigmoid function is selected as the activation function of the hidden layer neurons, i.e.:

[0034]

[0035] Where, <·> represents the inner product of Euclidean space, wj , b j are respectively the input weight and threshold of the j-th hidden node; the residual vector of the current network is defined as:

[0036]

[0037] the corresponding output residual value is

[0038]

[0039] M is the number of output nodes; if the output residual of the current network fails to meet the preset error tolerance requirement, that is e P = 0.001, the network will select a new hidden layer neuron node according to the inequality constraint condition (14), and obtain the L-th hidden layer node parameter g L (w L and b L ). Wherein, the inequality constraint condition for the network to select hidden layer node parameters is:

[0040]

[0041] in the formula: 0<ε<1 and 0<q<1 are error reduction factors, set ε=0.015, q=0.8; h L (X t ) = g L (w L , X t , b L ) represents the output of the L-th hidden node, which can be calculated by formula (10-11), e L-1,m is the output residual of the m-th output layer node when the network constructs L-1 hidden layer nodes, m=1,2,…,M, then for the t-th subtask, the output weight β of the hidden layer node after constructing the L-th hidden node is t :

[0042]

[0043] then, after the L-th hidden layer node is established, the output of the network is:

[0044]

[0045] at this time, the output residual of the network is:

[0046]

[0047] then, calculate the output error value of the current network according to formula (13) and judge whether the output error of the network meets the preset error requirement, that is If the condition is met, the SCNs construction is complete; otherwise, continue to build the network by adding new hidden layer node parameters g(w,b) according to the inequality constraint (14) to reduce the network output error. At this time, the number of hidden nodes in the subnetwork is L = L + 1, until the termination condition is met. or L≥L max ,L max =50.

[0048] Step 3.5: The output integration layer integrates the outputs of each sub-network to obtain the final overall output of the network. In modular stochastic configuration networks, the output integration strategy is closely related to the characteristics of task decomposition. Therefore, a weighted average method is used in this output strategy to integrate the results of each sub-network; assuming that for the nth input sample x... n The probability of being assigned to each subtask is like Then set it to Where t = 1, 2, ..., K is the number of subtasks, and v p The threshold for each task is set, with a value ranging from [0,1], representing the minimum membership value of a sample belonging to each cluster. Here, v is set to... p =0.4; then for The following processing is performed to obtain the final output weights for the t-th task corresponding to this sample.

[0049]

[0050] Therefore, for the nth input sample, the corresponding network's final output y n for:

[0051]

[0052] in, Indicates sample x n The output on task t can be calculated using formula (16); the unselected subnetworks do not contribute to the overall output of the network; therefore, the final output of the network corresponding to the nth sample can be calculated using formula (19), which is the weighted sum of the samples assigned to the corresponding subnetworks.

[0053] Step 3: Based on Step 2, the constructed modular random configuration network is obtained, and the ammonia nitrogen concentration in the effluent of the urban sewage treatment process is measured based on the constructed model to obtain the measurement results of the effluent ammonia nitrogen concentration.

[0054] This invention differs from traditional neural network methods and has the following characteristics:

[0055] This invention addresses the challenge of real-time and accurate detection of effluent ammonia nitrogen concentration in urban wastewater treatment processes due to the complex influence of multiple operating conditions and uncertainties. It proposes a soft-sensor modeling method based on a modular stochastic configuration network. This method simulates the human brain's "divide and conquer" approach to handle complex, multi-condition nonlinear modeling tasks, thereby achieving real-time and accurate detection of effluent ammonia nitrogen concentration. Specifically, this soft-sensor method first extracts variables with significant impact on effluent ammonia nitrogen concentration through grey relational analysis to improve the model's measurement performance; then, it reduces the modeling complexity of a single model by decomposing the complex task; further, it learns from sub-tasks using an improved stochastic configuration network to improve the overall learning efficiency and approximation accuracy of the modular network. This soft-sensor method has the following characteristics:

[0056] 1) Use grey relational analysis to screen water quality parameter variables to reduce model complexity and computational load;

[0057] 2) In the task decomposition layer, fuzzy mean clustering is used to decompose the data generated by the system into tasks, so that the data of each sub-task has a high degree of similarity and consistent distribution to improve the learning efficiency and generalization performance of the sub-model.

[0058] 3) Improve the overall learning efficiency and modeling performance of the modular stochastic configuration network by using an improved stochastic configuration algorithm as a sub-model to model each sub-task.

[0059] 4) By using an improved randomly constructed network as the sub-model, the sub-model can autonomously construct the corresponding network structure and parameters according to the sub-task to be processed. Therefore, the sub-network ensures the model's general approximation ability while possessing good autonomous learning and data dependency characteristics; Attached Figure Description

[0060] Figure 1 This is a basic strategy diagram of the M-SCNs soft measurement model of the present invention;

[0061] Figure 2 A basic structural diagram of a modular random configuration network;

[0062] Figure 3 This is a graph showing the test results of ammonia nitrogen concentration in the effluent for this example;

[0063] Figure 4 This is a graph showing the error in the ammonia nitrogen concentration test of the effluent in this example. Detailed Implementation

[0064] Relevant actual water quality data from a wastewater treatment plant were collected, including: influent total phosphorus (TP1), anaerobic terminal oxidation-reduction potential (OPR1), aerobic front-end dissolved oxygen concentration (DO1), temperature, aerobic terminal dissolved oxygen concentration (DO2), suspended solids concentration (TSS), pH value, effluent OPR2, effluent nitrate nitrogen concentration (NO3-N), effluent ammonia nitrogen concentration, and effluent total phosphorus (TP2). A total of 498 sets of experimental data were collected for research. The main steps are as follows:

[0065] Step 1: Selection of Auxiliary Variables

[0066] The following grey relational analysis is used to measure the correlation between the relevant water quality parameters collected during the above wastewater treatment process and the effluent ammonia nitrogen concentration, so as to obtain auxiliary variables with strong correlation with effluent ammonia nitrogen and thus improve the modeling accuracy.

[0067] The water quality parameters obtained, excluding the effluent ammonia nitrogen concentration, are used as comparison sequences, and the effluent ammonia nitrogen concentration is used as the reference sequence. Each sequence contains N samples. X0 = {X0(1), X0(2), ..., X0(N)} represents the reference sequence selected in the grey relational analysis, and the comparison sequence is X... i ={X i (1),X i (2),...,X i (N)}, i = 1, 2, 3, ..., d; d is the number of comparison sequences. To more accurately analyze the correlation coefficient between the reference sequence and each comparison sequence, we here... i The data in the reference sequence X0 were normalized as follows to mitigate the impact of different units on the analysis results.

[0068]

[0069]

[0070] By analyzing the reference sequence X0 and the comparison sequence X i After normalization, a new set of reference sequence x0 and comparison sequence x are obtained. i Wherein, the grey relational coefficient χ between the i-th comparison sequence and the reference sequence at time n. i (n) can be calculated using the following formula:

[0071]

[0072] Where x0(n) is the sample value of the reference sequence at time n, x i (n) represents the sample value of the comparison sequence at time n; let Δ i (n)=|x0(n)-x iIf (n)|, then the grey relational coefficient between the i-th comparison sequence and the reference sequence (i.e., the ammonia nitrogen concentration in the effluent) is:

[0073]

[0074] Where η = 0.5 is the resolution coefficient. Therefore, the average correlation coefficients at different times are used to compare the overall correlation between the sequences, obtaining the correlation between the comparison sequence and the reference sequence, calculated as follows:

[0075]

[0076] To facilitate the analysis and comparison of the correlation between various water quality parameters and effluent ammonia nitrogen concentration, the correlation between the parameters obtained from formula (24) and ammonia nitrogen concentration is statistically analyzed. The larger the correlation coefficient, the higher the correlation between the parameter variable and the effluent ammonia nitrogen concentration. Therefore, the correlation between the water quality parameters calculated by formula (24) and effluent ammonia nitrogen concentration is statistically analyzed, and six variables with relatively high correlation with effluent ammonia nitrogen concentration are obtained: aerobic front-end DO1, effluent pH, anaerobic terminal ORP1, influent TP, effluent NO3-N, and temperature T, which are used as auxiliary variables for research.

[0077] Step 2: Design a soft measurement model for effluent ammonia nitrogen concentration based on a modular stochastic configuration network;

[0078] The construction process of this soft measurement model includes five modules: input layer, task decomposition layer, task allocation layer, sub-network module, and output integration layer. The main construction process and specific functional implementation are detailed below:

[0079] Step 2.1: The input module is used to transmit data; here, the water quality data X after feature extraction is transmitted to the task decomposition module, where the number of nodes in the input layer is 6, consistent with the number of input auxiliary variables. This layer contains a total of 6 nodes, where each node represents a corresponding water quality variable.

[0080] Step 2.2: Based on the k-means method, set the cluster centers C and the number of clusters K for the samples, and use them to set the fuzzy membership function to fuzzify the input data. Here, the Gaussian function is chosen as the fuzzy membership function. In the fuzzy membership layer, each node represents a corresponding linguistic variable value. The output of the node in this layer is the membership function of the input component belonging to the fuzzy set. Therefore, the i-th input variable of the n-th input sample... The corresponding fuzzy membership degree is It can be obtained by calculating the Gaussian membership function, where the center of the Gaussian membership function is the cluster center C obtained by k-means, k = 1, 2, ... K, representing the fuzzy rule, which is the same as the number of clusters K.

[0081] In the task decomposition layer, fuzzy mean clustering is used to cluster the acquired data. The goal is to break down a complex learning task into several relatively simple subtasks; using C... k x represents the cluster centers of each category. n Let μ represent the nth input sample. nk Indicates sample x n For the k-th cluster center C k The membership degree. Then, the samples can be clustered by minimizing the following loss function J:

[0082]

[0083] Where N is the number of samples, K is the number of clusters, and α is a constant, usually set to α = 2. ||*|| represents the distance metric for data similarity; here, the most common Euclidean distance is used. Membership degree μ nk This represents the degree to which the nth sample belongs to the kth class, and it is calculated using the following formula:

[0084]

[0085] Where C j Let C represent the j-th cluster center, j = 1, 2, ..., K; and the corresponding k-th cluster center C. k It can be calculated using the following formula:

[0086]

[0087] Therefore, after the sample is decomposed by the task decomposition module, multiple clusters and cluster centers are obtained.

[0088] Step 2.3: In the task allocation layer, the membership values ​​of each sample data to different clusters obtained by the fuzzy mean clustering method used by the task decomposition module are regarded as the initial weight coefficients for the sample to belong to each cluster; then, an allocation threshold v is set in the task allocation layer. p Its value ranges from [0,1], representing the lowest membership degree of a sample belonging to each cluster. Here, v is set to... p =0.4; At the same time, the task allocation layer will count the number of times the sample is assigned to each task and record the corresponding weight coefficients, which will facilitate the weighted integration by the output layer.

[0089] Step 2.4: Construct corresponding sub-models using a randomized network for the data samples in each sub-task;

[0090] Assume the task allocation module assigns the acquired samples {X,Y} to the corresponding subtasks t, t=1,2,…K, where K is the number of subtasks and the number of clusters. The data for each task is represented as {X}.t ,Y t};in, Let be the input data of the network in the t-th subtask, and let be the corresponding output data. Nt represents the number of samples assigned to task t. Assume that a randomly configured network with L-1 hidden nodes has been constructed for the t-th subtask, given the objective function... The network output can then be represented as:

[0091]

[0092] Among them, X t Let β represent the input data corresponding to the t-th subtask. j =[β j1 ,β j2 ,...,β jM [] represents the output weight of the j-th hidden node corresponding to this sub-network; h j Let be the output vector of the j-th hidden node, which is represented as:

[0093]

[0094] g(·) is the activation function of the hidden layer neurons. In a randomly configured network, the sigmoid function is selected as the activation function of the hidden layer neurons, i.e.:

[0095]

[0096] Where, <·> represents the inner product of Euclidean space, w j b j Let be the input weights and threshold of the j-th hidden node, respectively; the residual vector of the current network is defined as:

[0097]

[0098] The corresponding output residual value is

[0099]

[0100] M is the number of output nodes; if the current network's output residual The preset error tolerance requirement was not met, i.e. The network will then select new hidden layer neurons based on the inequality constraints, thus obtaining the parameter g of the Lth hidden layer node. L (w L and b L The inequality constraints for selecting hidden layer node parameters in the network are as follows:

[0101]

[0102] In the formula: 0<ε<1 and 0<q<1 are error reduction factors, where ε=0.015 and q=0.8; h L (X t )=g L (w L ,X t ,b L ) represents the output of the L-th hidden node, which can be calculated by formula (10-11), e L-1,m is the output residual of the m-th output layer node when the network constructs L-1 hidden layer nodes, m=1,2,...,M, then for the t-th subtask, the output weight β of the hidden layer nodes after constructing the L-th hidden node is t :

[0103]

[0104] Then, the output of the network after the L-th hidden layer node is established is:

[0105]

[0106] At this time, the output residual of the network is:

[0107]

[0108] Then, calculate the current network output error value according to formula (32) and judge whether the network output error meets the preset error requirement, that is, if the requirement is met, the construction of SCNs is completed; otherwise, continue to add new hidden layer node parameters g(w,b) according to the inequality constraint (33) to construct the network and reduce the output error of the network, at this time, the number of hidden nodes of the sub-network is L=L+1, until the termination condition is satisfied or L≥L max ,L max =50.

[0109] Step 2.5: Perform integration processing on the output of each sub-network through the output integration layer to obtain the overall output of the final network.

[0110] Assume that for the n-th input sample x n , the probability of being assigned to each subtask is if then set it as where t=1,2,…K is the number of subtasks, v p is the allocation threshold of each task, its value range is the interval [0,1], representing the minimum membership value of the sample belonging to each cluster, where v p =0.4; then for The following processing is performed to obtain the final output weights for the t-th task corresponding to this sample.

[0111]

[0112] Therefore, for the nth input sample, the corresponding network's final output y n for:

[0113]

[0114] in, Indicates sample x n The output on task t can be calculated using formula (35); the unselected subnetworks do not contribute to the overall output of the network; therefore, the final output of the network corresponding to the nth sample can be calculated using formula (38), which is the weighted sum of the sample assigned to each corresponding subnetwork.

[0115] Step 3: Based on Step 2, the constructed modular random configuration network is obtained, and the ammonia nitrogen concentration in the effluent of the urban sewage treatment process is measured based on the constructed model to obtain the measurement results of the effluent ammonia nitrogen concentration.

Claims

1. A soft measurement method for effluent ammonia nitrogen concentration based on modular random configuration networks is characterized by, Includes the following steps: Step 1: Collect relevant water quality data during the wastewater treatment process; The following 11 water quality parameters were collected from a wastewater treatment plant during wastewater treatment: total phosphorus TP1 in the influent, oxidation-reduction potential (OPR1) at the end of the anaerobic process, dissolved oxygen concentration (DO1) at the beginning of the aerobic process, temperature, dissolved oxygen concentration (DO2) at the end of the aerobic process, suspended solids concentration (TSS), pH value, oxidation-reduction potential (OPR2) in the effluent, nitrate nitrogen concentration (NO3-N) in the effluent, ammonia nitrogen concentration in the effluent, and total phosphorus TP2 in the effluent. Step 2: Select auxiliary variables; use the following grey relational analysis method to measure the correlation between the relevant water quality parameters collected in the above wastewater treatment process and the effluent ammonia nitrogen concentration, so as to obtain auxiliary variables with strong correlation with effluent ammonia nitrogen. The obtained water quality parameters, excluding the effluent ammonia nitrogen concentration, were used as comparison sequences, with the effluent ammonia nitrogen concentration used as a reference sequence, and each sequence contained N samples. This represents the reference sequence selected in the grey relational analysis, and the comparison sequence is... Let i = 1, 2, 3, ..., d; d is the number of comparison sequences; to accurately analyze the correlation coefficient between the reference sequence and each comparison sequence, we will consider the following for each comparison sequence. and reference sequence The data in the analysis were normalized as follows to mitigate the impact of different units on the analysis results; ; ; By using the reference sequence and comparison sequences After normalization, a new set of reference sequences is obtained. and comparison sequences Wherein, the grey relational coefficient between the i-th comparison sequence and the reference sequence at time n. Calculated using the following formula: ; in, The sample value of the reference sequence at time n. To compare the sample values ​​of the sequence at time n, =0.5 is the resolution coefficient; This represents the value that minimizes the difference between all comparison sequences and the reference sequence; This represents the value that differs most from the reference sequence among all comparison sequences; denoted as... The grey relational coefficient between the i-th comparison sequence and the reference sequence (i.e., the ammonia nitrogen concentration in the effluent) is: ; Therefore, the correlation coefficients at different times are averaged to make an overall comparison of the correlation between the sequences, and the correlation between the comparison sequence and the reference sequence is obtained, as calculated below: ; To facilitate comparison and analysis of the correlation between various water quality parameters and effluent ammonia nitrogen concentration, the correlation between various parameters obtained by formula (5) and ammonia nitrogen concentration is statistically analyzed. The larger the correlation coefficient, the higher the correlation between the parameter variable and effluent ammonia nitrogen concentration. Therefore, the correlation between various water quality parameters calculated by formula (5) and effluent ammonia nitrogen concentration is statistically analyzed. Six variables with a high correlation with effluent ammonia nitrogen concentration are obtained: aerobic front-end DO1, effluent pH, anaerobic end ORP1, influent TP, effluent NO3-N and temperature T. These are used as auxiliary variables for research. The input data after feature extraction is represented by X and Y represents the corresponding effluent ammonia nitrogen concentration. Step 3: Design a soft measurement model for effluent ammonia nitrogen concentration based on a modular stochastic configuration network; The construction process of this soft measurement model includes five modules: input layer, task decomposition layer, task allocation layer, sub-network module, and output integration layer. The following details the model construction process and specific functional implementation: Step 3.1: The input layer is used to transmit data X; here, the data is transmitted to the task decomposition module through the input layer, where the number of nodes in the input layer is 6, which is the same as the number of input auxiliary variables; Step 3.2: The task decomposition layer processes the input data, aiming to break down a complex learning task into several relatively simple subtasks. This module uses fuzzy mean clustering to cluster the acquired dataset. (The last part, "using C," appears to be an unrelated fragment and is left untranslated.) k x represents the cluster centers of each category. n Let μ represent the nth input sample. nk Indicates sample x n For the k-th cluster center C k The membership degree is then used to cluster the samples by minimizing the following loss function J: ; Where N is the number of samples; K is the number of clusters; α is a constant, set to α=2; ||∗|| represents the distance metric for data similarity, here using Euclidean distance; membership degree μ nk This represents the degree to which the nth sample belongs to the kth class, and it is calculated using the following formula: ; Where C j Let C represent the j-th cluster center, j=1,2,…,K; and the corresponding k-th cluster center C. k Calculated by the following formula: ; Therefore, after the sample is decomposed by the task decomposition module, multiple clusters and cluster centers are obtained. Step 3.3: In the task allocation layer, the membership values ​​of each sample data to different clusters obtained by the fuzzy mean clustering method used by the task decomposition module are first regarded as the initial weight coefficients for the sample to belong to each cluster; then, an allocation threshold v is set in the task allocation layer. p Its value ranges from [0,1], representing the lowest membership degree of a sample belonging to each cluster. Here, v is set to... p =0.4; If the weight coefficient of a sample is greater than this threshold, the sample is assigned to the corresponding subtask; At the same time, the task allocation layer will count the number of times the sample is assigned to each task and record the corresponding weight coefficient, which is convenient for the output layer to perform weighted integration. Step 3.4: Use a randomly configured network as a sub-network of the modular neural network to construct corresponding sub-models for each sub-task obtained after task decomposition; Assume the task allocation module assigns the acquired samples {X,Y} to the corresponding subtasks t, t=1,2,…K, where K is the number of subtasks and the number of clusters. The data for each task is represented as {X}. t , Y t };in, Let be the input data of the network in the t-th subtask, and let be the corresponding output data. Nt represents the number of samples assigned to task t; assuming that a randomly configured network with L-1 hidden nodes has been constructed for the t-th subtask, given the objective function... The network output is: ; Among them, X t This represents the input data corresponding to the t-th subtask. The output weight of the j-th hidden node corresponding to the subnetwork is given by h; M is the number of output nodes; h is the output weight of the j-th hidden node corresponding to the subnetwork. j Let be the output vector of the j-th hidden node, which is represented as: ; Let be the activation function for the hidden layer neurons. In a randomly configured network, the sigmoid function is selected as the activation function for the hidden layer neurons, i.e.: ; Where <ꞏ> represents the inner product of Euclidean space, w j b j Let be the input weights and threshold of the j-th hidden node, respectively; the residual vector of the current network is defined as: ; The corresponding output residual value is ; M is the number of output nodes; if the current network's output residual The preset error tolerance requirement was not met, i.e. , Then the network will select new hidden layer neuron nodes according to the inequality constraint (14), and obtain the parameters of the Lth hidden layer node. ( , The inequality constraint for selecting hidden layer node parameters in the network is as follows: ; In the formula: 0<ε<1 and 0<q<1 are error reduction factors, where ε=0.015 and q=0.8; represents the output of the L-th hidden node, which is calculated by formula (10) and formula (11), is the output residual of the m-th output layer node when the network constructs L-1 hidden layer nodes, m=1,2,…,M, then for the t-th subtask, the output weight of the network hidden layer node after constructing the L-th hidden node is: ; Then, the network output after the Lth hidden layer node is established is: ; The network output residual at this point is: ; Then, the output error value of the network is calculated according to formula (13). And determine whether the network's output error meets the preset error requirements, i.e. If the condition is met, the SCNs construction is complete; otherwise, continue to add new hidden layer node parameters g(w, b) according to the inequality constraint (14) to reduce the network output error. At this time, the number of hidden nodes in the subnetwork is L=L+1, until the termination condition is met. , or , ; Step 3.5: The output integration layer integrates the outputs of each sub-network to obtain the final overall network output. In modular stochastic configuration networks, the output integration strategy is closely related to the characteristics of task decomposition. Therefore, a weighted average method is used in this output strategy to integrate the results of each sub-network. Assume that for the nth input sample x... n The probability of being assigned to each subtask is ,like Then set it to Where t=1,2,…K is the number of subtasks, v p The threshold for each task is set, with a value ranging from [0,1], representing the minimum membership value of a sample belonging to each cluster. Here, v is set to... p =0.4; then for The following processing is performed to obtain the final output weights for the t-th task corresponding to this sample. : ; Therefore, for the nth input sample, the corresponding network's final output y n for: ; in, Indicates sample x n The output on task t is calculated using formula (16); the unselected subnetworks do not contribute to the overall output of the network; therefore, the final output of the network corresponding to the nth sample is calculated using formula (19), which is the weighted sum of the samples assigned to the corresponding subnetworks. Step 4: Based on Step 3, the constructed modular random configuration network is obtained, and the ammonia nitrogen concentration in the effluent of the urban sewage treatment process is measured based on the constructed model to obtain the measurement results of the effluent ammonia nitrogen concentration.