Soft measurement and intelligent control method for sewage treatment process

By using improved isolated forest and RPD-XGBoost algorithms to filter variables, and combining HDformer and IHER-SAC models, high-precision soft measurement and intelligent control of the wastewater treatment process were achieved. This solved the problems of poor real-time performance and low efficiency in existing technologies, and improved the stability and economy of wastewater treatment.

CN120804924APending Publication Date: 2025-10-17HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510816386.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing wastewater treatment technologies suffer from poor real-time performance, low efficiency, and high cost, making it difficult to meet the requirements of modern wastewater treatment for high precision and high efficiency. Furthermore, in developing countries, inadequate facilities and a shortage of professional technicians lead to the direct discharge of wastewater without effective treatment, resulting in serious environmental pollution.

Method used

By using an improved isolated forest detection data outlier and combining it with the RPD-XGBoost algorithm to screen auxiliary variables, an HDformer soft measurement model was established. Based on the IHER-SAC deep reinforcement learning wastewater control model, a high-precision soft measurement and intelligent control of dissolved oxygen and nitrate nitrogen was achieved by adjusting the internal circulation flow and oxygen transfer coefficient through an intelligent agent.

Benefits of technology

It improves the accuracy and stability of wastewater treatment, reduces operating costs, can dynamically respond to water quality fluctuations, and ensures the stable and efficient operation of wastewater treatment plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804924A_ABST
    Figure CN120804924A_ABST
Patent Text Reader

Abstract

The invention discloses a soft measurement and intelligent control method for a sewage treatment process, and the method comprises the steps: collecting water quality parameter data in the sewage treatment process, improving an isolated forest through an attention mechanism, and detecting a data abnormal value; an RPD-XGBoost algorithm is used for screening auxiliary variables influencing the concentration of dissolved oxygen and nitrate nitrogen to serve as input data; establishing a soft measurement model of dissolved oxygen and nitrate nitrogen based on HDform, and outputting the concentrations of dissolved oxygen and nitrate nitrogen; a deep reinforcement learning sewage control model based on IHER-SAC is established, and the internal circulation flow and the oxygen transfer coefficient are controlled to adjust the concentration of nitrate nitrogen and dissolved oxygen in the sewage treatment process; the control strategy can be adjusted in real time according to the change of the sewage treatment system, higher sewage treatment efficiency is realized, and the operation cost of a sewage plant is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a soft measurement and intelligent control method, in particular to a soft measurement and intelligent control method for a sewage treatment process. BACKGROUND

[0002] With the aggravation of global water resource shortage and the increasing severity of water pollution, the demand for sewage treatment has risen sharply. According to the United Nations Environment Programme, global sewage production is expected to increase by 30% by 2030 compared with the current situation, and the sewage treatment industry is under great pressure. Traditional sewage treatment methods mainly rely on manual experience and laboratory analysis, which have the disadvantages of poor real-time performance, low efficiency and high cost, and are difficult to meet the requirements of modern sewage treatment for high precision and high efficiency. In developing countries, the problem of imperfect sewage treatment facilities and lack of professional technical personnel is particularly prominent, resulting in a large amount of sewage being directly discharged without effective treatment, causing serious pollution to the environment.

[0003] Existing sewage treatment technologies mostly use basic sensors to monitor water quality parameters such as chemical oxygen demand (COD) and biochemical oxygen demand (BOD) sensors, and each sensor operates independently, with data lacking effective integration. In terms of process control, PLC systems are often used to execute fixed control logic according to preset programs, making it difficult to flexibly adjust the treatment process according to real-time water quality changes. Existing sewage treatment technologies have problems such as difficulty in coordinating multiple parameters, rigid control strategy, insufficient real-time response, and poor stability, which are manifested in the following aspects: monitoring data cannot form a complete characterization system, control precision is low, adaptability to water quality fluctuations is poor, and system failure rate is high. These defects result in obvious shortcomings in sewage treatment efficiency, quality and sustainability of existing systems. SUMMARY

[0004] The purpose of the present application is to provide a soft measurement and intelligent control method for a sewage treatment process to improve the precision, stability and long-term operation efficiency of sewage treatment, overcome the monitoring and control problems in existing sewage treatment processes, and achieve efficient, intelligent and sustainable development of sewage treatment.

[0005] Technical solution: The method according to the present application comprises:

[0006] S1, collecting water quality parameter data in the sewage treatment process, detecting data outliers through an improved isolated forest, and performing data cleaning, wherein the improved isolated forest is improved by an attention mechanism;

[0007] S2, using an RPD-XGBoost algorithm to screen auxiliary variables affecting the concentrations of dissolved oxygen and nitrate nitrogen as input data;

[0008] S3, a soft measurement model of dissolved oxygen and nitrate nitrogen is established based on HDformer, the auxiliary variables screened in S2 are used as input parameters, and the concentrations of dissolved oxygen and nitrate nitrogen are output;

[0009] S4, a deep reinforcement learning sewage control model based on IHER-SAC is established, the soft measurement results obtained in S3 are used as the state space, the water quality of effluent and the total energy consumption are reduced as optimization objectives, and the target function is calculated as a reward;

[0010] S5, the model deep reinforcement learning sewage control model is used to control the internal circulation flow and the oxygen transfer coefficient to adjust the concentrations of nitrate nitrogen and dissolved oxygen in the sewage treatment process.

[0011] Preferably, the S1 comprises:

[0012] S11, water quality parameters in the sewage treatment process are collected;

[0013] S12, feature weights are dynamically allocated through an attention mechanism to enhance the influence of water quality parameters on anomaly detection;

[0014] S13, multiple isolated trees are constructed based on weighted features, and data space is recursively segmented to construct an isolated forest improved by an attention mechanism;

[0015] S14, the weighted path length is calculated, the path length of the sample in each tree is counted, and the anomaly sensitivity is adjusted in combination with the attention weight;

[0016] S15, the anomaly score is calculated according to the path length, and the threshold is dynamically adjusted;

[0017] S16, missing values and abnormal values are processed respectively, and data trends are retained.

[0018] Preferably, the S12 formula is as follows:

[0019]

[0020] wherein W represents a learnable weight matrix, b represents a bias term, α i represents feature attention weight, ⊙ represents element-wise multiplication, represents a weighted feature vector, x i represents a sample;

[0021] The formula of the isolated forest improved by the attention mechanism in S13 is as follows:

[0022]

[0023] wherein, represents a weighted feature, f represents a randomly selected feature of each tree, and p represents a segmentation point.

[0024] Preferably, S2 comprises:

[0025] S21, define input auxiliary variables and output target variables, divide the training set and the test set;

[0026] S22, train two XGBoost regression models respectively to predict the concentrations of dissolved oxygen and nitrate nitrogen, model dissolved oxygen and nitrate nitrogen respectively, and define the objective function;

[0027] S23, generate decision trees by gradient boosting algorithm iteration to optimize the objective function;

[0028] S24, calculate the initial gain, calculate the total gain of the features for the dissolved oxygen and nitrate nitrogen models respectively, introduce RPD evaluation, calculate the relative prediction deviation of the predicted values on the test set, take RPD as the weight to adjust the feature gain, and correct the gain;

[0029] S25, integrate the feature importance of the two XGBoost regression models to screen variables that significantly affect dissolved oxygen and nitrate nitrogen.

[0030] Preferably, the objective function formula of S22 is as follows:

[0031]

[0032] Wherein, θ represents a set of model parameters, represents a loss function, y i represents the true value, represents the predicted value, γ represents the regularization coefficient of the number of leaf nodes, T k represents the number of leaf nodes of the kth tree, λ represents the L2 regularization coefficient of the leaf weight, w k represents the leaf node weight vector of the kth tree, and n represents the number of samples;

[0033] The RPD evaluation formula of S24 is as follows:

[0034]

[0035] Wherein, SD represents the standard deviation, and RMSE represents the root mean square error;

[0036] The gain correction formula is as follows:

[0037]

[0038] Wherein, RPD DO and represents a proportion factor, and represents the adjusted gain.

[0039] The formula of the S25 screening condition is as follows:

[0040] Reserved

[0041] Wherein, τ represents a gain threshold, and α represents a weight coefficient, represents a comprehensive gain score of feature j.

[0042] Preferably, the S3 comprises:

[0043] S31, the auxiliary variables screened out by the RPD-XGBoost algorithm which have greater influence on dissolved oxygen and nitrate nitrogen are taken as inputs of the HDformer soft measurement model;

[0044] S32, the HDformer model is constructed, and a multi-head attention mechanism is used to capture the nonlinear relationship between the input variables;

[0045] S33, a hierarchical Transformer structure is used to enhance local and global feature fusion;

[0046] S34, residual connection is used to aggregate multi-layer features, and the concentrations of dissolved oxygen and nitrate nitrogen are outputted, and the aggregated features are mapped to an output space to obtain soft measurement values of dissolved oxygen and nitrate nitrogen.

[0047] Preferably, the formula of the S32 is as follows:

[0048]

[0049] Wherein, Q, K, and V represent query, key, and value matrices, which are obtained by input linear transformation, d k represents the dimension of the key vector;

[0050] The formula of the S33 is as follows:

[0051] Z l =LayerNorm(X l-1 +FFN(Attention(X l-1 )));

[0052] Wherein, Z l represents the output of the lth layer, FFN represents a feedforward network, and LayerNorm represents layer normalization.

[0053] Preferably, the S4 comprises:

[0054] S41, a state space is defined, the state space is composed of the concentrations of dissolved oxygen and nitrate nitrogen outputted by the soft measurement model, and a deep reinforcement learning wastewater control model based on IHER-SAC is constructed;

[0055] S42, define action space, agent controls process by adjusting internal circulation flow and oxygen transfer coefficient;

[0056] S43, design target function, integrate out-of-specification rate and energy consumption, introduce IHER mechanism to optimize sparse reward;

[0057] S44, IHER-SAC optimizes policy network and Q network to maximize cumulative reward and policy entropy.

[0058] Preferably, the S41 formula is as follows:

[0059]

[0060] Where s t represents the state vector at time t, including the current dissolved oxygen concentration DO t and nitrate nitrogen concentration

[0061] The S42 formula is as follows:

[0062] a t = [Q r,t , K La,t ];

[0063] Where a t represents the action vector, Q r,t represents the internal circulation flow, and K La,t represents the oxygen transfer coefficient.

[0064] The S43 formula is as follows:

[0065]

[0066] Where DO target and represent the set target concentration, E t represents the total energy consumption, α, β, γ represent the weight coefficients, I(s t ∈G) represents the virtual target reward introduced by IHER, and r t represents the reward.

[0067] The S44 formula is as follows:

[0068]

[0069] Where D is the experience replay buffer, containing real experience and virtual experience generated by IHER, α is the entropy regularization coefficient, represents the policy network parameter.

[0070] Preferably, the S5 includes:

[0071] S51, the policy network outputs an action, and adjusts the system parameters;

[0072] S52, update the policy network parameters according to the reward, and enhance the training stability by using the IHER mechanism.

[0073] Advantages: Compared with the prior art, the present application has the following remarkable advantages: 1. By constructing an IHER-SAC-based deep reinforcement learning sewage control model, the soft measurement results are included in the state space, and the effluent water quality is improved and the energy consumption is reduced as the optimization target, the agent dynamically adjusts the internal circulation flow and the oxygen transfer coefficient, and realizes intelligent, efficient and energy-saving control of the sewage treatment process; 2. By collecting multi-parameter water quality data and using an isolated forest improved by an attention mechanism for anomaly detection and cleaning, the data quality is ensured from the source, and accurate data support is provided for subsequent measurement and control; 3. By using the RPD-XGBoost algorithm to screen key auxiliary variables, the factors that significantly affect the dissolved oxygen and nitrate nitrogen are accurately located, and the dimensionality is effectively reduced while the relevance and representativeness of the model input are strengthened; 4. By constructing a soft measurement model using HDformer, its unique deep learning architecture can deeply mine the spatiotemporal characteristics of data, realize high-precision soft measurement of dissolved oxygen and nitrate nitrogen concentration, and improve the timeliness and accuracy of water quality state perception; 5. The whole method can dynamically adjust the control strategy according to the real-time changes of the sewage treatment system, effectively cope with water quality fluctuations and working condition changes, ensure the stable and economic operation of the sewage treatment plant, significantly improve the sewage treatment efficiency, and reduce the operating cost. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 It is a schematic diagram of the overall structure of the present application;

[0075] Figure 2 It is a schematic diagram of data processing by the improved isolated forest of the present application;

[0076] Figure 3 It is a schematic diagram of the soft measurement process of dissolved oxygen and nitrate nitrogen concentration of the present application. DETAILED DESCRIPTION

[0077] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings.

[0078] The sewage treatment process soft measurement and intelligent control method comprises the following steps:

[0079] S1, collect water quality parameter data in the sewage treatment process, detect data outliers by an improved isolated forest, and clean the data, the improved isolated forest is an isolated forest improved by an attention mechanism.

[0080] The specific steps are as follows:

[0081] S11. Use sensors or online instruments to collect key water quality parameters in real time during the sewage treatment process to ensure that the data covers the entire process cycle. The formula is as follows:

[0082] X={x1,x2,…,x n};

[0083]

[0084] Among them, X represents the data set, each sample x i Including pH, temperature (T), dissolved oxygen (DO), conductivity (EC), ammonia nitrogen Nitrate nitrogen and other parameters to ensure coverage of key water quality indicators.

[0085] S12. Dynamically assign feature weights through the attention mechanism to enhance the impact of key parameters (such as ammonia nitrogen and dissolved oxygen) on anomaly detection. The formula is as follows:

[0086]

[0087] Among them, W represents the learnable weight matrix, b represents the bias term, and a i represents the feature attention weight, ⊙ represents element-wise multiplication, represents the weighted eigenvector.

[0088] S13. Constructing attention isolation forest based on weighted features Construct multiple isolated trees and recursively partition the data space. The formula is as follows:

[0089]

[0090] in, Represents weighted features, f represents randomly selected features for each tree, p represents the split point, and recursive splitting is performed until the sample is isolated or the depth limit l is reached. maxx .

[0091] S14. Calculate the weighted path length and count the path length of the sample in each tree Combined with the attention weight to adjust the abnormal sensitivity, the formula is as follows:

[0092]

[0093] Among them, h k Indicates the path length of the k-th tree. The shorter the path ( The smaller the , the more likely the sample is abnormal, and the attention weight makes it easier to capture abnormalities in key parameters.

[0094] S15, dynamic threshold anomaly judgment, calculate the anomaly score based on the path length, and dynamically adjust the threshold. The formula is as follows:

[0095]

[0096] where H(n) represents the harmonic number, and the threshold value θ can be adjusted adaptively according to the data distribution.

[0097] S16, missing values and outliers are processed respectively, and the data trend is preserved, and the formula is as follows:

[0098]

[0099] where the first formula is for missing value filling, using linear interpolation; the second formula is for outlier correction, using sliding window mean, and the window size w is time correlation (such as 6 hours), and only normal values (s<θ) are used to calculate the mean.

[0100] S2, use RPD-XGBoost algorithm to screen auxiliary variables affecting dissolved oxygen and nitrate nitrogen concentration as input data.

[0101] The specific steps are as follows:

[0102] S21, define input auxiliary variables (such as pH, temperature, flow, etc.) and output target variables (dissolved oxygen DO, nitrate nitrogen , divide the training set and test set, and the formula is as follows:

[0103] X={x1,x2,…,x m};

[0104] Y DO ={y DO,1 ,y DO,2 ,…,y DO,n};

[0105]

[0106] where X represents the auxiliary variable set, containing m samples, each sample has multiple features (such as pH, temperature, conductivity, internal circulation flow, etc.); Y DO represents the actual measured value set of dissolved oxygen (DO), containing n samples corresponding to the DO concentration;

[0107] represents the actual measured value set of nitrate nitrogen , containing n samples corresponding to the nitrate nitrogen concentration.

[0108] S22, RPD-XGBoost model training, two XGBoost regression models are trained respectively to predict dissolved oxygen (DO) and nitrate nitrogen The concentration, dissolved oxygen and nitrate nitrogen were modeled independently to avoid the interference of feature importance; the objective function was defined to optimize the model by minimizing the prediction error and regularization term, as follows:

[0109]

[0110] where θ represents the set of model parameters, represents the loss function, y i represents the true value, represents the predicted value, γ represents the regularization coefficient of the number of leaf nodes, used to punish the complex tree structure, T k represents the number of leaf nodes of the kth tree, λ represents the L2 regularization coefficient of the leaf weight, used to prevent overfitting, w k represents the leaf weight vector of the kth tree, and n represents the number of samples.

[0111] S23, generate decision trees by gradient boosting algorithm iteration to optimize the objective function, as follows:

[0112]

[0113] where, is the predicted value of sample i after the tth iteration, represents the predicted value of the (t-1)th iteration; η represents the learning rate (step size), which controls the contribution weight of each tree (usually set to 0.01-0.3); f t (x i ) represents the prediction function of the tth tree, which is generated by selecting the optimal feature split point by the greedy algorithm.

[0114] S24, improve feature importance evaluation by RPD. First, calculate the initial gain, calculate the total gain of feature x j for dissolved oxygen and nitrate nitrogen models, as follows:

[0115]

[0116] where, represents the set of nodes where feature x j is used for splitting, ΔL k is the loss reduction after node k splitting.

[0117] S25, introduce RPD evaluation, calculate the relative prediction deviation of the predicted value on the test set, measure the stability of the model, as follows:

[0118]

[0119] where SD represents the standard deviation, RMSE represents the root mean square error, and RPD>2.5 indicates that the model has excellent prediction ability.

[0120] S26、Gain correction, adjust the feature gain as the weight of RPD, enhance the robustness, the formula is as follows:

[0121]

[0122] Where RPD DO and is the proportion factor, and is the adjusted gain.

[0123] S27, auxiliary variable screening, comprehensive feature importance of two models, screening variables that significantly affect both dissolved oxygen and nitrate nitrogen, screening condition formula is as follows:

[0124] Reserved

[0125] Where τ represents the gain threshold, only features that exceed the threshold in both models are retained (such as τ = 0.7); α represents the weight coefficient, used to balance the importance of dissolved oxygen and nitrate nitrogen;

[0126] is the comprehensive gain score of feature j, the top k features with the highest score are selected as input; variables that play a key role in measuring both dissolved oxygen and nitrate nitrogen are retained to improve the robustness of subsequent soft measurement models.

[0127] S3, based on HDformer to establish the soft measurement model of dissolved oxygen and nitrate nitrogen, using the auxiliary variables screened in S2 as input parameters, output the concentration of dissolved oxygen and nitrate nitrogen.

[0128] The specific steps are as follows:

[0129] S31, the auxiliary variables screened by RPD-XGBoost algorithm that have greater impact on dissolved oxygen and nitrate nitrogen are used as input of HDformer soft measurement model, the formula is as follows:

[0130] X = [x1, x2, …, x n ] T ;

[0131] Where x i is the i-th auxiliary variable (such as pH, temperature, DO, ammonia nitrogen, etc.), there are n samples.

[0132] S32, build HDformer model, capture the nonlinear relationship between input variables through multi-head attention mechanism, the formula is as follows:

[0133]

[0134] where Q, K, V represent Query, Key, Value matrix, which are linearly transformed from input X, d k denotes the dimension of key vector, which is used to scale dot-product attention.

[0135] S33, Hierarchical Transformer structure (HDformer) is adopted to enhance the fusion of local and global features, as follows:

[0136] Z l = LayerNorm(X l-1 + FFN(Attention(X l-1 ))).

[0137] where Z l denotes the output of the l-th layer, FFN denotes Feed-Forward Network, which contains two fully connected layers and ReLU activation, and LayerNorm is layer normalization to stabilize the training process.

[0138] S34, multi-layer features are aggregated through residual connection, as follows:

[0139]

[0140] where a l denotes the learnable weight coefficient, which balances the importance of different layers, and L denotes the total number of layers.

[0141] The update formula of a l can be expressed as:

[0142]

[0143] where η denotes the learning rate, which controls the update step size; denotes the loss function, which represents the difference between the model predicted value and the true value, and each weight coefficient a l is updated according to the gradient of the loss function on it, and the gradient reflects the influence of the current weight coefficient on the prediction error, and the learning rate η controls the update amplitude, in this way, a l is constantly adjusted to achieve the purpose of optimizing the model performance.

[0144] S35, the output dissolved oxygen and nitrate nitrogen concentration, the aggregated features are mapped to the output space to obtain the soft measurement values of dissolved oxygen and nitrate nitrogen, as follows:

[0145]

[0146] where W odenotes the output layer weight matrix, b o denotes the bias term.

[0147] and are the estimated values of dissolved oxygen and nitrate nitrogen concentration of the model output.

[0148] S4, a deep reinforcement learning sewage control model based on IHER-SAC is established, the soft measurement results obtained in S3 are used as the state space, the effluent water quality is improved and the total energy consumption is reduced as the optimization goal, and the target function is calculated as the reward.

[0149] The specific steps are as follows:

[0150] S41, a sewage control model based on IHER-SAC is constructed, the state space is defined, and the state space is composed of the dissolved oxygen DO and nitrate nitrogen concentration output by the soft measurement model, and the formula is as follows:

[0151]

[0152] Where, s t denotes the state vector at time t, including the current dissolved oxygen concentration DO t and nitrate nitrogen concentration

[0153] S42, define the action space, the agent controls the process by adjusting the internal circulation flow Q r and oxygen transfer coefficient K La The formula is as follows:

[0154] a t =[Q r,t ,K La,t ];

[0155] Where, a t denotes the action vector, Q r,t denotes the internal circulation flow (m 3 / h), and K La,t denotes the oxygen transfer coefficient (h -1 ).

[0156] S43, design the target function (reward function), the reward function integrates the effluent quality compliance rate and energy consumption, and introduces the IHER mechanism to optimize the sparse reward, and the formula is as follows:

[0157]

[0158] Where, DO target and denote the set target concentration, E t denotes the total energy consumption (Q r,t and K​La,t positive correlation), a, b, g represent weight coefficients, I(s t G) represents the virtual target reward introduced by IHER (if the state s t approaches the target set G, an additional reward is given), which is an indicator function (also called an identity function), when the state s t belongs to the target set G, the value of I(s t G) is 1, at this time an additional reward g (g is a weight coefficient, which determines the size of this additional reward) is given, when the state s t does not belong to the target set G, the value of I(s t G) is 0, and this part of the additional reward is not given.

[0159] Sparse reward refers to the fact that the agent only receives rewards in a small number of specific events or states during interaction, and the reward signal does not appear frequently. Here, the virtual target reward I(s t G) introduced by IHER is the key to optimizing sparse rewards, when the state s t approaches the target set G, an additional reward g is given, and this additional reward is set to solve the problem of possible reward sparsity, to encourage the agent to develop towards the state of the target set, so that the agent can get timely feedback when the state approaches the target, so as to better optimize the behavior policy to achieve the design goal of comprehensive water quality compliance rate and energy consumption.

[0160] S44, IHER-SAC optimization strategy network and Q network Maximize cumulative reward and policy entropy, formula as follows:

[0161]

[0162] Where D is the experience replay buffer, containing real experience and virtual experience generated by IHER, a is the entropy regularization coefficient, which controls the exploration intensity.

[0163] S5, use the model to control the depth reinforcement learning wastewater control model to control the internal circulation flow and the oxygen transfer coefficient to adjust the nitrate nitrogen and dissolved oxygen concentration in the wastewater treatment process.

[0164] The specific steps are as follows:

[0165] S51, action generation and execution, policy network outputs action a t = [Q r,t , K La,t ], adjusts system parameters, formula as follows:

[0166]

[0167] wherein, and are the policy sub-networks of inner circulation flow and oxygen transfer coefficient, respectively, Q r,t is the inner circulation flow (m 3 / h), K La,t is the oxygen transfer coefficient (h -1 ).

[0168] S52, dynamically optimizing, updating the policy network parameters t according to the reward r and enhancing the training stability by using the IHER mechanism, as follows:

[0169]

[0170] wherein, η is the learning rate, and the IHER mechanism: in the training process, the agent not only learns the real experience, but also learns the optimal policy in the virtual target state, improving the sample utilization rate.

Claims

1. A soft measurement and intelligent control method for sewage treatment process, characterized in that: The following steps are involved: S1. Collect water quality parameter data from the sewage treatment process, detect data outliers using a modified isolation forest algorithm that uses an attention mechanism to improve the isolation forest, and perform data cleaning. S2. Use the RPD-XGBoost algorithm to screen auxiliary variables that affect dissolved oxygen and nitrate nitrogen concentrations as input data; S3, based on HDformer, a soft sensing model of dissolved oxygen and nitrate nitrogen is established, using the auxiliary variables selected in S2 as input parameters to output the concentrations of dissolved oxygen and nitrate nitrogen; S4. Establish a deep reinforcement learning sewage control model based on IHER-SAC, use the soft measurement results obtained in S3 as the state space, improve the effluent water quality and reduce the total energy consumption as the optimization goals, and calculate the objective function as the reward; S5. Use the model to perform deep reinforcement learning on the sewage control model to control the internal circulation flow and the oxygen transfer coefficient to adjust the nitrate nitrogen and dissolved oxygen concentrations during the sewage treatment process.

2. The method according to claim 1, characterized in that Said S1 comprises: S11. Collect water quality parameters during sewage treatment; S12, dynamically assign feature weights through the attention mechanism to enhance the influence of water quality parameters on anomaly detection; S13. Build multiple isolation trees based on weighted features, recursively partition the data space, and construct an isolation forest improved by the attention mechanism. S14. Calculate the weighted path length, count the path length of the sample in each tree, and adjust the abnormal sensitivity in combination with the attention weight; S15. Calculate the anomaly score based on the path length and dynamically adjust the threshold; S16. Process missing values ​​and outliers separately and retain data trends.

3. The method according to claim 2, characterized in that The S12 formula is as follows: Among them, W represents the learnable weight matrix, b represents the bias term, α i represents the feature attention weight, ⊙ represents element-wise multiplication, represents the weighted eigenvector, x i represents a sample; The isolation forest formula improved by the attention mechanism described in S13 is as follows: in, represents weighted features, f represents the random selection of features for each tree, and p represents the split point.

4. The method according to claim 1, wherein The S2 includes: S21. Define input auxiliary variables and output target variables, and divide the training set and test set; S22. Train two XGBoost regression models to predict dissolved oxygen and nitrate nitrogen concentrations, model dissolved oxygen and nitrate nitrogen separately, and define the objective function. S23, iteratively generate a decision tree through the gradient boosting algorithm to optimize the objective function; S24. Calculate the initial gain, calculate the total gain of the features for the dissolved oxygen and nitrate nitrogen models respectively, introduce RPD evaluation, calculate the relative prediction deviation of the predicted value on the test set, use RPD as a weight to adjust the feature gain, and correct the gain; S25. Integrate the feature importance of the two XGBoost regression models to screen variables that have significant effects on both dissolved oxygen and nitrate nitrogen.

5. The method according to claim 4, characterized in that The objective function formula of S22 is as follows: Among them, θ represents the set of model parameters, represents the loss function, y i represents the true value, represents the predicted value, γ represents the regularization coefficient of the number of leaf nodes, T k represents the number of leaf nodes in the k-th tree, λ represents the L2 regularization coefficient of the leaf weight, k represents the leaf node weight vector of the k-th tree, and n represents the number of samples; The RPD evaluation formula described in S24 is as follows: Where SD represents standard deviation and RMSE represents root mean square error; The gain correction formula is as follows: Among them, RPD DO and represents the scale factor, and Indicates the adjusted gain. The formula of the S25 screening condition is as follows: reserve Among them, τ represents the gain threshold, α represents the weight coefficient, represents the comprehensive gain score of feature j.

6. The method according to claim 1, characterized in that The S3 includes: S31. The auxiliary variables with greater impact on dissolved oxygen and nitrate nitrogen selected by the RPD-XGBoost algorithm are used as the input of the HDformer soft sensing model; S32. Build the HDformer model to capture the nonlinear relationship between input variables through the multi-head attention mechanism; S33, using a hierarchical Transformer structure to enhance local and global feature fusion; S34. Aggregate multiple layers of features through residual connections, output dissolved oxygen and nitrate nitrogen concentrations, map the aggregated features to the output space, and obtain soft measurement values ​​of dissolved oxygen and nitrate nitrogen.

7. The method according to claim 6, characterized in that The S32 formula is as follows: Among them, Q, K, V represent query, key, and value matrices, which are obtained by linear transformation of input, d k represents the dimension of the key vector; The S33 formula is as follows: Z l =LayerNorm(X l-1 +FFN(Attention(X l-1 ))); Among them, Z l Represents the output of the lth layer, FFN represents the feedforward network, and LayerNorm is layer normalization.

8. The method according to claim 1, characterized in that The S4 includes: S41. Define the state space, which is composed of the dissolved oxygen and nitrate nitrogen concentrations output by the soft sensor model, and build a deep reinforcement learning sewage control model based on IHER-SAC; S42, defining the action space, the agent controls the process by adjusting the internal circulation flow and oxygen transfer coefficient; S43. Design the objective function, comprehensively consider the water quality compliance rate and energy consumption, and introduce the IHER mechanism to optimize the sparse reward; S44 and IHER-SAC optimize the policy network and Q network to maximize the cumulative reward and policy entropy.

9. The method according to claim 8, characterized in that The S41 formula is as follows: Among them, s t Represents the state vector at time t, including the current dissolved oxygen concentration DO t and nitrate nitrogen concentration The S42 formula is as follows: a t =[Q r,t ,K La,t ]; Among them, a t represents the action vector, Q r,t Indicates the internal circulation flow, K La,t represents the oxygen transfer coefficient; The S43 formula is as follows: Among them, DO target and Indicates the set target concentration, E t represents the total energy consumption, α, β, γ represent the weight coefficients, I(s t ∈G) represents the virtual target reward introduced by IHER, r t Indicates reward. The S44 formula is as follows: Where D is the experience replay buffer, which contains real experience and virtual experience generated by IHER, α is the entropy regularization coefficient, Represents policy network parameters.

10. The method according to claim 1, characterized in that The S5 includes: S51, the policy network outputs actions and adjusts system parameters; S52. Update the policy network parameters based on the reward and use the IHER mechanism to enhance training stability.