Bypass ash water washing dechlorination system
By building a bypass grey water washing and chlorine removal system, integrating physical dynamics modeling and reinforced learning control, the traditional grey water chlorine removal system cannot flexibly cope with changes in working conditions and resource waste, and achieve efficient, safe and intelligent chlorine concentration control, improving dechlorination efficiency and resource utilization.
Patent Information
- Application Number
- CN202510402163.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing grey water chlorine removal system lacks bypass design and cannot flexibly respond to changes in working conditions. The control strategy relies on fixed parameters and cannot perceive dynamic state changes, resulting in low control accuracy, serious resource waste, and failure to effectively integrate physical reaction models and intelligent control models, making it difficult to meet the efficient, safe and intelligent industrial needs.
Build a bypass grey water washing and chlorine removal system, integrate physical dynamic modeling and reinforcement learning control, and achieve accurate perception and dynamic adjustment of working conditions through data acquisition, timing modeling, dynamic modeling and depth deterministic strategy gradient controller, and has the advantages of sensitive response, strong adaptability, stable control, and high resource utilization.
The system's ability to adapt to chlorine concentration fluctuations has been improved, safety and continuous operation capabilities have been enhanced, dechlorination efficiency has been improved, operating costs have been reduced, and the system has been achieved with high generalization ability and learning efficiency.
Smart Images

Figure CN120335296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial wastewater treatment and intelligent control, and particularly relates to a bypass ash water washing dechlorination system. Background Art
[0002] Chlorine is a common industrial waste gas pollutant, which widely exists in industrial processes such as chemical industry, electric power, metallurgy, and papermaking. The effective treatment of chlorine-containing waste gas has become a key topic in the field of waste gas treatment. As a typical alkaline absorption method, ash water dechlorination has the advantages of low operating cost, flexible operation mode, and easy integration with the system, and is widely used to neutralize chlorine-containing tail gas. However, in actual engineering applications, traditional ash water dechlorination systems generally face multiple technical bottlenecks and are difficult to meet the modern treatment requirements of efficient, safe, intelligent, and low-consumption coordinated control.
[0003] Currently, the mainstream ash water dechlorination methods mainly rely on synchronous dosing or constant flow washing modes. Usually, the treated ash water is directly connected to the main reaction system for mixing and absorption. Although this method has a simple structure, its core defect is that it cannot achieve fine control of the ash water dosage and reaction conditions, which easily leads to over-dosing or insufficient reaction. On the one hand, the concentration of chlorine-containing media often has large fluctuations in actual working conditions. If the treatment system has a lag in response or limited adjustment ability, it is very likely to cause a decrease in dechlorination efficiency and even safety hazards such as chlorine leakage. On the other hand, if a conservative strategy is adopted to configure excessive ash water, it will cause waste of resources, secondary pollution, and a significant increase in operating costs. In addition, traditional ash water dechlorination systems have limited capabilities in terms of feedback response ability, parameter adaptive adjustment, and operating state perception, lack data-driven support, and are difficult to meet the urgent needs of modern industry for high-stability and high-intelligent control systems.
[0004] In existing technical solutions, some studies have tried to use fixed control strategies or rule bases to adjust the ash water dosing. For example, multiple fixed intervals are set according to the high and low chlorine concentrations, corresponding to different set values of ash water flow rates. Such adjustment methods based on thresholds or empirical models are effective under specific stable working conditions, but for the actual production process with dynamic change characteristics, their adaptability and robustness have significant deficiencies. In addition, due to the lack of kinetic modeling and real-time learning mechanisms for the reaction process between ash water and chlorine, the control strategy cannot effectively capture the coupling characteristics between multiple variables inside the system, and the control effect is easily affected by interference parameters and fluctuates.
[0005] In recent years, with the in-depth application of artificial intelligence technology in the field of industrial process modeling and control, some research has begun to explore the combination of machine learning and traditional control methods to achieve process state recognition, strategy optimization, and system adaptive control. For example, neural networks are used to predict reaction efficiency, or rule adjustment is achieved based on fuzzy control. However, most of these methods still rely on a large number of static training samples, are difficult to adapt to real-time changing working conditions, and most of them are not deeply coupled with the physical reaction mechanism, resulting in poor interpretability and weak stability, making it difficult to be implemented in complex industrial systems.
[0006] In addition, existing grey water treatment systems generally adopt the main process injection method, lacking a bypass control structure in operation and unable to achieve local optimization adjustment without interrupting the main process. Once abnormal fluctuations occur in the chlorine concentration of the main system or the operating parameters are unbalanced, the overall processing capacity of the system is easily affected, and the operation safety and continuity are difficult to guarantee. Therefore, there is an urgent need for a grey water dechlorination system with a bypass structure that can integrate system operation data, kinetic models, and intelligent control algorithms to achieve precise response and resource optimization allocation under chlorine concentration fluctuations.
[0007] In summary, the existing technologies have the following significant defects in grey water dechlorination control: First, the traditional system structure lacks a bypass design, the system adjustment ability is limited, and it is unable to flexibly respond to changes in the main process working conditions; Second, the dechlorination strategy depends on fixed parameters or empirical rules and cannot perceive the dynamic state changes of the system, resulting in low control accuracy and serious resource waste; Third, the physical reaction model and the learning-based intelligent control model are not effectively integrated, lacking system adaptive optimization ability and control interpretability, and it is difficult to meet the industrial requirements of green, intelligent, and stable operation. Therefore, it is necessary to propose a bypass grey water washing dechlorination system based on physical model perception and deep reinforcement learning to achieve efficient utilization of grey water resources, precise self-adjustment of control strategies, and closed-loop optimization of system operation, thereby improving dechlorination efficiency, reducing operation costs, and enhancing the intelligent response ability and industrial adaptability of the system. Summary of the Invention
[0008] An object of the present invention is to propose a bypass grey water washing dechlorination system. The present invention integrates physical kinetic modeling and reinforcement learning control methods to construct a bypass grey water treatment system to achieve intelligent washing and dechlorination control of chlorine-containing media. Through time-series modeling and coordinated optimization of control strategies, the system can accurately perceive changes in working conditions, dynamically adjust key parameters, and has the advantages of sensitive response, strong self-adaptability, stable control, and high resource utilization rate.
[0009] A bypass grey water washing dechlorination system according to an embodiment of the present invention includes:
[0010] A data acquisition module for acquiring operation data of the bypass grey water treatment system;
[0011] A data preprocessing module for normalizing the operation data to generate time series data that can be used for modeling;
[0012] A time series modeling module for extracting time series features of the system operation status based on the Transformer network;
[0013] A kinetic modeling module for constructing a kinetic model of the reaction between greywater and chlorine, calculating the chlorine removal rate, and predicting the residual chlorine concentration;
[0014] A control strategy module for fusing time series features and kinetic features, training with a deep deterministic policy gradient controller, and outputting a bypass greywater control strategy;
[0015] An execution control module for adjusting the operation of bypass greywater according to the control strategy to achieve the chlorine washing treatment of greywater;
[0016] A feedback optimization module for collecting the processing results, updating the kinetic model and control strategy, and achieving closed-loop optimal control.
[0017] Optionally, the modules are implemented by the following methods:
[0018] S1. Collect the operation data of the bypass greywater treatment system;
[0019] S2. Preprocess the operation data to generate standardized time series data;
[0020] S3. Input the standardized time series data into a time series modeling network with a Transformer encoder as the main structure to extract the time series features of the system operation status;
[0021] S4. Construct a kinetic model based on the chlorine-containing medium concentration, greywater pH value, reaction unit temperature, and reaction residence time in the operation data, calculate the chlorine removal rate per unit time, and predict the residual chlorine concentration;
[0022] S5. Jointly input the time series features extracted by the Transformer network and the chlorine removal features output by the kinetic model into a deep deterministic policy gradient controller to train the deep deterministic policy gradient controller including an Actor network and a Critic network;
[0023] S6. Output a greywater bypass control strategy by the deep deterministic policy gradient controller;
[0024] S7. Control the operation of the bypass greywater treatment system according to the greywater bypass control strategy to complete the chlorine washing treatment process of the chlorine-containing medium;
[0025] S8. Collect the operation result data of the reaction system after the washing treatment to generate feedback data;
[0026] S9. Update the parameters of the kinetic model according to the feedback data, and at the same time update the deep deterministic policy gradient controller.
[0027] Optionally, the operating data includes the concentration of chlorine-containing medium, the pH value of the ash water, the ash water flow rate, the reaction unit temperature, the reaction residence time, and the outlet chlorine concentration.
[0028] Optionally, the preprocessing includes missing value filling, outlier removal, normalization processing, and standardization processing.
[0029] Optionally, S3 includes the following specific steps:
[0030] S31. Input the standardized time series data X = {x i (t j )} into the time series modeling network with a Transformer encoder as the backbone structure, where x i (t j ) represents the standardized time series data value of the i-th variable at time step t j ;
[0031] S32. Perform a linear mapping on the standardized time series data, map it to an embedding space of a fixed dimension, and add a positional encoding vector to each time step to obtain a time embedding representation:
[0032] e(t j ) = W1x(t j ) + b1 + p(t j );
[0033] Among them, e(t j ) represents the time embedding representation, W1 represents the linear mapping weight matrix, x(t j ) represents the standardized time series data, b1 represents the bias vector, and p(t j ) represents the positional encoding vector, describing the time step feature vector after adding the positional encoding;
[0034] S33. Input the sequence composed of the time embedding representations of all time steps into the Transformer encoder module, and calculate the output of each attention head:
[0035]
[0036] Among them, head k (t j ) represents the output of the k-th attention head at time step t j , softmax represents normalization, Q k and K k and Vk represent the query, key, and value matrices respectively, T represents the transpose operation, and d k represents the feature dimension of a single attention head;
[0037] S34. Concatenate the outputs of all attention heads to obtain the multi-head attention representation, and extract high-dimensional features through a linear transformation and a feed-forward neural network:
[0038] h(t j ) = FFN(A(t j )) = max(0, A(t j ))W2 + b2)W3 + b3;
[0039] where h(t j ) represents the high-dimensional feature vector at time step t j , A(t j ) represents the multi-head attention representation, describing the concatenation result of the multi-head attention, FFN represents the feed-forward neural network, max represents taking the maximum value, max(0, ·) represents the non-linear activation function, W2 and W3 represent the weight matrices of the feed-forward network, and b2 and b3 represent the biases of the feed-forward network;
[0040] S35. Apply a residual connection and layer normalization operation to the output h(t j ) at each time step to form the final encoded feature vector, and use the final encoded feature vector as the high-dimensional expression of the system operating state.
[0041] Optionally, S4 includes the following specific steps:
[0042] S41. Obtain the initial concentration of the chlorine-containing medium, the pH value of the grey water, the temperature of the reaction unit, and the reaction residence time as the input variables of the reaction kinetics model;
[0043] S42. Establish a first-order reaction rate model between the grey water and the chloride, and define the differential relationship of the chlorine concentration changing with time:
[0044]
[0045] where C(t) represents the chlorine concentration at time t represents the differential of the chlorine concentration with respect to time t, and k eff represents the reaction rate constant corrected by pH;
[0046] S43. Calculate the reaction rate constant according to the Arrhenius formula based on the temperature of the reaction unit, and calculate the alkalinity adjustment factor based on the pH value to correct the reaction rate constant:
[0047]
[0048] k eff = k·α pH , α pH = log 10 (1 + 10 pH-7 );
[0049] where k represents the reaction rate constant, A represents the pre-exponential factor, E a represents the activation energy of the reaction, e represents the natural constant, R represents the gas constant, P represents the temperature of the reaction unit, k eff represents the reaction rate constant after pH correction, and α pH represents the pH enhancement factor, which is a dimensionless constant;
[0050] S44. Calculate the predicted value of the chlorine concentration during the reaction residence time under the condition of the initial concentration of the chlorine-containing medium:
[0051]
[0052] where C pred represents the predicted value of the chlorine concentration, τ represents the reaction residence time, and C0 represents the initial concentration of the chlorine-containing medium;
[0053] S45. Calculate the chlorine removal rate based on the predicted value C pred and the input parameters:
[0054]
[0055] where R rem represents the average chlorine removal rate per unit time;
[0056] S46. Use the predicted value C pred of the chlorine concentration and the chlorine removal rate R rem as the chlorine removal characteristics and input them into the subsequent control strategy training model for the state space of the reinforcement learning controller.
[0057] Optionally, the S5 includes the following specific steps:
[0058] S51. Extract the system state encoding feature vector at the current time step. The feature vector is output by the Transformer network and combined with the predicted value C pred of the chlorine concentration and the average chlorine removal rate R rem per unit time calculated by the kinetic model, and splice them to form a state vector;
[0059] S52. Construct a deep deterministic policy gradient controller. The deep deterministic policy gradient controller includes an Actor network and a Critic network. The Actor network is used to output a control action a t according to the current state s t, where the Critic network is used to estimate the value function Q(s t , a t ) of the state-action pair (s t , a t );
[0060] S53. Define the action vector a t = [v t , q t , f t , τ t , where v t is the opening and closing state of the bypass valve, q t is the rotational speed of the ash water pump, f t is the ash water injection flow rate, and τ t is the residence time of the reaction unit;
[0061] S54. In the policy training stage, based on the current Actor policy, output the action a t , input the current state and action into the system environment to obtain the immediate reward and the next state, and combine them to form a quadruple (s t , a t , r t , s t+1 ), store it in the experience replay pool, and randomly sample a batch of parameters for training the Actor network and the Critic network from the experience replay pool;
[0062] S55. Construct an objective value function containing the penalty for the deviation of the dynamic model, and calculate the target Q value for the state at the next time step:
[0063] y t = r t + γQ ′ (s t+1 , μ ′ (s t+1 )) - λ·|C pred - C real |;
[0064] where y t represents the target Q value, r t represents the immediate reward returned by the environment at time step t, γ represents the discount factor used to balance the current reward and the future value, s t+1 represents the state at the next time step, μ ′ (s t+1 ) represents the action output by the target policy network, and Q ′ (s t+1 , μ ′ (s t+1 )) represents the target Critic network's evaluation of the state s t+1The value estimation of the actions output by the target policy network, λ represents the weight of the dynamic model error penalty, and C pred represents the predicted value of the chlorine concentration, and C real represents the true outlet chlorine concentration collected by the sensor;
[0065] S56. Construct a loss function and update the current Critic network parameters using the Adam optimizer by minimizing the loss function:
[0066]
[0067] where L represents the mean square error loss function, N represents the sample batch size for each round of training, represents the state-action value estimation output by the current Critic network, represents the target Q value;
[0068] S57. Update the Actor network parameters using the policy gradient method to maximize the state-action value function, and the optimization objective function is:
[0069]
[0070] where J represents the policy objective function, represents the state of the target Critic network at time step t and the actions output by the target policy network of the value estimation, represents the target Q value;
[0071] S58. Repeat the processes of sampling, training, and parameter updating to achieve the training and convergence of the greywater bypass control strategy with the perception of dynamic model error.
[0072] Optionally, the greywater bypass control strategy includes the opening and closing states of the bypass valve, the rotational speed of the greywater pump, the greywater injection flow rate, and the residence time of the reaction unit.
[0073] Optionally, the operation result data includes the treated chlorine concentration, the dechlorination efficiency, and the control response change.
[0074] Optionally, S9 includes the following specific steps:
[0075] S91. Obtain feedback data, where the feedback data includes the treated chlorine concentration, the dechlorination efficiency, and the control response change. The treated chlorine concentration is the real-time measured concentration at the outlet of the reaction unit, and the dechlorination efficiency is calculated from the initial chlorine concentration and the outlet chlorine concentration;
[0076] S92. Compare the post-treatment chlorine concentration in the feedback data with the chlorine concentration predicted by the kinetic model, calculate the prediction error of the kinetic model, and use the prediction error to evaluate the performance deviation of the kinetic model;
[0077] S93. Fine-tune the parameters of the kinetic model based on the prediction error, set the set of parameters to be optimized, update the parameters of the kinetic model using a loss function based on the mean square error, and update the deep deterministic policy gradient controller simultaneously.
[0078] The beneficial effects of the present invention are as follows:
[0079] First of all, a reinforced learning bypass ash water chlorine removal system based on physical model perception provided by the present invention fully integrates kinetic mechanism modeling and data-driven control strategies, and has significant innovation and practical application value in both structural design and control methods. By setting up a bypass ash water treatment system, the addition and reaction of ash water operate independently outside the main process without interfering with the original process chain, realizing flexible adjustment and partition decoupling of the system, improving the adaptability to the fluctuating emission of chlorine-containing media, enhancing the safety and continuous operation ability of the system. The reaction kinetic model constructed by combining working condition parameters such as the temperature, pH value, and residence time of the reaction unit can dynamically predict the residual chlorine concentration and chlorine removal rate, making the control strategy not only rely on empirical data, but also have physical interpretability and prediction foresight.
[0080] Secondly, in terms of the control strategy, the present invention introduces a time series modeling network with Transformer as the backbone to accurately capture the time-dependent relationship and state evolution characteristics among multiple variables in the ash water treatment system, and jointly constitutes a state space with kinetic characteristics to guide the deep deterministic policy gradient controller to train the ash water bypass control strategy, effectively realizing the continuity, self-adaptability, and dynamic optimality of the action output. Compared with traditional rule-based or fuzzy logic control methods, this strategy has higher generalization ability and learning efficiency, and can timely adjust key parameters such as ash water flow rate, pump speed, and reaction time according to working condition fluctuations, improving the dechlorination efficiency while significantly reducing alkali consumption and ash water waste.
[0081] Finally, the present invention establishes a feedback optimization module, and jointly updates the kinetic model and the control strategy network based on the error-driven mechanism between the actual operation results and the model predictions, realizing the closed-loop adaptive co-optimization between the physical model perception and the reinforcement learning model, fundamentally improving the intelligent level and long-term stability of the control system. The system has the ability of continuous learning and dynamic optimization, can gradually approach the optimal control path during actual operation, taking into account dechlorination efficiency, response speed, and energy consumption balance, and provides an integrated, intelligent, and low-carbon new solution for the treatment of chlorine-containing industrial tail gas. Description of the Drawings
[0082] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:
[0083] Figure 1 It is a schematic diagram of the overall structure of a bypass ash water chlorine washing system proposed by the present invention;
[0084] Figure 2 It is a flowchart of a method for a bypass ash water chlorine washing system proposed by the present invention;
[0085] Figure 3 It is a schematic diagram of the structure and training process of a deep deterministic policy gradient controller for a bypass ash water chlorine washing system proposed by the present invention. Detailed implementation manners
[0086] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0087] Referring to Figures 1-3 , a bypass ash water chlorine washing system includes:
[0088] A data acquisition module for acquiring the operation data of the bypass ash water treatment system;
[0089] A data preprocessing module for performing standardization processing on the operation data to generate time series data that can be used for modeling;
[0090] A time series modeling module for extracting the time series features of the system operation status based on the Transformer network;
[0091] A kinetic modeling module for constructing a kinetic model of the reaction between ash water and chlorine, calculating the chlorine removal rate, and predicting the residual chlorine concentration;
[0092] A control strategy module for fusing time series features and kinetic features, training with a deep deterministic policy gradient controller, and outputting a bypass ash water control strategy;
[0093] An execution control module for adjusting the operation of the bypass ash water according to the control strategy to achieve ash water chlorine washing treatment;
[0094] A feedback optimization module for collecting the processing results, updating the kinetic model and the control strategy, and achieving closed-loop optimization control.
[0095] The bypass ash water washing dechlorination system proposed by the present invention constructs a modular structure, integrating data acquisition, time series modeling, kinetic modeling, reinforcement learning control, and closed-loop optimization functions, breaking through the bottlenecks of the rigid control strategy, response lag, and resource waste of traditional dechlorination systems, and realizing the intelligent, adaptive, and efficient operation of the dechlorination process of chlorine-containing media.
[0096] In this embodiment, the modules are implemented through the following methods:
[0097] S1. Collect the operation data of the bypass ash water treatment system;
[0098] S2. Preprocess the operation data to generate standardized time series data;
[0099] S3. Input the standardized time series data into a time series modeling network with a Transformer encoder as the main structure to extract the time series features of the system operation state;
[0100] S4. Build a kinetic model based on the chlorine-containing medium concentration, ash water pH value, reaction unit temperature, and reaction residence time in the operation data, calculate the chlorine removal rate per unit time, and predict the residual chlorine concentration;
[0101] S5. Jointly input the time series features extracted by the Transformer network and the chlorine removal features output by the kinetic model into a deep deterministic policy gradient controller to train the deep deterministic policy gradient controller including an Actor network and a Critic network;
[0102] S6. Output the ash water bypass control strategy by the deep deterministic policy gradient controller;
[0103] S7. Control the operation of the bypass ash water treatment system according to the ash water bypass control strategy to complete the ash water washing dechlorination process of the chlorine-containing medium;
[0104] S8. Collect the operation result data of the reaction system after the washing dechlorination treatment to generate feedback data;
[0105] S9. Update the parameters of the kinetic model according to the feedback data, and at the same time update the deep deterministic policy gradient controller.
[0106] The present invention defines the control logic flow between modules clearly, establishes a complete closed loop of data - model - control, improves the flow efficiency of internal information flow and model-driven ability of the system, enables the modules to operate in coordination, and enhances the real-time response ability and control accuracy of the system to changes in the operation state.
[0107] In this embodiment, the operation data includes chlorine-containing medium concentration, ash water pH value, ash water flow rate, reaction unit temperature, reaction residence time, and outlet chlorine concentration.
[0108] The present invention constructs a complete set of input variables by collecting multi-dimensional operating data, including parameters such as chlorine-containing medium concentration, pH value, temperature, flow rate and residence time, thereby enhancing the model's ability to recognize changes in operating conditions and providing high-quality input support for accurate modeling and strategy training.
[0109] In this implementation, the preprocessing includes missing value filling, outlier removal, normalization, and standardization.
[0110] The present invention improves the quality and consistency of the original operating data by introducing processing steps such as data missing filling, anomaly elimination, normalization and standardization, reduces the interference of outliers on the modeling and training process, and enhances the stability and generalization ability of the model.
[0111] In this implementation, S3 includes the following specific steps:
[0112] S31, standardize the time series data X = {x i (t j )} is input to the temporal modeling network with Transformer encoder as the backbone structure, where x i (t j ) means that at time step t j The standardized time series data value of the i-th variable above;
[0113] S32. Linearly map the standardized time series data to an embedding space of fixed dimension, and add a position encoding vector to each time step to obtain a time embedding representation:
[0114] e(t j )=W1x(t j )+b1+p(t j );
[0115] Among them, e(t j ) represents the time embedding representation, W1 represents the linear mapping weight matrix, x(t j ) represents the standardized time series data, b1 represents the bias vector, p(t j ) represents the position encoding vector, describing the time step feature vector after adding the position encoding;
[0116] S33. Input the sequence of time embedding representations of all time steps into the Transformer encoder module and calculate the output of each attention head:
[0117]
[0118] Among them, head k (tj ) represents the output of the k-th attention head at time step t j , where softmax represents normalization, Q k and K k and V k represent the query, key, and value matrices respectively, T represents the transpose operation, and d k represents the feature dimension of a single attention head;
[0119] S34. Concatenate the outputs of all attention heads to obtain the multi-head attention representation, and extract high-dimensional features through a linear transformation and a feed-forward neural network:
[0120] h(t j ) = FFN(A(t j )) = max(0, A(t j ))W2 + b2)W3 + b3;
[0121] where h(t j ) represents the high-dimensional feature vector at time step t j , A(t j ) represents the multi-head attention representation, describing the concatenation result of the multi-head attention, FFN represents the feed-forward neural network, max represents taking the maximum value, max(0, ·) represents the non-linear activation function, W2 and W3 represent the weight matrices of the feed-forward network, and b2 and b3 represent the biases of the feed-forward network;
[0122] S35. Apply the residual connection and layer normalization operations to the output h(t j ) at each time step to form the final encoded feature vector, and use the final encoded feature vector as the high-dimensional expression of the system operating state.
[0123] The present invention adopts a time series modeling network with a Transformer encoder as the core, which has strong capabilities of capturing long time series dependencies and multi-variable attention modeling, extracts key state features during the system operation process, improves the time series modeling accuracy and feature expression ability, and provides high-quality input for subsequent control strategy training.
[0124] In this embodiment, S4 includes the following specific steps:
[0125] S41. Obtain the initial concentration of the chlorine-containing medium, the pH value of the grey water, the temperature of the reaction unit, and the reaction residence time as the input variables of the reaction kinetics model;
[0126] S42. Establish a first-order reaction rate model between the grey water and the chloride, and define the differential relationship of the chlorine concentration changing with time:
[0127]
[0128] where C(t) represents the chlorine concentration at time t, represents the differential of chlorine concentration with respect to time t, and k eff represents the reaction rate constant after pH correction;
[0129] S43. Calculate the reaction rate constant using the Arrhenius formula based on the reaction unit temperature, and calculate the alkalinity adjustment factor based on the pH value to correct the reaction rate constant:
[0130]
[0131] k eff = k·α pH , α pH = log 10 (1 + 10 pH-7 );
[0132] where k represents the reaction rate constant, A represents the pre-exponential factor, E a represents the reaction activation energy, e represents the natural constant, R represents the gas constant, P represents the reaction unit temperature, and k eff represents the reaction rate constant after pH correction, and α pH represents the pH enhancement factor, which is a dimensionless constant;
[0133] S44. Calculate the predicted value of chlorine concentration during the reaction residence time under the condition of the initial concentration of the chlorine-containing medium:
[0134]
[0135] where C pred represents the predicted value of chlorine concentration, τ represents the reaction residence time, and C0 represents the initial concentration of the chlorine-containing medium;
[0136] S45. Calculate the chlorine removal rate based on the predicted value C pred and the input parameters:
[0137]
[0138] where R rem represents the average chlorine removal rate per unit time;
[0139] S46. Use the predicted value C pred of chlorine concentration and the chlorine removal rate R rem as chlorine removal characteristics and input them into the subsequent control strategy training model for the state space of the reinforcement learning controller.
[0140] The present invention constructs a reaction kinetic model by integrating variables such as the pH value of grey water, reaction temperature and time, calculates the chlorine removal rate and concentration prediction value, realizes the modeling and physical interpretation of the mechanism of the dechlorination process, provides a control basis with high credibility for the reinforcement learning strategy, and enhances the interpretability and prediction ability of the system control.
[0141] In this embodiment, S5 includes the following specific steps:
[0142] S51. Extract the system state encoding feature vector at the current time step. The feature vector is output by the Transformer network and combined with the predicted chlorine concentration value C pred calculated by the kinetic model rem and the average chlorine removal rate R per unit time
[0143] to splice and form a state vector; t S52. Construct a deep deterministic policy gradient controller. The deep deterministic policy gradient controller includes an Actor network and a Critic network. The Actor network is used to output a control action a t according to the current state s t , and the Critic network is used to estimate the value function Q(s t , a t ) of the state-action pair (s t );
[0144] S53. Define the action vector a t = [v t , q t , f t , τ t , where v t is the opening and closing state of the bypass valve, q t is the rotational speed of the grey water pump, f t is the grey water injection flow rate, and τ t is the residence time of the reaction unit;
[0145] S54. In the policy training stage, based on the current Actor policy, output the action a t , input the current state and action into the system environment to obtain an immediate reward and the next state, combine them to form a quadruple (s t , a t , r t , s t+1 ), store it in the experience replay pool, and randomly sample a batch of parameters for training the Actor network and the Critic network from the experience replay pool;
[0146] S55. Construct an objective value function containing the kinetic model deviation penalty, and calculate the target Q value for the state at the next time step:
[0147] y t = r t + γQ ′ (s t+1 , μ ′ (s t+1 )) - λ·|C pred - C real |;
[0148] where y t represents the target Q value, r t represents the immediate reward returned by the environment at time step t, γ represents the discount factor used to balance the current reward and future value, s t+1 represents the state at the next time step, μ ′ (s t+1 ) represents the action output by the target policy network, Q ′ (s t+1 , μ ′ (s t+1 )) represents the value estimation of the target Critic network for the state s t+1 at the next time step and the action output by the target policy network, λ represents the penalty weight for the dynamic model error, C pred represents the predicted value of the chlorine concentration, C real represents the true outlet chlorine concentration collected by the sensor;
[0149] S56. Construct the loss function and update the current Critic network parameters using the Adam optimizer by minimizing the loss function:
[0150]
[0151] where L represents the mean squared error loss function, N represents the sample batch size for each round of training, represents the state-action value estimation output by the current Critic network, represents the target Q value;
[0152] S57. Update the Actor network parameters using the policy gradient method to maximize the state-action value function, and the optimization objective function is:
[0153]
[0154] where J represents the policy objective function, represents the value estimation of the target Critic network for the state at time step t and the action output by the target policy network, represents the target Q value;
[0155] S58. Repeat the sampling, training, and parameter update processes to achieve the training and convergence of the greywater bypass control strategy that incorporates the error perception of the fusion dynamics model.
[0156] The present invention concatenates the output features of the time series modeling and the predicted features of the dynamics model into a state vector, inputs it into the deep deterministic policy gradient controller, and optimizes the policy training process in combination with the physical error penalty mechanism, thereby realizing the control of the greywater dosing strategy that integrates data learning and physical perception, and improving the policy training efficiency and control decision-making quality.
[0157] In this embodiment, the greywater bypass control strategy includes the opening and closing state of the bypass valve, the rotational speed of the greywater pump, the greywater injection flow rate, and the residence time of the reaction unit.
[0158] The control strategy output of the present invention covers core parameters such as the opening and closing of the bypass valve, pump rotational speed, flow rate, and reaction time, has the ability of multi-dimensional linkage adjustment, realizes the dynamic and fine-grained control of the greywater dechlorination process, meets the flexible control requirements under complex working conditions, and improves the overall control accuracy and execution flexibility of the system.
[0159] In this embodiment, the operation result data includes the post-treatment chlorine concentration, dechlorination efficiency, and control response change.
[0160] The present invention comprehensively evaluates the system control effect and feedback status by collecting operation results such as the post-treatment chlorine concentration, dechlorination efficiency, and control response change, constructs a data-driven closed-loop feedback link, provides accurate basis for model and controller updates, and improves the stability and optimization ability of system operation.
[0161] In this embodiment, S9 includes the following specific steps:
[0162] S91. Obtain feedback data, where the feedback data includes the post-treatment chlorine concentration, dechlorination efficiency, and control response change. The post-treatment chlorine concentration is the real-time measured concentration at the outlet of the reaction unit, and the dechlorination efficiency is calculated from the initial chlorine concentration and the outlet chlorine concentration;
[0163] S92. Compare the post-treatment chlorine concentration in the feedback data with the chlorine concentration predicted by the dynamics model, calculate the prediction error of the dynamics model, and use the prediction error to evaluate the performance deviation of the dynamics model;
[0164] S93. Fine-tune the parameters of the dynamics model based on the prediction error, set the set of parameters to be optimized, update the parameters of the dynamics model using the loss function based on the mean square error, and update the deep deterministic policy gradient controller simultaneously.
[0165] The present invention establishes a joint update mechanism for the kinetic model and control strategy based on operation feedback data, realizes continuous learning and adaptive optimization of the system during the actual operation process, effectively improves the robustness and regulation ability of the system to chlorine concentration fluctuations, and forms an intelligent dechlorination control system with the characteristics of self-learning, self-adjustment, and self-closed loop.
[0166] Example 1:
[0167] To verify the feasibility of the present invention in implementation, the present invention is applied to the chlorine-containing waste gas scrubbing section of an exhaust gas treatment station in a fine chemical industrial park in Jiangsu Province. This section mainly receives the tail gas discharged from multiple chemical synthesis reaction workshops. The chlorine concentration in the tail gas changes frequently and fluctuates greatly affected by the reaction load and raw material types. The traditional treatment system adopts the synchronous injection method and performs dosing control based on a rule model, having problems such as large fluctuations in dechlorination efficiency, excessive consumption of lye, and system response lag. The average dechlorination efficiency is 84.3%, and even short-term chlorine over-standard emissions occur in the high-load stage.
[0168] The present invention upgrades and transforms this treatment process by constructing a bypass ash water dechlorination system. One set of ash water bypass treatment unit is deployed in the project, equipped with an independent ash water pump, bypass valve, and reaction control unit, and integrated with data acquisition, preprocessing, time series modeling, kinetic modeling, and reinforcement learning control modules. The system uses ash water as the treatment medium, recycling the alkaline waste liquid generated by the workshop process to reduce the dosing cost. By arranging multiple groups of on-line sensors to real-time collect key operation parameters such as the chlorine-containing tail gas concentration, ash water pH value, ash water flow rate, reaction unit temperature, residence time, and outlet chlorine concentration, the data is input into a time series modeling network with a Transformer as the backbone structure after standardized preprocessing to extract the time series characteristics of the current operation condition; and jointly constituting the control state space with the chlorine removal rate and predicted chlorine concentration output by the reaction kinetic model, and inputting it into a deep deterministic policy gradient controller to generate the optimal bypass control strategy, controlling the ash water pump speed, injection flow rate, valve state, and residence time.
[0169] During the implementation process, the system runs for 18 hours every day, continuously monitors and collects operation data for more than 21 days. In the initial stage, the system does not use intelligent control and only adopts the manual timing switch and flow rate setting method for treatment. Subsequently, the intelligent bypass ash water control method proposed by the present invention is introduced.
[0170] Table 1 Monitoring data of cable branch boxes
[0171]
[0172] As can be seen from the experimental data table, the bypass ash water chlorine removal system proposed by the present invention exhibits overall performance superior to that of traditional manual control systems during actual operation. In terms of chlorine removal efficiency, the average chlorine removal efficiency of the traditional control system was 84.2% during the period from May 1st to May 3rd, while the average chlorine removal efficiency of the system of the present invention reached 94.6% from May 4th to May 7th, with an increase of more than 10 percentage points. This fully demonstrates that the present invention has higher processing capabilities and response accuracy in multi-variable dynamic control and state recognition. Especially in the context of significant fluctuations in the chlorine concentration of the tail gas, it can still maintain a stable and efficient chlorine removal effect.
[0173] In terms of resource utilization, the average ash water consumption per unit of the traditional system was 11.2 L / ton of treated gas, while the average ash water consumption per unit of the system of the present invention was 7.3 L / ton, with a decrease in unit resource consumption of more than 34.8%. This result indicates that while achieving a higher chlorine removal efficiency, the system of the present invention effectively reduces the processing cost, has good resource-saving and economic advantages, and is particularly suitable for industrial scenarios with large-scale continuous operation.
[0174] In terms of control response ability, the average response time of the system of the present invention is only 36 seconds, far lower than the 105-second response cycle of the traditional system. This indicates that after the intelligent control strategy real-time identifies changes in the system state, it can quickly give adjustment actions, effectively avoiding the problem of lag in the chlorine removal reaction and ensuring the smooth operation of the system and the continuous compliance of tail gas emissions.
[0175] In terms of system stability and operating condition adaptability, especially for the highly fluctuating emission scenario simulated during the period from May 10th to May 14th, the traditional system had 2 instances of exceeding the standard of chlorine concentration at the outlet within five days, with 4 times of manual intervention, and the average number of times of adapting to changes per day was only 2 times. While the system of the present invention, under the condition of fully autonomous operation, did not have any event of exceeding the chlorine standard, without the need for manual intervention, and could automatically complete approximately 19 dynamic adjustments per day. This fully reflects the adaptive optimization ability of the system of the present invention driven by reinforcement learning, which can continuously optimize the control strategy as the system operates, cope with complex and unstable emission environments, and ensure the long-term safe and efficient operation of the system.
[0176] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. A bypass ash water washing chlorine removal system, characterized in that, It includes: A data acquisition module for acquiring the operation data of the bypass grey water treatment system; A data preprocessing module for performing standardized processing on the operation data to generate time series data that can be used for modeling; A time series modeling module for extracting the time series features of the system operation status based on the Transformer network; A kinetic modeling module for constructing a kinetic model of the reaction between grey water and chlorine, calculating the chlorine removal rate and predicting the residual chlorine concentration; A control strategy module for fusing the time series features and kinetic features, training and outputting a bypass grey water control strategy using a deep deterministic policy gradient controller; An execution control module for adjusting the operation of the bypass grey water according to the control strategy to achieve chlorine removal treatment of the grey water wash; A feedback optimization module for collecting the processing results, updating the kinetic model and control strategy, and realizing closed-loop optimization control.
2. The bypass ash water washing chlorine removal system according to claim 1, wherein The modules are implemented through the following methods: S1. Acquire the operation data of the bypass grey water treatment system; S2. Preprocess the operation data to generate standardized time series data; S3. Input the standardized time series data into a time series modeling network with a Transformer encoder as the main structure to extract the time series features of the system operation status; S4. Based on the chlorine-containing medium concentration, grey water pH value, reaction unit temperature, and reaction residence time in the operation data, construct a kinetic model, calculate the chlorine removal rate per unit time, and predict the residual chlorine concentration; S5. Jointly input the time series features extracted by the Transformer network and the chlorine removal features output by the kinetic model into a deep deterministic policy gradient controller to train the deep deterministic policy gradient controller including an Actor network and a Critic network; S6. Output a grey water bypass control strategy by the deep deterministic policy gradient controller; S7. Control the operation of the bypass grey water treatment system according to the grey water bypass control strategy to complete the grey water wash treatment process of the chlorine-containing medium; S8. Collect the operation result data of the reaction system after the wash treatment to generate feedback data; S9. According to the feedback data, update the parameters of the kinetic model and simultaneously update the deep deterministic policy gradient controller.
3. The bypass ash water washing chlorine removal system according to claim 2, characterized in that, The operation data includes chlorine-containing medium concentration, grey water pH value, grey water flow rate, reaction unit temperature, reaction residence time, and outlet chlorine concentration.
4. A bypass ash water washing chlorine removal system according to claim 2, characterized in that, The preprocessing includes missing value filling, outlier removal, normalization processing, and standardization processing.
5. The bypass ash water washing chlorine removal system according to claim 2, wherein The S3 includes the following specific steps: S31. Input the standardized time series data X = {x i (t j )} into the time series modeling network with a Transformer encoder as the backbone structure, where x i (t j ) represents the standardized time series data value of the i-th variable at time step t j . S32. Perform a linear mapping on the standardized time series data, map it to an embedding space with a fixed dimension, and add a position encoding vector to each time step to obtain a time embedding representation; e(t j ) = W1x(t j ) + b1 + p(t j ); Among them, e(t j ) represents the time embedding representation, W1 represents the linear mapping weight matrix, x(t j ) represents the normalized time series data, b1 represents the bias vector, p(t j ) represents the position encoding vector, describing the time step feature vector after adding the position encoding; S33. Input the time embedding representations of all time steps into the Transformer encoder module as a sequence to calculate the output of each attention head; Among them, head k (t j ) represents the output of the k-th attention head at time step t j , softmax represents normalization, Q k and K k and V k represent the query, key, and value matrices respectively, T represents the transpose operation, d k represents the feature dimension of a single attention head; S34. Concatenate the outputs of all attention heads to obtain a multi-head attention representation, and extract high-dimensional features through a linear transformation and a feed-forward neural network; h(t j ) = FFN(A(t j )0 = max(0, A(t j )W2 + b2)W3 + b3; Among them, h(t j ) represents the high-dimensional feature vector at time step t j . A(t j ) represents the multi-head attention representation, describing the concatenation result of multi-head attention. FFN represents the feed-forward neural network, max represents taking the maximum value, max(0,·) represents the non-linear activation function, W2 and W3 represent the weight matrices of the feed-forward network, and b2 and b3 represent the biases of the feed-forward network; S35. Apply residual connection and layer normalization operations to the output h(t j ) to form the final encoded feature vector, and use the final encoded feature vector as the high-dimensional representation of the system operating state.
6. The bypass ash water washing chlorine removal system according to claim 2, characterized in that, The S4 includes the following specific steps: S41. Obtain the initial concentration of the chlorine-containing medium, the pH value of the ash water, the temperature of the reaction unit, and the reaction residence time as the input variables of the reaction kinetic model; S42. Establish a first-order reaction rate model between the ash water and the chloride, and define the differential relationship of the chlorine concentration changing with time: where C(t) represents the chlorine concentration at time t, represents the differential of the chlorine concentration with respect to time t, and k eff represents the reaction rate constant after pH correction; S43. Calculate the reaction rate constant using the Arrhenius formula according to the temperature of the reaction unit, and calculate the alkalinity adjustment factor based on the pH value to correct the reaction rate constant: k eff = k·α pH ,α pH = log 10 (1 + 10 pH-7 ); Among them, k represents the reaction rate constant, A represents the pre-exponential factor, E a represents the activation energy of the reaction, e represents the natural constant, R represents the gas constant, P represents the temperature of the reaction unit, k eff represents the reaction rate constant after pH correction, α pH represents the pH enhancement coefficient, which is a dimensionless constant; S44. Calculate the predicted value of the chlorine concentration within the reaction residence time under the condition of the initial concentration of the chlorine-containing medium: Among them, C pred represents the predicted value of chlorine concentration, τ represents the reaction residence time, and C0 represents the initial concentration of the chlorine-containing medium; S45. Calculate the chlorine removal rate based on the predicted value C pred and the input parameters: wherein, R rem represents the average chlorine removal rate per unit time; S46. Use the predicted value C of the chlorine concentration pred and the chlorine removal rate R rem as chlorine removal characteristics and input them into the subsequent control strategy training model for the state space of the reinforcement learning controller.
7. The bypass ash water washing chlorine removal system according to claim 2, characterized in that, The said S5 includes the following specific steps: S51. Extract the system state encoding feature vector at the current time step. The feature vector is output by the Transformer network and combined with the predicted chlorine concentration value C calculated by the kinetic model pred and the average chlorine removal rate R per unit time rem , and splice them to form a state vector S52. Construct a deep deterministic policy gradient controller, which includes an Actor network and a Critic network. The Actor network is used to output a control action a according to the current state s t . The Critic network is used to estimate the value function Q(s t , a t ) of the state-action pair (s t , a t ) t ; S53. Define the action vector a t = [v t , q t , f t , τ t , where v t is the opening and closing state of the bypass valve, q t is the rotation speed of the ash water pump, f t is the ash water injection flow rate, and τ t is the residence time of the reaction unit; S54. During the policy training phase, based on the current Actor policy, output the action a t , input the current state and action into the system environment to obtain the immediate reward and the next state, and combine them to form a quadruple (s t , a t , r t , s t+1 ), store it in the experience replay pool, and randomly sample a batch from the experience replay pool to train the parameters of the Actor network and the Critic network; S55. Construct an objective value function containing the penalty for the deviation of the kinetic model, and calculate the target Q value for the state of the next time step: y t = r t + γQ ′ (s t+1 , μ ′ (s t+1 )) - λ·|C pred - C real |; Among them, y t represents the target Q value, r t represents the immediate reward returned by the environment at time step t, γ represents the discount factor, which is used to balance the current reward and future value, s t+1 represents the state at the next time step, μ ′ (s t+1 ) represents the action output by the target policy network, Q ′ (s t+1 , μ ′ (s t+1 )) represents the value estimation of the state s t+1 at the next time step and the action output by the target policy network, λ represents the penalty weight of the dynamic model error, C pred represents the predicted value of the chlorine concentration, C real represents the true outlet chlorine concentration collected by the sensor; S56. Construct a loss function, and update the current Critic network parameters using the Adam optimizer by minimizing the loss function: where L represents the mean squared error loss function, N represents the sample batch size for each round of training, represents the state-action value estimate output by the current Critic network, represents the target Q value; S57. Update the Actor network parameters using the policy gradient method to maximize the state-action value function, and the optimization objective function is: where J represents the policy objective function, represents the value estimate of the state at time step t by the target Critic network and the action output by the target policy network ; represents the target Q value; S58. Repeat the processes of sampling, training, and parameter update to achieve the training and convergence of the ash water bypass control strategy with the perception of the kinetic model error.
8. The bypass ash water washing chlorine removal system according to claim 2, wherein The said ash water bypass control strategy includes the opening and closing state of the bypass valve, the rotation speed of the ash water pump, the ash water injection flow rate, and the residence time of the reaction unit.
9. The bypass ash water washing chlorine removal system according to claim 2, wherein The said operation result data includes the treated chlorine concentration, the dechlorination efficiency, and the control response change.
10. A bypass ash water washing chlorine removal system according to claim 2, characterized in that The said S9 includes the following specific steps: S91. Obtain the feedback data, which includes the treated chlorine concentration, the dechlorination efficiency, and the control response change. The treated chlorine concentration is the real-time measured concentration at the outlet of the reaction unit, and the dechlorination efficiency is calculated from the initial chlorine concentration and the outlet chlorine concentration; S92. Compare the treated chlorine concentration in the feedback data with the chlorine concentration predicted by the kinetic model, calculate the prediction error of the kinetic model, and use the prediction error to evaluate the performance deviation of the kinetic model; S93. Fine-tune the parameters of the kinetic model based on the prediction error, set the set of parameters to be optimized, use the loss function based on the mean square error to update the parameters of the kinetic model, and update the deep deterministic policy gradient controller at the same time.