A method for exhaust gas purification under multi-objective optimization
By employing hardware synchronization, data repair, and deep reinforcement learning methods, the control strategy of the exhaust gas purification system was optimized, solving the multi-objective optimization problem of exhaust gas purification technology in dynamic environments and achieving a comprehensive improvement in purification efficiency and energy consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2026-03-31
AI Technical Summary
Existing waste gas purification technologies are unable to cope with real-time changes in parameters such as waste gas composition, flow rate, and temperature, resulting in large fluctuations in purification efficiency, difficulty in controlling energy consumption and operating costs, and a lack of multi-objective optimization and dynamic adaptability.
Hardware time synchronization is performed using the IEEE 1588 protocol to correct outliers in exhaust gas data. Multi-scale state spaces and control actions are constructed. A deep reinforcement learning network architecture is combined to optimize the control method of the exhaust gas purification system. High-fidelity simulation is performed using a deep deterministic policy gradient algorithm and an adaptive weight network to optimize purification efficiency and energy consumption.
It achieves efficient and stable operation of the exhaust gas purification system in dynamic environments, comprehensively optimizes purification efficiency, energy consumption and operating costs, and enhances the robustness and adaptability of the system.
Smart Images

Figure CN120257815B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of waste gas treatment technology under deep learning, and in particular to a waste gas purification method under multi-objective optimization. Background Technology
[0002] As a crucial component of environmental protection, waste gas purification technology has garnered widespread attention due to the accelerating pace of industrialization and continuous urbanization. Traditional waste gas purification methods primarily include adsorption, absorption, catalytic conversion, and thermal incineration, all of which have achieved certain successes in reducing pollutant emissions. For instance, adsorption uses materials such as activated carbon to adsorb pollutants, absorption utilizes chemical solutions to neutralize harmful gases, and catalytic conversion accelerates the conversion of pollutants into harmless substances using catalysts. However, with the expansion of industrial production scale and increasingly stringent emission standards, the complexity of waste gas composition and the dynamic changes in emission conditions place higher demands on the adaptability of purification systems. In recent years, with the rapid development of computer technology and artificial intelligence, machine learning and reinforcement learning algorithms have been gradually introduced into the optimization control of waste gas purification systems. Among these, neural network-based predictive models and deep reinforcement learning algorithms have been used to optimize purification efficiency and energy consumption, but related research is still in its early stages, especially in multi-objective optimization and real-time control under dynamic environments, where systematic solutions have yet to be formed.
[0003] Currently, existing waste gas purification technologies have the following shortcomings in practical applications: First, traditional waste gas purification systems mostly rely on static models or empirical formulas, making it difficult to cope with real-time changes in parameters such as waste gas composition, flow rate, and temperature. This leads to significant fluctuations in purification efficiency and makes it difficult to effectively control energy consumption and operating costs. For example, single-objective optimization control strategies often focus only on purification efficiency or energy consumption, ignoring the interrelationships and trade-offs between purification efficiency, energy consumption, and operating costs, making it difficult to improve the overall system performance. Second, the time synchronization and data quality issues of hardware devices in waste gas purification systems also limit the accuracy and robustness of control strategies. For example, clock deviations between hardware devices may lead to asynchronous data acquisition, and the presence of abnormal data further reduces the reliability of the model. Third, existing technologies are still insufficient in handling multi-scale state changes and control action optimization in dynamic environments, lacking solutions that can comprehensively consider multi-objective optimization, dynamically adapt to environmental changes, and possess high-precision data processing capabilities. Therefore, there is an urgent need for a waste gas purification method that can improve purification efficiency, reduce energy consumption and operating costs, and possess high robustness and dynamic adaptability to meet the actual needs of industrial waste gas treatment. Summary of the Invention
[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a multi-objective optimization method for exhaust gas purification to solve the problems mentioned in the background art.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a waste gas purification method under multi-objective optimization, comprising:
[0007] Acquire exhaust gas data collected by each hardware component in the exhaust gas purification system, and based on the exhaust gas data collected by each hardware component, perform synchronization operations for each hardware component and abnormal value repair operations for the exhaust gas data.
[0008] Based on the hardware synchronization operations and outlier repair operations of the discarded data after execution, a multi-scale state space and control actions are defined to capture the dynamic characteristics of the exhaust gas purification system. Based on the dynamic characteristics of the exhaust gas purification system and combined with the deep deterministic policy gradient algorithm, a deep reinforcement learning network architecture is constructed to determine the optimization objective of the exhaust gas purification system.
[0009] Based on the optimization objectives of the exhaust gas purification system, a high-fidelity simulation of the exhaust gas purification system is conducted. Based on the simulation results, the control method of the exhaust gas purification system is optimized, thereby improving the exhaust gas purification efficiency.
[0010] As a preferred embodiment of the waste gas purification method under multi-objective optimization described in this invention, the synchronous operation of each hardware component includes:
[0011] Using the IEEE 1588 protocol, and through the hardware timestamp and master-slave clock dynamic compensation mechanism in the IEEE 1588 protocol, sub-microsecond time synchronization operations are performed on each piece of hardware in the exhaust gas purification system to obtain the operating status data of each piece of hardware.
[0012] As a preferred embodiment of the waste gas purification method under multi-objective optimization described in this invention, the outlier repair operation of the waste gas data includes:
[0013] By predicting the pollutant concentration in the current exhaust gas data using historical exhaust gas data, the predicted value of the pollutant concentration in the current exhaust gas data is obtained.
[0014] Regarding the pollutant concentration c in the current exhaust gas data i (t), if satisfying σ i If the standard deviation of the historical pollutant concentration is used, then outliers in the current exhaust gas concentration are corrected using the concentration gradient continuity interpolation method; otherwise, no processing is performed.
[0015] As a preferred embodiment of the waste gas purification method under multi-objective optimization described in this invention, the concentration gradient continuity interpolation method includes:
[0016] Define a dynamic repair window, centered on the outlier value of pollutant concentration in the current exhaust gas data at time t, extending forward and backward by one time step Δt, and selecting the pollutant concentration data c from the exhaust gas data at the previous time step. i (t-Δt) and the pollutant concentration data c in the exhaust gas data at the next time step. i (t+Δt);
[0017] Based on the pollutant concentration data c from the previous moment's exhaust gas data. i (t-Δt) and the pollutant concentration data c in the exhaust gas data at the next time step. i (t+Δt) is used to estimate the pollutant concentration curvature through a Savitzky-Golay filter, resulting in pollutant concentration data in the exhaust gas data after outlier correction.
[0018] As a preferred embodiment of the waste gas purification method under multi-objective optimization described in this invention, the method includes: defining a multi-scale state space and control actions based on the synchronized hardware operations and outlier repair operations of the waste data after execution, thereby capturing the dynamic characteristics of the waste gas purification system, including:
[0019] Based on discarded data and the operational status data of each hardware component, a multi-scale state space s is constructed. t ∈S, where S is the state space, s t It is the state vector of the exhaust gas purification system at the current time t;
[0020] Let the sequence of states over the past N times at the current time t be represented as {s} t-N ,s t-N+1 ,…,s t-1}, by calculating the trend of state change Δs t-i =s t-i -s t-i-1 For i = 1, 2, ..., N, construct the enhanced state vector S. t Then the enhanced state vector S t Represented as: s t =[s t ,Δs t-1 ,Δs t-2 ,…,Δs t-N];
[0021] Based on discarded data, multi-scale control actions are constructed, and the range of the multi-scale control actions is limited by the physical limits of the equipment or safety constraints.
[0022] As a preferred embodiment of the multi-objective optimization method for exhaust gas purification described in this invention, the following is included: A deep reinforcement learning network architecture is constructed based on the dynamic characteristics of the exhaust gas purification system and combined with a deep deterministic policy gradient algorithm to determine the optimization objective of the exhaust gas purification system, including:
[0023] Using a deep deterministic policy gradient algorithm, a continuous action vector a is input through an actor network. t and the enhanced state vector S t The continuous action vector a t The number of multi-scale control actions is determined by the defined parameters;
[0024] The actor network consists of a three-layer fully connected network with dimensions of 128 and 64 for hidden layer 1 and hidden layer 2, respectively. The activation function is ReLU, and the output layer uses the tanh function to normalize the output to the range [-1,1].
[0025] Based on the normalized output, an adaptive weight network is constructed, and an instantaneous reward function r is defined in the adaptive weight network. t The instant reward function r t It includes the purification efficiency, energy consumption, and operating cost of the exhaust gas purification system, as well as the weight values corresponding to the purification efficiency, energy consumption, and operating cost;
[0026] The weight values are determined by the neural network f. θ Dynamically generated, the neural network f θ The input is a multi-scale state space and the instantaneous reward function r. t The output is the weighted value corresponding to the purification efficiency, energy consumption, and operating cost, and satisfies ∑w i =1, but the initial weight values are uniformly distributed;
[0027] The adaptive weighted network is a two-layer fully connected network with an input dimension of (N+1)·(S). dim +r dim ), S dim For S t The dimension, r dim For r t The dimensions are: 64 for the hidden layer, ReLU for the activation function, and softmax for the output layer.
[0028] As a preferred embodiment of the multi-objective optimization method for waste gas purification described in this invention, the method includes: performing high-fidelity simulation of the waste gas purification system under high-fidelity conditions based on the optimization objectives of the waste gas purification system, including:
[0029] Using open-source CFD tools, a physical model was built to simulate the flow and chemical reactions of exhaust gases in an exhaust gas system. The physical model was input with an enhanced state vector S. t and continuous action vector a t Output the next enhanced state vector S t+1 and instant reward function r t ;
[0030] An adversarial sample generator is constructed and fused with the physical model to generate combinations of exhaust gas components under different scenarios, i.e., simulation results.
[0031] As a preferred embodiment of the multi-objective optimization-based waste gas purification method of the present invention, the adversarial sample generator includes:
[0032] Consider the error in establishing the physical model and the input continuous action vector a t The deviation is used to calculate the Q value for the continuous action vector a. t The gradient of the input continuous action vector a, and the gradient of the input continuous action vector a. t Inject noise.
[0033] As a preferred embodiment of the multi-objective optimization method for waste gas purification described in this invention, the control method of the waste gas purification system is optimized based on simulation results, including:
[0034] Based on the simulation results, an anti-regularization loss assessment is performed to obtain the robustness under the simulation result loss. By minimizing the robustness, the exhaust gas purification system can maintain optimal control under different scenarios, thereby improving the exhaust gas purification efficiency.
[0035] Compared with existing technologies, the beneficial effects of the invention are as follows:
[0036] 1. This invention constructs an enhanced state vector, which comprehensively considers the current state and historical state of the exhaust gas purification system, and can dynamically capture the dynamic characteristics of the exhaust gas purification system, thereby accurately obtaining the dynamic information of the exhaust gas purification system.
[0037] 2. Based on the deep deterministic policy gradient algorithm, an adaptive weight network is constructed to achieve comprehensive optimization of purification efficiency, energy consumption and operating costs in the waste gas purification system. Furthermore, the adaptive weight network can dynamically adjust the weights of each objective according to the operating status of the waste gas purification system, ensuring that while improving purification efficiency, the energy consumption and operating costs of waste gas purification are effectively reduced.
[0038] 3. By establishing a physical model and an adversarial sample generator, a high-fidelity simulation environment was constructed, which accurately simulated the dynamic process of waste gas flow and chemical reaction in the waste gas purification system. In addition, by evaluating the adversarial regularization loss of the simulation results, the control method of the waste gas purification system under extreme conditions was optimized, the stability of the waste gas purification system under different scenarios was enhanced, and the high efficiency of purification effect in long-term operation was ensured. Attached Figure Description
[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0040] Figure 1 This is a flowchart illustrating the overall process of a multi-objective optimization method for exhaust gas purification according to an embodiment of the present invention.
[0041] Figure 2 This is a time series comparison chart of the purification efficiency and energy consumption of the waste gas purification method under multi-objective optimization according to an embodiment of the present invention.
[0042] Figure 3 This is a multi-dimensional performance comparison chart of the waste gas purification method under multi-objective optimization according to an embodiment of the present invention. Detailed Implementation
[0043] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0044] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0045] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0046] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0047] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0048] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0049] Example 1
[0050] Reference Figure 1 This is the first embodiment of the present invention, which provides a waste gas purification method under multi-objective optimization, including:
[0051] S1. Obtain the exhaust gas data collected by each hardware in the exhaust gas purification system, and perform synchronization operations for each hardware and abnormal value repair operations for the exhaust gas data based on the exhaust gas data collected by each hardware.
[0052] Specifically, the concentrations of pollutants (NOx, CO, PM, unit: ppm) in the exhaust gas purification system are obtained using an FTIR spectrometer and a particulate matter sensor; the exhaust gas temperature (unit: °C) is obtained using a thermocouple; and the exhaust gas flow rate (unit: m³) is obtained using a mass flow meter. 3 / h), catalyst temperature (unit: °C) is obtained by thermocouple, and chemical reagent injection rate (unit: L / h) is obtained by ultrasonic flow meter;
[0053] Specifically, the exhaust gas data consists of the concentration of pollutants in the exhaust gas, the exhaust gas temperature, the exhaust gas flow rate, the catalyst temperature, and the amount of chemical reagents injected.
[0054] It needs to be explained that in scenarios where the composition of exhaust gas fluctuates rapidly (such as exhaust gas from chemical plant reactors), the actual exhaust gas parameters of the exhaust gas purification system will change significantly within Δt. At this time, the exhaust gas purification system may still issue control commands based on asynchronous data. In this case, the issuance of control commands will lead to a mismatch between the real-time exhaust gas parameters. Therefore, it is necessary to eliminate the phase difference, that is, to eliminate the time difference between various hardware components.
[0055] Furthermore, using the IEEE 1588 protocol, the hardware in the exhaust gas purification system is synchronized at the sub-microsecond level through hardware timestamps (PHY layer marking data packet time) and a master-slave clock dynamic compensation mechanism (including link delay measurement and clock frequency correction) to obtain the operating status data of each hardware component.
[0056] It should be noted that sub-microsecond (<1 microsecond) synchronization accuracy can be achieved through the hardware timestamp and master-slave clock dynamic compensation mechanism in the IEEE 1588 protocol. This is crucial for eliminating phase differences, especially when the device sampling interval (e.g., 1Hz) needs to be strictly aligned.
[0057] Furthermore, by using historical exhaust gas data to predict the pollutant concentration in the current exhaust gas data, the predicted value of the pollutant concentration in the current exhaust gas data is obtained.
[0058] Specifically, the Autoregressive Differential Moving Average (ARIMA) model is used to predict the pollutant concentration in the current exhaust gas data based on historical exhaust gas data, thus obtaining the predicted value of the pollutant concentration in the current exhaust gas data.
[0059] Furthermore, regarding the pollutant concentration c in the current exhaust gas data... i (t), if satisfying If the outlier values of pollutant concentration in the current exhaust gas data are corrected by the concentration gradient continuity interpolation method, then no action is taken.
[0060] Where, σ i Expressed as the standard deviation of historical pollutant concentrations;
[0061] It should be noted that if the pollutant concentration c in the current exhaust gas data... i If (t) follows a normal distribution, then 99.7% of the normal values should fall within the range of 0.5%. If the value exceeds the specified range, it is considered abnormal.
[0062] Specifically, the concentration gradient continuity interpolation method is as follows:
[0063] Define a dynamic repair window, centered on the outlier value of pollutant concentration in the current exhaust gas data at time t, extending forward and backward by one time step Δt, and selecting the pollutant concentration data c from the exhaust gas data at the previous time step. i (t-Δt) and the pollutant concentration data c in the exhaust gas data at the next time step. i (t+Δt);
[0064] Based on the pollutant concentration data c from the previous moment's exhaust gas data. i (t-Δt) and the pollutant concentration data c in the exhaust gas data at the next time step. i (t+Δt) is used to estimate the pollutant concentration curvature through a Savitzky-Golay filter, resulting in pollutant concentration data in the exhaust gas data after outlier correction.
[0065] Specifically, the pollutant concentration data in the exhaust gas data after outlier correction. Represented as:
[0066]
[0067] in, This is expressed mathematically as a representation of the continuity of data using the Savitzky-Golay filter, and its function is to calculate the second derivative of the fitted curve. Simultaneously estimate the pollutant concentration curvature;
[0068] For example, suppose the pollutant concentration data at a certain time t is abnormal, and the data in the window is: [c(t-2)=10ppm, c(t-1)=12ppm, c(t)=50ppm, c(t+1)=14ppm, c(t+2)=16ppm]
[0069] There are a total of 5 data points. We can see that c(t) = 50ppm is an outlier. Therefore, we use a Savitzky-Golay filter, fit the data within the window using a quadratic polynomial, and solve for the coefficient matrix b0, b1, b2 using the least squares method. We then extract twice the value of b2. As the curvature of the current pollutant concentration, the second derivative is substituted into... The formula yields:
[0070]
[0071] As can be seen, the outlier was corrected from 50ppm to 13.5ppm, which is consistent with the trend of the data before and after, while retaining the curvature information of the current pollutant concentration.
[0072] Specifically, if the outlier at time t is in a continuous state, meaning that even if the pollutant concentration data in the exhaust gas data at the previous and next time moments cannot be selected by extending the time by Δt, then the time by Δt is gradually extended until the pollutant concentration data in the exhaust gas data at the previous and next time moments is found.
[0073] S2. Based on the hardware synchronization operations and outlier repair operations of the discarded data after execution, define a multi-scale state space and control actions to capture the dynamic characteristics of the exhaust gas purification system. Based on the dynamic characteristics of the exhaust gas purification system and combined with the deep deterministic policy gradient algorithm, construct a deep reinforcement learning network architecture to determine the optimization objective of the exhaust gas purification system.
[0074] Furthermore, based on the pollutant concentration, exhaust gas temperature, exhaust gas flow rate, catalyst temperature, chemical reagent injection volume, and the operational status data of each hardware component in the waste data, a multi-scale state space s is constructed. t ∈S, where S is the state space, s t It is the state vector of the exhaust gas purification system at the current time t;
[0075] Furthermore, the sequence of states over the past N times at the current time t is represented as {s}. t-N ,s t-N+1 ,…,s t-1}, by calculating the trend of state change Δs t-i =s t-i -s t-i-1 For i = 1, 2, ..., N, construct the enhanced state vector S. t Then the enhanced state vector S t Represented as: S t =[s t ,Δs t-1 ,Δs t-2 ,…,Δs t-N ];
[0076] Specifically, S = {S t |S t =[s t ,Δs t-1 ,Δs t-2 ,…,Δs t-N ],s t ∈R d ,Δs t-i ∈R d}, where d is the dimension of the state vector (i.e., the number of state variables in the exhaust gas purification system), and R represents the set of real numbers;
[0077] It should be noted that the enhanced state vector S tIt not only includes the state vector of the exhaust gas purification system at the current time t, but also the state change trend of the past N current times t, so it can better capture the trend of the exhaust gas purification system's state evolution over time and history.
[0078] Specifically, based on the catalyst temperature, exhaust gas flow rate, and chemical reagent injection amount in the waste data, a multi-scale control action is constructed, assuming its action range is catalyst temperature adjustment ΔT∈[-5°·W]. max ,5°·W max The adjustment of exhaust gas flow rate ΔF ∈ [-10%·F] max 10% F max The change in the amount of chemical reagent injected ΔH ∈ [-20%·H] max 20% H max ];
[0079] Among them, W max F max H max It needs to be determined based on the physical limits or safety constraints of the equipment;
[0080] Furthermore, a deep deterministic policy gradient algorithm is used, with continuous action vectors a input through the actor network. t and the enhanced state vector S t Continuous action vector a t The number of multi-scale control actions is determined by the number of actions defined. For example, if the multi-scale control actions considered above are three types, then a t = [ΔT, ΔF, ΔH];
[0081] Specifically, the actor network consists of three fully connected layers, with hidden layer 1 and hidden layer 2 having dimensions of 128 and 64 respectively, using ReLU as the activation function, and the output layer using the tanh function to normalize the output to the range [-1,1].
[0082] Specifically, the three-layer structure of the actor network is represented as: input layer → hidden layer 1 (dimension 128) → hidden layer 2 (dimension 64) → output layer;
[0083] Specifically, in the deep reinforcement learning network architecture, the learning rate is set to 0.001 for the actor network, 0.002 for the critic network, and a discount factor of 0.99.
[0084] Furthermore, based on the normalized output, an adaptive weight network is constructed, and an instantaneous reward function r is defined in the adaptive weight network. t Instant reward function r t It includes the purification efficiency, energy consumption, and operating cost of the exhaust gas purification system, as well as the corresponding weight values for purification efficiency, energy consumption, and operating cost;
[0085] Specifically, the defined instant reward function r t This can be expressed as a mathematical formula:
[0086] r t =w1·r eff +w2·r energy +w3·r cost
[0087] Where w1, w2, and w3 represent the purification efficiency r, respectively. eff Energy consumption r energy Operating costs cost The corresponding weight value;
[0088] Furthermore, the weight values are determined by the neural network f. θ Dynamically generated, neural network f θ The inputs are a multi-scale state space and an instantaneous reward function r. t The output consists of weighted values for purification efficiency, energy consumption, and operating costs, satisfying ∑w i =1, but the initial weight values are uniformly distributed.
[0089] Specifically, the adaptive weight network is a two-layer fully connected network with an input dimension of (N+1)·(S). dim +r dim ), S dim For the enhanced state vector S t The dimension, r dim For the instant reward function r t The dimension of the hidden layer is 64, the activation function is ReLU, and the output layer uses the softmax function.
[0090] Specifically, the optimizer for this adaptive weight network is Adam, with a learning rate of 0.001;
[0091] S3. Based on the optimization objectives of the exhaust gas purification system, conduct high-fidelity simulation of the exhaust gas purification system. Based on the simulation results, optimize the control method of the exhaust gas purification system to improve the exhaust gas purification efficiency.
[0092] Furthermore, open-source CFD tools (such as OpenFOAM) are used to build a physical model to simulate the flow and chemical reactions of exhaust gases in the exhaust gas system. The input to the physical model is the enhanced state vector S. t and continuous action vector a t The output is the next enhanced state vector S. t+1 and instant reward function r t ;
[0093] Specifically, the physical models include thermodynamic equations, chemical reaction kinetics, and mass conservation equations;
[0094] Furthermore, an adversarial sample generator is constructed and fused with a physical model to generate combinations of exhaust gas components under different scenarios, i.e., simulation results.
[0095] Specifically, the fusion method can be expressed by the mathematical formula as follows:
[0096] S t+1 =α·S phy (S t ,a t )+(1-α)·G(S t ,a t )
[0097] Among them, S phy Let G represent the physical model, G represent the adversarial sample generator, and α represent the selection coefficient between the physical model and the adversarial sample generator. If α = 1, the physical model is used entirely; if α = 0, the adversarial sample generator is used entirely. In the initial stage, let α be between 0.5 and 1, which means that the initial simulation stage needs to rely more on the physical model.
[0098] Specifically, through an adversarial example generator, the error in building the physical model and the input continuous action vector a are considered. t The deviation is used to calculate the Q-value for the continuous action vector a. t The gradient of the input continuous action vector a, and the gradient of the input continuous action vector a. t Injected noise, expressed mathematically as:
[0099]
[0100] Where, Δa t Represented as a continuous action vector a t The perturbation is ∈, where ∈ is the magnitude of the perturbation, and sign is the sign function, which binarizes the gradient direction (taking +1 or -1) to avoid the excessive influence of the gradient magnitude on the perturbation. The output Q(S) of the Critic network is represented as... t ,a t The gradient of the action variable;
[0101] Specifically, the standard deviation of the injected noise is initially 0.1, and the attenuation rate is 0.995.
[0102] Specifically, the Q-value is obtained through the Critic network, which evaluates the continuous action vector 'a' output in the actor network. t Good or bad;
[0103] It should be noted that by taking into account the perturbation of the continuous action vector, the control error of the exhaust gas purification system under the worst case or the inaccuracy of the physical model (such as catalyst deactivation, sudden change in exhaust gas composition, etc.) is simulated, so that the simulation results are close to the actual effect of the exhaust gas purification system.
[0104] Furthermore, based on the simulation results, an anti-regularization loss assessment is performed to obtain the robustness under the simulation result loss. By minimizing the robustness, the exhaust gas purification system can maintain optimal control under different scenarios, thereby improving the exhaust gas purification efficiency.
[0105] Specifically, the robustness under simulation result loss is obtained through mathematical formulas:
[0106]
[0107] Among them, L robust Represented as robustness under simulation result loss, The expression is represented by the mathematical expectation operator, which represents a weighted average over all possible cases; D represents a dataset storing past experiences, also known as an experience replay buffer, which contains... t ,a t ,r t ,S t+1 Its experience replay buffer size is 10 7 Batch size is 128; S t ~D represents the enhanced state vector S t Sampling is performed from the experience playback buffer D;
[0108] It should be noted that, through This allows the loss assessment against regularization to be not limited to the current state encountered by the exhaust gas purification system, but to cover various operating conditions that may occur during the long-term operation of the exhaust gas purification system.
[0109] Furthermore, when the adversarial example generator G is at ||Δa t Within the range ||≤∈, find the perturbation with the largest increase in Q value and then update L. robust Δa t This allows us to obtain robustness while minimizing the loss of simulation results.
[0110] Example 2
[0111] Reference Figure 2 and Figure 3 This is the second embodiment of the present invention, which provides a waste gas purification method under multi-objective optimization, including: verifying the application effect of the present invention in a real industrial scenario and comparing it with traditional waste gas purification technologies (adsorption and catalytic methods); the waste gas purification system of a sintering workshop in a steel plant was selected as the test object, and the system processes waste gas with a flow rate of 1200 m³ / h. 3 / h, the main pollutants include nitrogen oxides (NOx), carbon monoxide (CO) and particulate matter (PM). The optimization goal is to improve purification efficiency, reduce energy consumption and operating costs, and ensure the stability of the system in dynamic environments.
[0112] The experimental equipment configuration is as follows: FTIR spectrometer (Thermo Fisher Nicolet iS50), used for real-time monitoring of NOx and CO concentrations, with a measurement unit of ppm and a resolution of 0.1 ppm; particulate matter sensor (TSI 3330), used to measure PM concentration, with a unit of mg / m³. 3 Accuracy ±0.5mg / m 3 Thermocouple (Omega K type) is used to monitor exhaust gas temperature and catalyst bed temperature, measured in °C, with a measurement range of -50 °C to 1350 °C; Mass flow meter (Alicat M-100SLPM-D) is used to monitor exhaust gas flow rate, measured in m³ / s. 3 / h, accuracy ±0.8%; ultrasonic flow meter (Siemens FUS1010) for measuring the injection volume of chemical reagent (urea solution) in L / h, accuracy ±1%; computing platform equipped with Intel Core i7 processor, 32GB RAM and NVIDIA RTX 3080 GPU for running deep reinforcement learning models and CFD simulations.
[0113] The above experimental equipment configuration was simulated using OpenFOAM, and data from the first 0-50 minutes was used for verification on an hourly basis. The experimental results are referenced. Figure 2 As shown, the purification efficiency of the method of the present invention fluctuates between 87.3% and 95%, with an average of approximately 92%. This indicates that even under a sudden increase in flow rate, the exhaust gas purification system can quickly recover to over 90% when using the present invention. In contrast, the traditional method's efficiency ranges from 73% to 88%, with an average of approximately 85%. This suggests that the traditional method recovers more slowly after a sudden increase in flow rate, indirectly verifying the stability of the exhaust gas purification system using the present invention under dynamic environments. Furthermore, regarding energy consumption, the system consumes between 42.1 and 44.7 kW, with an average of approximately 42 kW, indicating relatively small fluctuations in energy consumption. In contrast, the traditional method consumes between 42.5 and 55 kW, with an average of approximately 50 kW, indicating higher energy consumption. This clearly demonstrates the energy efficiency advantage of the present invention.
[0114] In addition, through experiments Figure 3 By projecting purification efficiency, energy consumption, and response time into a three-dimensional space for comprehensive analysis, it can be seen that the scatter distribution of the proposed solution is basically located in the region of high purification efficiency (90%–94%), low energy consumption (40–43 kW), and short response time (4–5 min), while the scatter distribution of the traditional solution is basically located in the region of lower purification efficiency (80%–90%), higher energy consumption (48–53 kW), and longer response time (5–7 min). Although the addition of response time resulted in slight deviations in purification efficiency and energy consumption, it did not affect the overall experimental simulation results. Figure 2 and Figure 3 The parameters are relatively similar;
[0115] In summary, the present invention outperforms traditional solutions in terms of purification efficiency, stability, energy consumption, and response time. It not only achieves multi-objective balance optimization but also maintains the stability control of the exhaust gas purification system under dynamic environments, providing a reference for improving the performance of exhaust gas purification systems.
[0116] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0117] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0120] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0121] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for exhaust gas purification under multi-objective optimization, characterized by, The method comprises the following steps: acquiring exhaust data collected by each hardware in an exhaust purification system, performing each hardware synchronization operation and an outlier repair operation of the exhaust data based on the exhaust data collected by each hardware; the outlier repair operation of the exhaust data comprises: Predicting the concentration of pollutants in current exhaust gas data from historical exhaust gas data, to obtain a predicted value of the concentration of pollutants in the current exhaust gas data ; For the pollutant concentration in the current exhaust data , if the following condition is satisfied , , which is the standard deviation of the pollutant historical concentration, the abnormal value of the pollutant concentration in the current exhaust data is repaired by the concentration gradient continuity interpolation method, otherwise no processing is performed; the concentration gradient continuity interpolation method comprises: A dynamic repair window is defined, with the abnormal value of the pollutant concentration in the current exhaust data at the current time t as the center, extending one Δt time step forward and backward, and selecting the pollutant concentration data in the exhaust data at the previous time and the pollutant concentration data in the exhaust data at the next time . Based on selecting the pollutant concentration data in the exhaust data at the previous moment and the pollutant concentration data in the exhaust data at the next moment , estimating the pollutant concentration curvature through a Savitzky-Golay filter to obtain the pollutant concentration data in the exhaust data after repairing the abnormal value ; Specifically, the pollutant concentration data in the exhaust gas data after repairing the abnormal value is represented as: wherein, Savitzky-Golay filter on the data continuity, the role is to calculate the second derivative of the fitted curve Simultaneously estimate the pollutant concentration curvature; According to the performed each hardware synchronization operation and the outlier repair operation of the exhaust data, a multi-scale state space and a control action are defined, the dynamic characteristics of the exhaust purification system are captured, a deep reinforcement learning network architecture is constructed based on the dynamic characteristics of the exhaust purification system combined with a deep deterministic policy gradient algorithm, and an optimization target of the exhaust purification system is determined; Through the optimization target of the exhaust purification system, simulation of the exhaust purification system in a high-fidelity environment is performed, and based on the simulation result, the control mode of the exhaust purification system is optimized, thereby improving the exhaust purification efficiency.
2. The method of exhaust gas purification under multi-objective optimization according to claim 1, characterized by, The each hardware synchronization operation comprises: Using the IEEE 1588 protocol, the hardware time stamp and the master-slave clock dynamic compensation mechanism in the IEEE 1588 protocol are used to perform sub-microsecond time synchronization operation on each hardware in the exhaust purification system, and the running state data of each hardware itself is obtained.
3. The exhaust gas purification method under multi-objective optimization according to claim 1 or 2, characterized by, According to the performed each hardware synchronization operation and the outlier repair operation of the exhaust data, a multi-scale state space and a control action are defined, the dynamic characteristics of the exhaust purification system are captured, including: According to the discarded data and the operation state data of each hardware, a multi-scale state space is constructed , S is a state space, is a state vector of the exhaust gas purification system at the current time t; The state sequence of the past N current time t is represented as , the enhanced state vector is constructed by calculating the state change trend The enhanced state vector is represented as: ; According to the exhaust data, a multi-scale control action is constructed, and the range of the multi-scale control action is limited by the device physical limit value or the safety constraint.
4. The exhaust gas purification method under multi-objective optimization according to claim 3, characterized by, According to the dynamic characteristics of the exhaust purification system combined with a deep deterministic policy gradient algorithm, a deep reinforcement learning network architecture is constructed, and an optimization target of the exhaust purification system is determined, including: using a deep deterministic policy gradient algorithm, by an actor network inputting a continuous action vector and an enhanced state vector , the continuous action vector is determined by the number of multi-scale control actions defined; The actor network is composed of three fully connected networks, the dimensions of the hidden layer 1 and the hidden layer 2 are 128 and 64 respectively, the activation function is ReLU, the output layer uses tanh function, and the output result is normalized to the range [-1, 1]; According to the normalized output result, an adaptive weight network is constructed, and an instant reward function is defined in the adaptive weight network , wherein the instant reward function contains the purification efficiency, energy consumption, operation cost of the exhaust evolution system, and weight values corresponding to the purification efficiency, energy consumption, and operation cost. The weight value is generated by a neural network The neural network is dynamically generated The input of the neural network is a multi-scale state space and the instant reward function The output of the neural network is a weight value corresponding to the purification efficiency, energy consumption, and operation cost, and satisfies But the initialized weight value is a uniform distribution The adaptive weight network is a two-layer fully connected network, the input dimension is , is dimension, is dimension; the hidden layer dimension is 64, the activation function is ReLU, and the output layer adopts the softmax function.
5. The method of purifying exhaust gas under multi-objective optimization according to claim 4, characterized by, Through the optimization target of the exhaust purification system, simulation of the exhaust purification system in a high-fidelity environment is performed, including: using an open source CFD tool, a physical model is built to simulate the flow and chemical reactions of the exhaust gas in the exhaust system, the physical model inputs an enhanced state vector and a continuous action vector , outputs a next enhanced state vector and an immediate reward function ; An adversarial sample generator is constructed to fuse with the physical model to generate different combinations of exhaust components under different scenarios, i.e. simulation results.
6. The exhaust gas purification method under multi-objective optimization according to claim 5, characterized by, The adversarial sample generator comprises: considering errors in building the physical model and deviations in the input continuous action vector , compute the gradient of the Q-value with respect to the continuous action vector , and inject noise into the input continuous action vector .
7. The method of purifying exhaust gas under multi-objective optimization according to claim 5, characterized by, Based on the simulation result, the control mode of the exhaust purification system is optimized, including: Based on the loss evaluation of the adversarial regularization of the simulation result, the robustness under the loss of the simulation result is obtained, the robustness is minimized, the exhaust purification system maintains optimal control under different scenarios, and the exhaust purification efficiency is improved.
Citation Information
Patent Citations
Central air conditioner purification method and system based on deep reinforcement learning
CN118836523A