Low carbon sub-micron particulate matter low pressure charge condensation method and device
By combining multi-scale spatiotemporal synchronous perception with deep reinforcement learning agents, the low-voltage charged coagulation technology can operate efficiently under low energy consumption, solving the problems of high energy consumption and poor adaptability to operating conditions in existing technologies, and improving the coagulation and capture efficiency of submicron particles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINESE RES ACAD OF ENVIRONMENTAL SCI
- Filing Date
- 2025-12-03
- Publication Date
- 2026-05-05
AI Technical Summary
Existing charged coalescing technology has high energy consumption, poor adaptability to operating conditions, and difficulty in achieving precise and efficient operation under low-pressure conditions, resulting in insufficient control accuracy.
By employing multi-scale spatiotemporal synchronous perception, multi-scale data fusion and state representation, deep reinforcement learning agents, and predictive collaborative decision-making methods, and through low-voltage multi-energy coupled strong charging and low-voltage pulse adaptive condensation, combined with a deep reinforcement learning agent using a multi-objective reward function with total power consumption as the penalty term, the charging and condensation parameters are adjusted in real time to achieve low-energy and high-efficiency operation.
While ensuring the coagulation effect of submicron particles, the overall operating voltage and power consumption of the coagulation system are reduced, operating costs and carbon emissions are decreased, and real-time response and dynamic adjustment to complex working conditions are achieved, thereby improving the coagulation effect and subsequent collection efficiency.
Smart Images

Figure CN121623955B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electrostatic precipitator technology, specifically to a method and apparatus for low-voltage electrostatic coagulation of low-carbon submicron particles. Background Technology
[0002] Submicron particulate matter (PM1.0) poses a serious threat to the environment and human health due to its small particle size, large specific surface area, and long residence time. Charged agglomeration technology is an effective means of treating this type of particulate matter. Its basic principle is to charge the particles and then use an electric field to cause them to collide and agglomerate into large-sized aggregates, which are then easier to capture.
[0003] Existing charged coalescence techniques typically employ high-voltage direct current or alternating current to apply a high-intensity electric field to the charged and coalescence regions. However, such techniques suffer from the following significant drawbacks:
[0004] High energy consumption and high cost: Continuous high-pressure operation results in huge overall power consumption of the system, which directly increases the operating cost of the equipment and indirect carbon emissions, which is not in line with the development trend of green and low-carbon.
[0005] Poor adaptability to operating conditions: The particulate matter concentration, particle size, temperature, humidity, and other parameters of actual industrial flue gas are dynamically changing. Existing technologies mostly use fixed or experience-based operating parameters, which are difficult to match with changing operating conditions in real time. This often leads to energy waste (over-powered equipment) when the particulate matter concentration is low, or insufficient condensation efficiency when the concentration is high.
[0006] Low-voltage operation results in low efficiency: To ensure sufficient charge and coagulation force, existing technologies generally rely on high voltage. Simply reducing the voltage will directly lead to insufficient charge and weakened coagulation, thus failing to guarantee the treatment effect.
[0007] Insufficient control precision: Existing systems mostly rely on macroscopic parameters such as outlet concentration for coarse feedback adjustment, lacking real-time and precise perception of the microscopic state of the coagulation process (such as aggregate structure and dynamic behavior), resulting in control lag and limited optimization effect. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides a low-pressure charged coalescence method and apparatus for low-carbon submicron particles, which solves the problems of high energy consumption, poor adaptability to operating conditions, and difficulty in achieving precise and efficient operation under low-pressure conditions in existing charged coalescence technologies.
[0009] To achieve the above objectives, the first aspect of the present invention provides a method for low-voltage charged coagulation of low-carbon submicron particles, comprising the following steps:
[0010] Step S1: Multi-scale Spatiotemporal Synchronous Sensing. This step is used to acquire multi-dimensional spatiotemporal data characterizing the environmental field, micromorphology, and mesoscopic distribution of the charged coagulation process in real time and comprehensively, and to perform spatiotemporal synchronous processing on the multi-dimensional spatiotemporal data. The sensing includes deploying sensor arrays in the charged region and the coagulation region to acquire local temperature field, local humidity field, local airflow velocity field, and local electric field intensity distribution in real time. In addition, a laser particle size spectrometer is set up at the downstream outlet of the coagulation region to acquire the particle size distribution function in real time. Simultaneously, a high-speed camera is deployed in conjunction with a sheet-like laser light source to perform slice-like imaging of the agglomerates in the coagulation region, and the spatial number density and fractal dimension are calculated in real time through image processing and three-dimensional reconstruction. To ensure accurate data alignment, this step provides a unified time reference by setting a global clock source and using a time synchronization protocol (such as IEEE-1588PTP) to attach a high-precision global timestamp to each sensing event, thereby performing time alignment and feature extraction on multi-source heterogeneous data.
[0011] Step S2: Multi-scale Data Fusion and State Representation. This step constructs the multi-dimensional spatiotemporal data after spatiotemporal synchronization into a dynamically evolving graph model over time, and extracts a unified state vector that fuses spatiotemporal dependencies and represents the global state of the system. The graph model includes a set of nodes consisting of sensor node sets and physical region node sets, and a set of edges mathematically represented by an adjacency matrix. The construction of the edges follows the principles of physical proximity and fluid dynamics directionality, and can capture the time-delay effects of local physical coupling and airflow transport. The node feature matrix is updated in real time, containing scalar data and macroscopic statistical indicators of each node. To extract meaningful state representations from high-dimensional complex data, this step uses a Spatiotemporal Graph Convolutional Network (STGCN) to process the time series of the graph model. STGCN models the spatial topology and temporal evolution of node features of the graph by stacking spatiotemporal convolutional blocks, outputting a high-dimensional feature matrix. Its update operation formula is expressed as:
[0012] ;
[0013] In the formula, For the first The hidden state matrix of the layer; For the first The hidden state matrix of the layer; It is an adjacency matrix with self-loops added. It is the adjacency matrix constructed in the previous step. yes The identity matrix; yes The degree matrix; It is the first Trainable weight matrix of a layer space graph convolutional layer; It is the first Layer-time convolution kernel; Indicates a temporal convolution operation; This involves modifying the linear unit activation function. Finally, graph pooling or flattening operations are used to aggregate the high-dimensional feature matrix into a fixed-dimensional state vector. As a response A comprehensive high-dimensional mathematical representation of the state of the system at all times.
[0014] Step S3: Predictive Collaborative Decision Making. This step utilizes a deep reinforcement learning (DRL) agent based on the unified state vector. The system employs multi-objective reward and future prediction to generate a globally optimal cooperative action vector that balances cohesion effects and system energy consumption. The DRL agent utilizes an algorithm based on the Actor-Critic framework (e.g., DDPG or PPO) to achieve this. As a basis for decision-making, the optimal collaborative control strategy is generated. At any given moment, the cooperative action vector output by the DRL agent. It is a multidimensional vector containing multiple continuous control variables:
[0015] ;
[0016] In the formula: This is the amplitude of the driving voltage in the charged region. This parameter is the peak voltage of the driving power supply applied to the low-voltage multi-energy coupled strongly charged region. The goal of adjusting this parameter is to control the local electric field strength around the discharge electrode of the charged region microstructure, thereby affecting the concentration of ions produced. and the saturation charge of particulate matter , It is also affected by factors such as particle size, and its value range is ( to The threshold value is determined by the electrical safety threshold and discharge characteristics of the coagulation system. To guide the convergence of the strategy network, the DRL agent employs a multi-objective composite reward function. This is used to simultaneously balance the coagulation effect and system energy consumption.
[0017] ;
[0018] In the formula, It is a reward function with respect to fractal dimension; It is a reward function for increasing particle size; It is a condensed system at time step Total power consumption; , , These are the weighting coefficients for each reward:
[0019] First item It is a reward for condensed structures. This function utilizes the microscale sensing module in step S1. The fractal dimension obtained from analysis .because An increased value indicates a more compact and stable aggregate structure, which is beneficial for subsequent trapping. Designed as Monotonically increasing functions (e.g., To guide intelligent agents in learning can produce high Value control strategy;
[0020] Second item This is a particle size increase reward, used to quantify the macroscopic effect of particulate matter migrating to larger particle sizes. This function utilizes the particle size distribution function obtained in step S1. The specific formula for the calculation is as follows:
[0021] ;
[0022] In the formula, It is a preset critical particle size (e.g., 1.0 μm). In time step Particle size at time The number concentration of particulate matter, this function Calculations were made for particle sizes larger than The particle number fraction, and the increase of this function value directly indicates that the coagulation effect is enhanced, thus providing a clear positive reward signal for the DRL agent.
[0023] And penalties The optimization objective used to achieve low-carbon and energy-saving results is: total power consumption and the action vector output by the DRL agent in the previous stage. Direct association, its constituent expression formula is:
[0024] ;
[0025] In the formula, It is the instantaneous power in the charged region, which is the driving voltage of the charged region. and drive frequency The function; It is the instantaneous power of the condensation region, which is the peak voltage of the pulse in the condensation region. Pulse frequency and pulse duty cycle The function; It is the sum of the constant power consumption of each sensor and data acquisition device in step S1.
[0026] Furthermore, the DRL agent combines model predictive control (MPC) based on a dynamic system model. Predict future states. At each decision time step, the agent predicts future states within a predefined prediction time domain. Internally solve for an optimal action sequence Its optimization objective is to maximize the expected cumulative reward over the time domain:
[0027] ;
[0028] In the formula, It is a discount factor; In the predicted state Next action Instant rewards obtained at that time; Through dynamic models The future state obtained by iterative roll-forward prediction, i.e. ,and Furthermore, the agent only performs the first action in the sequence. The output is executed in step S4, which uses a rolling time-domain strategy to overcome the time delay of the physical system and achieve predictive control.
[0029] Step S4: Low-voltage coordinated drive execution. This step parses the coordinated action vector into physical control commands for charging and condensation and executes them to achieve coordinated regulation of multiple physics fields. Specifically, the amplitude of the driving voltage in the charging region... and frequency It is fed to the driving power source in the charged region to control the intensity, density, and uniformity of the micro-discharge in order to maximize the ion concentration in the charged region. Thus minimizing the particulate charging time constant. This allows particulate matter to remain in the environment for a limited time. Maximize the charge within Charging time constant With ion concentration The relationship is:
[0030] ;
[0031] In the formula, It is the vacuum permittivity; It is the average ion concentration within the charged region. It is the amount of elementary charge; It is ion mobility; and, , and These are all physical constants or inherent properties of flue gas; therefore, minimizing... The only way is to maximize the average ion concentration. And ion concentration These micro-discharges (e.g., dielectric barrier discharges or surface discharges) occur within the low-voltage, multi-energy coupled, strongly charged region. The intensity, density, and uniformity of these micro-discharges are directly controlled by the voltage amplitude output from the power supply driving the charged region. and drive frequency Synergistic regulation aims to maximize the total condensation kernel function of particulate matter. The formula is:
[0032] ;
[0033] In the formula, It is a Brownian condensate nucleus. It is a Coulomb condensation nucleus: ;in It is a turbulent condensation nucleus; It is a turbulent condensation nucleus. ;
[0034] It is the Cunningham slip correction factor. It's the viscosity of the gas. It refers to the kinematic viscosity of the gas. By using the action vector uniformly output by the DRL agent, deep physical coupling and synergistic optimization of the charging and condensation processes are achieved to reach Pareto optimality in terms of both efficiency and energy consumption.
[0035] Step S5: Closed-Loop Iteration and Adaptive Optimization. This step continuously collects multi-dimensional spatiotemporal data and uses the accumulated data to continuously update the deep reinforcement learning agent, achieving adaptive adjustment across all operating conditions. In each running loop, the system generates and records experience tuples. The time-state vector. The DRL agent uses these experience tuples for online or offline training, continuously updating the internal parameters of the policy network and value network, enabling it to adapt to non-ideal characteristics and unmodeled dynamics in actual operation, and achieving robust and continuously evolving adaptive control for all operating conditions such as flue gas flow fluctuations, sudden changes in inlet particulate matter concentration, and humidity changes.
[0036] A second aspect of the present invention provides a low-pressure charged coagulation device for low-carbon submicron particles, characterized in that it comprises:
[0037] Sensing module: This module is used to acquire multidimensional spatiotemporal data characterizing the environmental field, microstructure, and mesoscopic distribution of the charged coalescence process in real time, and to perform spatiotemporal synchronization processing on the multidimensional spatiotemporal data. This module corresponds to and implements step S1 in the above method.
[0038] State representation module: This module is used to construct a dynamically evolving graphical model from the spatiotemporally synchronized multidimensional spatiotemporal data and extract a unified state vector that integrates spatiotemporal dependencies and represents the global state of the system. This module corresponds to and implements step S2 in the above method.
[0039] Decision-making module: Based on the unified state vector, multi-objective rewards, and future predictions, a deep reinforcement learning agent generates a globally optimal cooperative action vector that balances cohesion effects and system energy consumption. This module corresponds to and implements step S3 in the above method.
[0040] The drive execution module is used to parse the cooperative action vector into physical control commands for charging and condensation, and then execute them to achieve coordinated control of multiple physics fields. This module corresponds to and implements step S4 in the above method.
[0041] Optimization module: Based on the data acquired by the perception module after coordinated regulation, and utilizing accumulated multi-dimensional spatiotemporal data, the deep reinforcement learning agent is continuously updated to achieve adaptive adjustment under all operating conditions. This module corresponds to and implements step S5 in the above method.
[0042] This invention provides a method and apparatus for low-voltage electrocoagulation of low-carbon submicron particles. It has the following beneficial effects:
[0043] 1. This invention employs low-voltage multi-energy coupling strong charging and low-voltage pulse adaptive coagulation, combined with a deep reinforcement learning agent using a multi-objective reward function with total power consumption as the penalty term. This allows the system to autonomously find low-energy control strategies while ensuring the coagulation effect of submicron particles, thereby reducing the overall operating voltage and power consumption of the coagulation system, and reducing operating costs and indirect carbon emissions.
[0044] 2. This invention uses multidimensional spatiotemporal data and establishes a graphical model that dynamically evolves over time. This enables the coagulation system to respond to and predict changes in operating conditions in real time, dynamically adjust the charge and coagulation parameters, and continuously train the deep reinforcement learning agent through closed-loop iteration and empirical tuples. This ensures that the invention can maintain high efficiency and low energy consumption under various complex operating conditions.
[0045] 3. This invention acquires the fractal dimension and particle size distribution function of particles in real time, and merges the fractal dimension reward term and particle size increase reward term in the reward function. Therefore, it enables the deep reinforcement learning agent to make fine and forward-looking adjustments, thereby improving the merging effect and subsequent capture efficiency. Attached Figure Description
[0046] Figure 1 This is a schematic diagram of the method flow of the present invention;
[0047] Figure 2 This is a schematic diagram of the multidimensional spatiotemporal data synchronization processing of the present invention;
[0048] Figure 3 This is a schematic diagram illustrating the structure of the multi-objective composite reward function of the present invention;
[0049] Figure 4 This is a schematic diagram of the device architecture of the present invention. Detailed Implementation
[0050] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] To better understand the present invention, the above content will be described in detail below with reference to specific embodiments.
[0052] Please see the appendix Figure 1 - Appendix Figure 3 This invention provides a low-pressure charged coagulation method for low-carbon submicron particles, the method comprising the following steps:
[0053] The first step, S1, involves using multiple sensors to acquire, in real-time and comprehensively, various physical parameters characterizing the state of the charged coalescence process. Specifically, to obtain the environmental field parameters within the charged coalescence region, this method deploys a sensor array at key locations in the charged and coalescence regions (e.g., near-wall region, mains region, and near the electrode surface) to acquire information, including:
[0054] Used for measuring local temperature fields Thermocouple or thermistor sensor; used to measure local humidity field Capacitive or resistive humidity sensors; used to measure local airflow velocity fields. Hot-wire or microelectromechanical system (MEMS) flow velocity sensors; and sensors for measuring local electric field intensity distribution. A miniature electric field probe. Among them, For the sensor's spatial coordinates, For timestamps.
[0055] To monitor the agglomeration effect of particulate matter online, this method installs a laser particle size analyzer at the downstream outlet of the agglomeration zone. Specifically, the laser particle size analyzer, based on Mie scattering theory or Fraunhofer diffraction principle, performs non-invasive measurements of particulate matter flowing through the measurement area, thereby acquiring the particle size distribution function in real time. ,in Particle size, For timestamps. This function represents the time stamp in time. Time, different particle sizes The corresponding particulate matter number concentration.
[0056] To capture the microscopic dynamic characteristics of particle agglomerates during coagulation, this method employs multiple high-speed cameras combined with a sheet-like laser light source to generate a thin illumination plane within the coagulation zone. This plane is used for slice-like imaging of agglomerates passing through it. The acquired image sequence first requires image segmentation algorithms (such as thresholding or edge detection) to identify the two-dimensional projection of the agglomerates from the background. Subsequently, multi-view geometric or motion parallax algorithms are used to reconstruct the three-dimensional morphology of the agglomerates. Based on the reconstructed three-dimensional morphology, key microscopic parameters, including spatial number density, are then calculated. Through statistics Time, specific three-dimensional volume The fractal dimension is obtained by measuring the number of aggregates within the fractal. The formula used to characterize the density of aggregates is based on the relationship between the mass and radius of the aggregates, as shown below:
[0057] ;
[0058] In the formula, It is the number of primary particles in the aggregate; It is the radius of gyration of the aggregate, calculated from the reconstructed 3D point cloud data; It is the average particle size of the primary particles, which can be input as a priori parameter or obtained from the initial flue gas measurement; It is the fractal preconditioner, a constant whose value can be obtained through calibration experiments or literature. Therefore, by fitting... and logarithmic relation (i.e.) by The slope parameter is solved in real time.
[0059] In one embodiment, the data collected above has the characteristics of being multi-source and heterogeneous. Specifically, the data sampling rate, data dimension, and physical delay of each sensor are not the same. For example, the sampling rate of a high-speed machine vision system (e.g., 100Hz) is much higher than that of a laser particle size analyzer (e.g., 1Hz).
[0060] To ensure the accuracy of data fusion and state representation, it is necessary to perform spatiotemporal synchronization processing on these data streams from different sources and at different rates.
[0061] Specifically, this method provides a unified time reference by setting a high-precision global clock source. All sensor data acquisition interfaces, the laser particle size analyzer's control unit, and the high-speed machine vision system's image acquisition card maintain high-precision synchronization with this global clock source via a time synchronization protocol (e.g., IEEE-1588 Precision Time Protocol, PTP). Therefore, under this unified time reference, any sensing event (e.g., temperature measurement, particle size spectrum scanning, exposure of a frame of image) is appended with a high-precision global timestamp by its corresponding acquisition hardware at the time of its occurrence.
[0062] The data stream with global timestamps then needs to be aligned. The specific steps are as follows: First, define a unified data fusion time window. The length of this window is set to... This is a key parameter, and its value (e.g., 0.5 seconds) needs to be set according to the dynamic response characteristics of the system and the required control precision.
[0063] Subsequently, with The data stream is sliced and aligned using a step size. For a single time window ( High-frequency sensing data that generates multiple readings within the range (e.g., multiple calculated readings) (value), and then a time-weighted average or arithmetic average method is used to calculate a characteristic value that can represent the window period (e.g., ).
[0064] For low-frequency sensing data that is sampled less than once within a single time window (e.g., laser particle size analyzer data), ... Nearest neighbor interpolation or zero-order hold method can be used to assign the most recent valid measurement value to the current time window. These time-aligned and feature-extracted (e.g., averaging, interpolation) data will be used at each discrete time node. A complete, time-synchronized data snapshot is formed and used as the node feature matrix in step S2. The foundation for its construction.
[0065] After acquiring the time-aligned data, the present invention proceeds to step S2: Specifically, in order to effectively capture the complex spatial coupling relationships within the charged coagulation system and the time delay effect brought about by airflow transport, the present invention abstracts the geometric structure and sensor network of the coagulation system into a graphical model that dynamically evolves over time. .
[0066] The graphical model The set of nodes in Depend on The system consists of several nodes, specifically divided into a set of sensor nodes and a set of physical region nodes. Each sensor node corresponds one-to-one with the sensors deployed in step S1, representing a specific monitoring point in the condensation system, such as a temperature sensor at the inlet of the charged region or an electric field probe on the electrode surface. Physical region nodes correspond to specific functional space volumes within the condensation system, such as the core volume of the charged region, the functional volume of the condensation region, and the volume of the outlet monitoring area, used to accommodate macroscopic or microscopic statistical indicators within that region.
[0067] The set of edges in a graphical model Through the adjacency matrix Perform mathematical representation, where matrix elements Describes the nodes With nodes The strength of the physical association between them is determined by the principle that the construction of edges follows the principles of physical proximity and fluid dynamic directionality.
[0068] Regarding physical proximity, if the Euclidean distance between two nodes in the three-dimensional space of the condensed system is less than a preset interaction threshold, then it is assumed that there is a direct spatial coupling between them. It is assigned a non-zero value. This connection method can capture the influence of local thermal diffusion or electric field edge effects on adjacent regions.
[0069] Regarding the directionality of fluid dynamics, considering the physical fact that flue gas flows from upstream to downstream, the adjacency matrix... It is constructed as a directed graph matrix or an asymmetric matrix. If the nodes Located at node If the airflow is upstream and there is a fluid connection path, then establish a connection from... point to The directed edges represent the migration process of matter and energy with the airflow. This flow-direction-based graph construction allows subsequent graph convolution operations to simulate the transport lag characteristics of particulate matter in the reactor, that is, the upstream charge state will affect the downstream coagulation state after a certain time delay.
[0070] Node feature matrix The real-time data collected and synchronized in step S1 is characterized, where To unify the feature dimensions, that is: to unify the scalar data collected by each sensor node (such as... , , , This is mapped to the feature vector components of that node. For physical region nodes, such as condensation region nodes, the fractal dimension resolved by the visual system is... and spatial number density Fill in its eigenvector; for outlet region nodes, load the particle size distribution function. The values within a preset discrete particle size range, for example, are used as eigenvector components through histogram statistics or specific moments (such as number-average particle size and weight-average particle size).
[0071] To address the differences in data dimensions among different types of nodes, the system employs feature embedding or zero-padding techniques to map the original data of all nodes to the same dimension. This forms a standardized feature matrix. This matrix Together with the adjacency matrix It has been completely preserved. The spatial topology and physical state information at each time step provide a standardized input data structure for subsequent deep feature extraction using spatiotemporal graph convolutional networks.
[0072] Furthermore, in step S2, the graphical model evolves dynamically over time. Subsequently, in order to extract meaningful system state representations from this high-dimensional and complex data structure, this embodiment employs a Spatiotemporal Graph Convolutional Network (STGCN) to process the time series of the graph data. Specifically, the core of STGCN lies in its ability to simultaneously model the spatial topology and temporal evolution of node features of the graph.
[0073] STGCN is implemented by stacking spatiotemporal convolutional blocks. In the... In each spatiotemporal convolutional block, its hidden state The formula for the update operation is expressed as:
[0074] ;
[0075] In the formula, For the first The hidden state matrix of the layer; For the first The hidden state matrix of the layer; It is an adjacency matrix with self-loops added. It is the adjacency matrix constructed in the previous step. yes The identity matrix; yes The degree matrix; It is the first Trainable weight matrix of a layer space graph convolutional layer; It is the first Layer-time convolution kernel; Indicates a temporal convolution operation; It is a modified linear unit activation function.
[0076] And the formula inside Part, that is It performs a spatial graph convolution operation, the core of which is to utilize a normalized adjacency matrix. As a propagation operator, for Layer node features Perform weighted aggregation.
[0077] And due to the adjacency matrix The system already encodes physical adjacency relationships (such as the distance between sensors) and fluid dynamic relationships (such as upstream and downstream airflow). This spatial convolution operation can capture the complex (non-Euclidean) spatial dependencies within the condensing system. For example, it can aggregate the electric field intensity distribution of upstream charged nodes. Humidity distribution with downstream condensation zone nodes This characterizes the spatial coupling relationship between upstream charge efficiency and downstream coagulation environment.
[0078] And outside the formula The operation involves performing temporal convolution, which is applied to the temporal series output of spatial convolution (i.e., node features that have incorporated spatial neighborhood information). This is achieved by sliding the convolution kernel along the temporal dimension. This operation can capture the dynamic trends of the evolution of the characteristics of each node over time.
[0079] For example, this temporal convolution can identify the microscopic parameters of nodes in the condensation region. The fractal dimension shows an increasing, decreasing, or oscillating trend over past time steps. After deep feature extraction through multiple layers of spatiotemporal convolutional blocks, the STGCN model can output a high-dimensional feature matrix. Each row of the matrix represents a node and incorporates information about the (spatial) neighborhood and (temporal) historical evolution of that node.
[0080] To generate a unified feature vector that represents the global state of the system, this embodiment employs graph pooling operations (e.g., global average pooling or max pooling) or flattening operations to... Aggregates into a fixed-dimensional state vector The state vector Yes A comprehensive high-dimensional mathematical representation of the state of the system at all times.
[0081] Subsequently, the invention proceeds to step S3, which involves employing a deep reinforcement learning (DRL) agent with the global state vector generated in S2. This serves as the basis for decision-making, used to generate the optimal collaborative control strategy. The agent operates within a Markov Decision Process (MDP) framework, and its state space... From the input state vector definition.
[0082] Specifically, in this embodiment, the agent may employ an algorithm based on the Actor-Critic framework, such as the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm. The Actor-Critic framework may include a policy network (Actor) and a value network (Critic).
[0083] The policy network (Actor) is responsible for receiving the state. As input, it outputs an optimal action policy. The value network (Critic) also receives state vectors. As input, it is used to evaluate the merits of the current policy (e.g., outputting the state value function). or action value function This guides the parameter updates of the policy network. For the implementation of such DRL algorithms, those skilled in the art can construct them using well-known neural networks (such as multilayer perceptrons). The training and inference processes are well-known techniques in this field and will not be elaborated upon here.
[0084] Furthermore, this DRL agent in Cooperative action vector output at each time step It is a multidimensional vector containing multiple continuous control variables, defined as follows:
[0085] ;
[0086] In the formula: This is the amplitude of the driving voltage in the charged region, which is the peak voltage of the driving power supply applied to the low-voltage multi-energy coupled strongly charged region. The goal of adjusting this parameter is to control the local electric field strength around the discharge electrode of the charged region microstructure, thereby affecting the concentration of ions produced. and the saturation charge of particulate matter , It is also affected by factors such as particle size, and its value range is ( to The electrical safety threshold and discharge characteristics of the condensation system are used to determine the condensation system's electrical safety threshold.
[0087] This is the driving frequency of the charged region. This parameter is the frequency applied to the driving power supply of the charged region. The goal of adjusting this parameter is to control the morphology of the charged region discharge (e.g., the uniformity of micro-discharge in dielectric barrier discharge) and the ion generation efficiency. Its value range ( to The characteristics of the power supply system determine the specific characteristics of the power supply system.
[0088] This is the peak voltage of the pulse in the condensation region. This parameter is the peak voltage of the high-frequency pulse power supply applied to the low-voltage pulse adaptive condensation region. The goal of adjusting this parameter is to control the maximum intensity of the time-varying electric field in the condensation region, which directly affects the Coulomb attraction experienced by charged particles and the intensity of fluid disturbances induced by the electric field. Its value range ( to The breakdown voltage between the electrodes in the coagulation zone is determined.
[0089] This is the pulse frequency in the condensation zone. This parameter is the pulse repetition frequency of the pulsed power supply applied to the condensation zone. The goal of adjusting this parameter is to control the rate of electric field alternation, which affects the oscillatory behavior of particles in the electric field and the energy dissipation rate of induced microturbulence. Its value range is ( to Determined based on the hydrodynamic properties of particulate matter, such as relaxation time;
[0090] This is the pulse duty cycle in the coagulation region. This parameter represents the pulse duty cycle applied to the pulsed power supply in the coagulation region, i.e., the proportion of the pulse high-level time within one cycle. The goal of adjusting this parameter is to finely regulate the net time the electric field acts on the particles within a single cycle, as well as the average power injected into the coagulation system, to synergistically optimize the relative contributions of electrostatic coagulation and turbulent coagulation. Its value range ( to It is usually between 0 and 1.
[0091] Furthermore, in step S3, in order to guide the policy network (Actor) of the DRL agent to converge toward the optimal goal, the value network (Critic) needs an accurate scalar reward signal. To evaluate the agent in the state vector Next action The advantages and disadvantages of the reward function are determined by its multi-objective composition, which aims to balance the cohesion effect with system energy consumption. At time step The specific formula is:
[0092] ;
[0093] In the formula, It is a reward function with respect to fractal dimension; It is a reward function for increasing particle size; It is a condensed system at time step Total power consumption; , , These are the weighting coefficients for each reward;
[0094] reward function The first two items, namely and This constitutes a quantified reward for the cohesion effect.
[0095] First item It is a reward for condensed structures. This function utilizes the microscale sensing module in step S1. The fractal dimension obtained from analysis .because An increased value indicates a more compact and stable aggregate structure, which is beneficial for subsequent trapping. Designed as Monotonically increasing functions (e.g., To guide intelligent agents in learning can produce high Value control strategy;
[0096] Second item This is a particle size increase reward, used to quantify the macroscopic effect of particulate matter migrating to larger particle sizes. This function utilizes the particle size distribution function obtained in step S1. The specific formula for the calculation is as follows:
[0097] ;
[0098] In the formula, It is a preset critical particle size (e.g., 1.0 μm). In time step Particle size at time The number concentration of particulate matter, this function Calculations were made for particle sizes larger than The particle number fraction, and the increase of this function value directly indicates that the coagulation effect is enhanced, thus providing a clear positive reward signal for the DRL agent.
[0099] And penalties The optimization objective used to achieve low-carbon and energy-saving results is: total power consumption and the action vector output by the DRL agent in the previous stage. Direct association, its constituent expression formula is:
[0100] ;
[0101] In the formula, It is the instantaneous power in the charged region, which is the driving voltage of the charged region. and drive frequency The function; It is the instantaneous power of the condensation region, which is the peak voltage of the pulse in the condensation region. Pulse frequency and pulse duty cycle The function; It is the sum of the constant power consumption of each sensor and data acquisition device in step S1.
[0102] Therefore, by As a negative reward (penalty), the DRL agent is guided during training to find a way to achieve high [reward / penalty]. and At the same time, consume as little as possible The Pareto optimal strategy;
[0103] Weighting coefficient These coefficients play a balancing role in the multi-objective optimization process. They are positive real numbers and are set before training the DRL model based on specific process requirements and cost considerations. For example, when prioritizing ensuring that the outlet particulate matter meets standards, the coefficients are increased. and The value; when it is necessary to prioritize reducing operating costs, increase it. The value of .
[0104] Furthermore, in step S3, the optimization objective of the DRL agent is to maximize the defined multi-objective reward function. The long-term cumulative expectation. However, the physical process of charged coalescence exhibits significant nonlinear, inertial, and time-delay characteristics, that is, in Actions performed at all times (e.g., adjust) ), which affects the state of the coagulation system (e.g., or The full impact will require a period of evolution to materialize. The moment is fully observed in step S1.
[0105] Therefore, if the DRL agent is based only on the current state Making reactive decisions inherently leads to oscillations and suboptimal results in control strategies. To address this issue, this embodiment incorporates model predictive control (MPC) into the decision algorithm in step S3 to achieve predictive regulation.
[0106] Specifically, the realization of this predictive regulation relies on a dynamic system model. The model Able to determine the current state Actions output by the DRL agent Predict the (desired) state of the condensed system at the next moment. ,Right now: And this dynamic model It can be obtained by training a standalone neural network with time-series data processing capabilities (e.g., a recurrent neural network RNN or a long short-term memory network LSTM) based on historical running data, or by being learned online as part of a model-driven DRL agent.
[0107] At each decision time step The DRL agent first sets a fixed prediction time domain. (For example, (Time step).
[0108] Subsequently, the optimization objective of the DRL agent (i.e., the policy network actor) is no longer to output a single optimal action. Instead, it solves a problem in Optimal action sequence in the time domain .
[0109] Therefore, the objective function of this optimization is to maximize the value of the target value in ... Expected cumulative reward within the time domain Its definition is as follows:
[0110] ;
[0111] In the formula, It is a discount factor; In the predicted state Next action Instant rewards obtained at that time; Through dynamic models The future state obtained by iterative roll-forward prediction, i.e. ,and .
[0112] When the agent finds the optimal solution through the optimization algorithm... Maximize the optimal action sequence Subsequently, this invention follows the MPC rolling time-domain strategy: only the first action in the sequence is performed. The output is sent to the drive system execution in step S4. And in the next time step... Obtain the new real state from S2 (This status already includes execution) (the actual effect), and completely repeat the above. The optimization process in the time domain involves replanning the action sequence for the next time period.
[0113] In this way, DRL agents can Decisions at any moment The impact of this action on the future has been fully considered. All states within a step ( to The invention also incorporates the effects of rewards and incentives. This allows the invention to predict changes in operating conditions, enabling predictive control and effectively overcoming the time lag of the physical system, thus avoiding control overshoot and oscillation.
[0114] In step S4, the optimal action sequence determined by the DRL agent in step S3 is first executed. The first action in Then from this action vector In the process, the charge control command is extracted. and The command is then sent to the charged area drive power supply in the power supply and drive system. This power supply (e.g., a programmable high-frequency AC power supply or pulse power supply) operates according to the voltage amplitude set in the command. and drive frequency Electrical excitation is applied to the electrode array in the low-voltage, multi-energy coupled, strongly charged region of the coagulation system. The physical principle behind this modulation lies in the fact that particulate matter has a limited residence time. The amount of charge that can be obtained inside Following the field-diffusion charging theory, the formula is:
[0115] ;
[0116] In the formula, The residence time of particulate matter The amount of charge carried at the end; It is the average residence time of particulate matter as it passes through the charged region, which is mainly determined by the flue gas velocity. The geometry of the charged region determines the condition and can be considered a constraint under specific operating conditions. It is the saturated charge of particulate matter, and its value depends on factors such as particle size and applied electric field strength. It is the charging time constant.
[0117] Therefore, it can be concluded that in and Under relatively limited conditions (especially at low pressure and high flow rates), to achieve Maximizing (i.e., strong charging) hinges on ensuring that the charging time constant is minimized. minimize, The smaller the particle size, the faster it charges and the faster it can approach saturation charge in a short time.
[0118] Charge time constant (When the electric field is dominant) it is mainly determined by the ion concentration in the charged region, and the relationship is expressed by the following formula:
[0119] ;
[0120] In the formula, It is the vacuum permittivity; It is the average ion concentration within the charged region. It is the amount of elementary charge; It is ion mobility; and, , and These are all physical constants or inherent properties of flue gas; therefore, minimizing... The only way is to maximize the average ion concentration. And ion concentration These micro-discharges (e.g., dielectric barrier discharges or surface discharges) occur within the low-voltage, multi-energy coupled, strongly charged region. The intensity, density, and uniformity of these micro-discharges are directly controlled by the voltage amplitude output from the power supply driving the charged region. and drive frequency .
[0121] Therefore, the essential regulation logic of the DRL agent in step S3 can be determined as: the DRL algorithm (in a multi-objective reward function) Guided by [the system / mechanism], it learns and predicts, and outputs an optimal combination of parameters in real time. The goal of this combination is to drive the charged region under current operating conditions (by...). Under characterization and low-pressure constraints, the maximum effective ion concentration is generated. Thus achieving Minimize the particulate matter, making it more efficient in short-term applications. To maximize the charge within .
[0122] Following step S4, which involves low-voltage multi-energy coupling strong charge regulation of the charged region, this invention simultaneously analyzes the optimal action sequence. Coagulation control instructions in The instruction is then sent to the drive power supply system of the low-voltage pulse adaptive condensation zone.
[0123] The driving power supply system can be specifically implemented as a high-frequency, high-voltage pulse power supply, whose output terminal is connected to the electrode structure of the condensation region. This electrode structure can be an alternating electric field plate array or a multi-level grid electrode. The power supply system modulates the output waveform according to the received command: pulse peak voltage... The peak value of the electric field intensity applied between the electrodes in the condensation zone is determined; the pulse frequency Set the rate of electric field polarity reversal or pulse excitation; pulse duty cycle. This controls the duration of the high-pressure effect within a single cycle, thereby regulating the energy density injected into the coagulation zone. Furthermore, the core objective of this step is to enhance the collisional aggregation probability of charged submicron particles within the coagulation zone. From an aerosol mechanics perspective, the particle coagulation rate depends on the total coagulation kernel function. This embodiment aims to maximize the total cohesion kernel function through dynamic optimization by the DRL agent, and its formula is expressed as:
[0124] ;
[0125] In the formula, The particle size is and The total collision frequency coefficient between two particles; It is a Brownian condensation nucleus, which is mainly affected by temperature and can be regarded as a reference quantity in this control method; It is a Coulomb condensate nucleus, which characterizes collisions under the influence of electrostatic attraction; It is a turbulent condensation nucleus, characterizing the collision under the action of fluid turbulent shear and inertia;
[0126] The DRL agent adjusts the voltage parameters in the action vector. Direct intervention in Coulomb condensation nuclei The size. In a system of charged particles, The formula for expressing the contribution is:
[0127] ;
[0128] In the formula, , These are the charges of the two particles (determined by the preceding charging steps); It is the viscosity of the gas; It is the vacuum permittivity; It is the Cunningham slip correction factor; It is the effective electric field strength within the condensation region. Because... With applied voltage There is a positive correlation ( , (for interpole spacing), the agent improves It can generate stronger Coulomb attraction between oppositely charged particles or induce stronger dipole moment between like-charged particles, thereby significantly improving... ; For particle size and Coulombic condensation nuclei between two particles Size.
[0129] However, simply increasing the voltage is limited by the breakdown threshold and cannot induce fluid disturbances. Therefore, the DRL agent needs to simultaneously optimize the frequency. and duty cycle Microscopic turbulence is excited by utilizing electrohydrodynamics (EHD) effects, thereby enhancing turbulent condensation nuclei. In a turbulent field induced by a pulsed electric field, It can be represented by the following formula:
[0130] ;
[0131] In the formula, For particle size and Turbulent condensation nuclei between two particles Size; It is the kinematic viscosity of the gas; It is the turbulent energy dissipation rate. In this embodiment, It is not generated by an external fan, but by ion wind induced by a pulsed electric field and perturbations of dielectric force. Pulse parameters and It determines the frequency and duration of electric field energy injection into the fluid, directly controlling... Size.
[0132] In summary, DRL regulation is essentially a multi-physics collaborative optimization process: dynamically weighing factors within the action space based on the currently sensed particle concentration and fractal dimension. On one hand, by setting appropriate... Maintain sufficient electrostatic attraction to ensure On the one hand, it contributes; on the other hand, it adjusts and Maximize induced turbulence without causing electric field breakdown. thereby improving The contribution of this invention. Ultimately, this invention... and Establishing a positive synergistic effect between them makes the total condensation kernel function To achieve the theoretical maximum value under current operating conditions, and realize efficient, low-pressure, adaptive coagulation of submicron particles.
[0133] Furthermore, during the execution of step S4, the aforementioned low-voltage multi-energy coupled strong charge regulation and low-voltage pulse adaptive condensation regulation are not two independent entities, but rather synergistic, specifically manifested in...
[0134] The policy network (Actor) of the DRL agent at any decision time It is the global state vector generated based on step S2. (Having integrated all sensing information from the charged region and the condensation region), it outputs a unified, complete motion vector containing all control parameters. .
[0135] This decision-making mechanism utilizes the deep physical coupling between the charging and coalescence processes. Specifically, it is controlled by the charging region. Determined particulate matter charge It is a Coulomb condensation nucleus in the downstream condensation region. The key input parameters. The magnitude of the electrostatic coagulation effect directly determines the strength of the effect.
[0136] Simultaneously, control commands from the condensation zone. The determined pulsed electric field characteristics primarily regulate the turbulent condensation nuclei. (through influence) ).
[0137] The optimization objective of the DRL agent is to maximize the defined global, multi-objective cumulative reward. The reward It is due to the final export coagulation effect (by and (characterization) and total energy consumption That was decided.
[0138] Therefore, during training and decision-making, the agent learns a cooperative strategy of charging and cohesion, and does not simply pursue... Maximizing (which may lead to) Too high, pull down It is not simply about pursuing Maximize.
[0139] For example, when the system state This indicates that when the inlet particle concentration is low, the agent may determine that only a moderate amount of charge is needed. (Corresponding to a set of energy-saving) ) and moderate turbulence (Corresponding to a set of energy-saving) By combining these methods, export needs can be met. The goal.
[0140] Conversely, when This indicates harsh operating conditions (such as high flow rates leading to...) In a very short time, the agent may simultaneously output high-intensity charge commands (to achieve high charge in a short time). Internal rapid improvement ) and high-intensity coagulation pulse commands (to utilize and Strengthen together This enables robust control over harsh working conditions.
[0141] Therefore, global rewards can be achieved by uniformly outputting charge parameters and condensation parameters through DRL agents. The maximization of decision-making methods constitutes the coordinated control of charging and coalescence in the execution phase of step S4, ensuring that the present invention can operate on the Pareto optimal frontier of effect and energy consumption.
[0142] In one embodiment, after step S4, the invention proceeds to step S5: closed-loop iteration and adaptive optimization. Due to the changes in the charge state, particle size distribution, and agglomerate morphology of the particles after strong charging and adaptive coagulation control, the flue gas, with these altered states, is propelled by the airflow at the next moment... The system reaches the downstream sensing area. Therefore, after completing step S4, the control flow automatically returns to step S1.
[0143] Specifically, in At time S1, the state of the condensed system after being regulated by S4 will be re-acquired, thereby obtaining a new set of real-time data reflecting the control effect of the previous moment, such as a new fractal dimension. New particle size distribution and new environmental field parameters , wait).
[0144] This new set of sensory data will then proceed to step S2 to construct a new spatiotemporal map. and generate new state vectors .Should The data is then sent to S3, where the DRL agent, based on this new state and the prediction model, makes the decision for S4. Therefore, this invention can continuously cycle at each time step, forming a closed-loop operating mechanism for the system, ensuring the real-time responsiveness of the control strategy to changes in operating conditions.
[0145] Furthermore, the aforementioned closed-loop operation mechanism not only achieves real-time control, but more importantly, it also provides a continuous data flow for the adaptive updates of the DRL agent. In each operating cycle of the system (i.e., each time step)... Each of these generated and recorded a complete data experience tuple. .
[0146] In this tuple: It is step S2 in The state vector generated at each moment for decision-making; The DRL agent in step S3 is in The action vector output at each time step and executed by S4; yes The actions are constantly sensed and identified by S1 and S2. The new state vector after; It is based on Real-time sensing data (such as) and and combined energy consumption The calculated instant reward.
[0147] At this point, DRL intelligence can utilize these experience tuples. Training can be conducted online or offline to continuously update its policy network. And the internal parameters of the value network (Critic) (e.g., the weights of the neural network).
[0148] Furthermore, the DRL agent in this invention can be pre-trained offline in a simulation environment based on the dynamic system model M in step S3. This allows the agent to quickly learn the basic control strategy for the charge coalescence process before actual deployment, shortening the learning cycle of the actual system and improving safety. When the agent is deployed to the physical system, it can also be fine-tuned online using real-time acquired experience tuples, enabling its strategy to continuously adapt to non-ideal characteristics and unmodeled dynamics in actual operation, thereby achieving robust and continuously evolving full-condition adaptive behavior.
[0149] For example, in offline training mode, data is accumulated during operation, and when computing resources are idle, small batches of empirical tuples are randomly sampled from the buffer to update the network by performing optimization steps of an Actor-Critic algorithm (such as PPO or DDPG).
[0150] Through this continuous accumulation of experience and the adaptive update mechanism of self-training, the strategy of the DRL agent (Actor) It can continuously evolve, autonomously learning from long-term operational data the complex dynamic characteristics of the charge coalescence process and the optimal mapping relationship between control actions and long-term cumulative rewards.
[0151] Ultimately, the DRL model's adaptive update mechanism enables the control system to automatically adapt to various operating conditions without manual intervention or reprogramming, such as fluctuations in flue gas flow, sudden changes in inlet particulate matter concentration, and humidity changes. Through autonomous learning, it continuously optimizes its control strategy, always approaching the globally optimal operating point defined by the multi-objective reward function, achieving intelligent control that is adaptive under all operating conditions.
[0152] Please see the appendix Figure 4 This invention provides a low-pressure charged coagulation device for low-carbon submicron particles, the device comprising:
[0153] The sensing module is used to acquire multi-dimensional spatiotemporal data in real time, including environmental field, particulate matter size spectrum, agglomerate fractal dimension, and spatial number density, through sensor arrays (such as thermocouples, humidity sensors, flow rate sensors, and electric field probes) deployed in the charged and agglomerated regions, as well as downstream laser particle size spectrometers and high-speed camera systems. This module also includes a clock synchronization unit that uses a global clock source and time synchronization protocols (such as PTP) to append high-precision global timestamps to all acquired data and ensures temporal synchronization of data from different sources through data alignment processing.
[0154] The state representation module is used to construct a dynamically evolving graphical model from the spatiotemporal synchronization data processed by the sensing module. This model includes nodes representing sensors and physical regions, as well as edges representing physical proximity and fluid dynamic directionality. Furthermore, a spatiotemporal graph convolutional network (STGCN) is used to process the time series of this graphical model to capture complex spatiotemporal dependencies. Through graph pooling or flattening operations, high-dimensional features are aggregated into a unified state vector representing the global state of the system.
[0155] The decision-making module internally employs a deep reinforcement learning agent (e.g., based on the Actor-Critic framework) and takes a unified state vector output from the state representation module as input, combined with model predictive control (MPC) to predict future states. Simultaneously, it optimizes based on a multi-objective composite reward function (used to balance cohesion effects, such as fractal dimension and particle size increase, with total system power consumption), ultimately generating a globally optimal cooperative action vector at each decision time step. This vector contains multiple continuous control variables (such as voltage amplitude, frequency, and duty cycle) in the charged and cohesive regions.
[0156] The drive execution module is used to parse the optimal cooperative action vector output by the decision module (specifically, the first action in the model predictive control sequence) and convert it into physical control commands for the driving power supply in the charged region and the pulse power supply in the condensation region. Simultaneously, it achieves strong charging by coordinating the voltage and frequency of the charged region, and adaptively regulates the pulse peak voltage, frequency, and duty cycle of the condensation region to synergistically enhance Coulomb condensation and turbulent condensation effects, thus achieving coordinated control of multiple physics fields.
[0157] The optimization module is used to achieve closed-loop iteration and adaptive optimization. Simultaneously, it continuously utilizes new data collected by the perception module after control execution to construct experience tuples containing the current state, executed actions, obtained rewards, and the next state. This module uses these accumulated experience tuples to continuously update the internal parameters of the policy network and value network of the deep reinforcement learning agent in the decision-making module through online or offline training, enabling the entire ensemble system to learn autonomously and adapt to changes in all operating conditions.
Claims
1. A method for low-voltage electrocoagulation of low-carbon submicron particles, characterized in that, Includes the following steps: Step S1: Multi-scale spatiotemporal synchronous sensing is used to acquire multi-dimensional spatiotemporal data characterizing the environmental field, micromorphology, and meso-scale distribution of the charge coagulation process in real time, and to perform spatiotemporal synchronous processing on the multi-dimensional spatiotemporal data. Step S2: Multi-scale data fusion and state representation, used to construct a graph model that dynamically evolves over time from the multi-dimensional spatiotemporal data after spatiotemporal synchronization processing, and to extract a unified state vector that fuses spatiotemporal dependencies and represents the global state of the system. Step S3: Predictive collaborative decision-making, through a deep reinforcement learning agent based on the unified state vector, multi-objective rewards and future predictions, to generate a globally optimal collaborative action vector that balances cohesion effect and system energy consumption; Step S4: Low-voltage collaborative drive execution, used to parse the collaborative action vector into physical control commands for charging and condensation and execute them, so as to realize the collaborative regulation of multiple physics fields; Step S5: Closed-loop iteration and adaptive optimization, continuously collect multi-dimensional spatiotemporal data, and continuously update the deep reinforcement learning agent using the accumulated multi-dimensional spatiotemporal data to achieve adaptive adjustment for all working conditions.
2. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 1, characterized in that, Step S1 specifically includes: Sensor arrays are deployed in the charged region and condensation region to acquire in real time: local temperature field, local humidity field, local airflow velocity field, and local electric field intensity distribution; A laser particle size analyzer is installed at the downstream outlet of the agglomeration zone to acquire the particle size distribution function in real time; and a high-speed camera is deployed in conjunction with a sheet-like laser light source to perform slice-like imaging of the agglomerates in the agglomeration zone. Through image processing and three-dimensional reconstruction, the spatial number density and fractal dimension are calculated in real time.
3. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 1, characterized in that, Step S1 specifically also includes: By setting a global clock source to provide a unified time reference and using a time synchronization protocol, the acquired multidimensional spatiotemporal data is spatiotemporally synchronized, so that each perceived event is appended with a high-precision global timestamp.
4. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 1, characterized in that, The graphical model that dynamically evolves over time in step S2 specifically includes: The node set includes the sensor node set and the physical area node set; The edge set is mathematically represented by an adjacency matrix, and the construction of the edges follows the principles of physical proximity and hydrodynamic directionality.
5. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 4, characterized in that, Step S2 further includes: A spatiotemporal graph convolutional network is used to process the time series of the graph model. By stacking spatiotemporal convolutional blocks, the spatial topology and temporal evolution of the node features of the graph model are modeled simultaneously, and a high-dimensional feature matrix is output. The high-dimensional feature matrix is aggregated into a fixed-dimensional state vector by using graph pooling or flattening operations.
6. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 5, characterized in that, Step S3 specifically includes: A deep reinforcement learning agent based on the Actor-Critic framework is used to generate the optimal cooperative control strategy by using the state vector as the basis for decision-making. The cooperative action vector output by the deep reinforcement learning agent at the current moment is a multi-dimensional vector containing multiple continuous control variables, specifically including: the magnitude of the driving voltage in the charged region, the driving frequency in the charged region, the peak voltage of the pulse in the condensation region, the pulse frequency in the condensation region, and the pulse duty cycle in the condensation region.
7. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 6, characterized in that, Step S3 further includes: The reward function used by the deep reinforcement learning agent to guide the policy network towards the optimal objective is a multi-objective composite function, used to simultaneously balance cohesion effect and system energy consumption. The reward function includes: The cohesion structure reward term for fractal dimension is used to increase with the fractal dimension to guide the deep reinforcement learning agent to learn control strategies that generate clusters. The reward for increasing particle size is used to quantify the macroscopic effect of particulate matter migrating to larger particle sizes, and it increases with the increase of the number fraction of particulate matter larger than the preset critical particle size to characterize the enhancement of the coagulation effect. The penalty term for the total power consumption of the condensation system is used to reduce the reward as the total power consumption increases, so as to guide the agent to find a low-energy control strategy.
8. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 7, characterized in that, Step S3 further includes: The deep reinforcement learning agent combines model predictive control to predict future states based on a dynamic system model. At each decision time step, the deep reinforcement learning agent solves for an optimal action sequence within a preset prediction time domain, and outputs only the first action in the sequence to step S4 for execution.
9. The method for low-voltage electrocoagulation of low-carbon submicron particles according to claim 8, characterized in that, Step S5 specifically includes: The experience tuples generated and recorded in each cycle are recorded. The experience tuples include the current state, the current action, the current reward, and the next state. The deep reinforcement learning agent updates the internal parameters of the policy network and the value network based on the experience tuples.
10. A low-voltage electrocoagulation device for low-carbon submicron particles, based on the low-voltage electrocoagulation method for low-carbon submicron particles according to any one of claims 1-9, characterized in that, include: The sensing module is used to acquire multi-dimensional spatiotemporal data characterizing the environmental field, micromorphology, and mesoscopic distribution of the charge coagulation process in real time, and to perform spatiotemporal synchronization processing on the multi-dimensional spatiotemporal data. The state representation module is used to construct a graph model that dynamically evolves over time from the multidimensional spatiotemporal data after spatiotemporal synchronization processing, and to extract a unified state vector that integrates spatiotemporal dependencies and represents the global state of the system. The decision-making module generates a globally optimal cooperative action vector that balances cohesion effect and system energy consumption by using a deep reinforcement learning agent based on the unified state vector, multi-objective rewards, and future predictions. The drive execution module is used to parse the cooperative action vector into physical control commands for charging and condensation and execute them to achieve coordinated control of multiple physics fields. The optimization module, based on the data acquired by the perception module after coordinated regulation, and using the accumulated multidimensional spatiotemporal data, continuously updates the deep reinforcement learning agent to achieve adaptive adjustment for all working conditions.
Citation Information
Patent Citations
Intelligent energy-saving optimization method for electric precipitation system based on neural network
CN112099343A
Intelligent electric precipitation system
CN120984439A