Multi-mode reinforcement learning driven PM2.5 chemical component vertical profile inversion system and method

Through a deep reinforcement learning model driven by multimodal reinforcement learning, combined with multimodal data sources, the shortcomings of the PM2.5 inversion system in high-altitude blind spot coverage and real-time performance are solved, and high-precision and real-time vertical profile prediction of PM2.5 chemical components are achieved.

CN120123697AActive Publication Date: 2025-06-10INST OF ATMOSPHERIC PHYSICS CHINESE ACADEMY SCI

Patent Information

Application Number
CN202510589239.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-10
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The prior art has problems of data limitations and poor analysis accuracy in the prediction of spatiotemporal distribution and concentration inversion of PM2.5, especially in the coverage and real-time nature of high-altitude blind spots.

Method used

The PM2.5 chemical component vertical profile inversion system driven by multimodal reinforcement learning is adopted. Through a deep reinforcement learning model, data fusion and prediction are used to combine multimodal data from ground-based lidar, remote sensing satellite and ground monitoring stations.

Benefits of technology

It improves the prediction accuracy and coverage of the vertical profile of PM2.5 chemical components, shortens the inversion delay, achieves minute-level real-time response, and improves the system's adaptability and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123697A_ABST
    Figure CN120123697A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode reinforcement learning driven PM2.5 chemical component vertical profile inversion system and method. The system comprises a deep reinforcement learning model; the deep reinforcement learning model comprises a state space, an Action action space, a reward function, an Actor network and a Critic network; the deep reinforcement learning model is used for calculating a reward value at the current moment in combination with a reward function according to the difference between a vertical profile predicted value of the concentration of each PM2.5 chemical component output by the Action action space and an actual monitoring value; and according to an evaluation result of the strategy evaluation, continuously performing strategy iteration by using a gradient descent method until the deep reinforcement learning model is converged, and obtaining an optimal target deep reinforcement learning model. An Actor-Critic network dynamic optimization strategy is adopted, minute-level model updating is achieved in combination with edge calculation, and the problems of high-altitude blind areas, insufficient real-time performance and the like of a traditional model are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental monitoring, and in particular to a system and method for inverting the vertical profile of PM2.5 chemical components driven by multi-modal reinforcement learning. Background Art

[0002] With the increasing severity of air pollution, especially the impact of PM2.5 (fine particulate matter) on human health and the environment becoming more significant, the accurate monitoring and forecasting of PM2.5 have become one of the key research topics in the field of environmental monitoring. In recent years, significant progress has been made in the application of deep learning technology in air quality prediction and inversion. Traditional models, such as convolutional neural networks (CNNs) and long short-term memory networks (LSTMs), have been widely used in the prediction of the spatio-temporal distribution of PM2.5 and the inversion of its concentration. However, these traditional models have some limitations: Traditional deep learning models usually have problems of data limitations and poor analysis accuracy; the reasons for data limitations are as follows: traditional inversion models rely on single ground-based lidar data, which cannot cover the high-altitude blind area (such as above 6 km), and the spatio-temporal alignment accuracy between satellite data and ground observations is insufficient.

[0003] The poor analysis accuracy mainly stems from two aspects. One is the static model defect and insufficient real-time performance: existing deep learning models (such as CNN-LSTM) are mostly statically trained and difficult to adapt to the dynamic process of pollution diffusion, resulting in a lag in inversion during sudden pollution events. In addition, traditional optimization algorithms (such as NSGA-II) have a high computational complexity and cannot support minute-level real-time inversion.

[0004] Therefore, how to design a more accurate and adaptable inversion system is a key problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a system and method for inverting the vertical profile of PM2.5 chemical components driven by multi-modal reinforcement learning, which solves the above-mentioned technical problems pointed out in the prior art.

[0006] The present invention provides a system for inverting the vertical profile of PM2.5 chemical components driven by multi-modal reinforcement learning, including a deep reinforcement learning model; the deep reinforcement learning model includes a state space, an Action action space, a reward function, an Actor network, and a Critic network; The state space is used to perform data conversion on multi-modal input data and output a normalized state tensor; The Action action space is used to transform multi-modal input data and output a normalized action tensor to realize the predicted value of the vertical profile of the concentration of each chemical component of PM2.5; The Actor network is used to capture the long-range across height layers in the lidar vertical grid by using a multi-layer Transformer encoder; the capture function for the long-range across height layers of the self-attention mechanism is: ; Q: That is, the Query matrix, representing the query matrix; K: That is, the Key matrix, representing the key matrix; V: That is, the Value matrix, representing the content matrix; : is the attention mechanism; : Perform a dot product on the Query matrix and the Key matrix to calculate the similarity between the Query matrix and the Key matrix; : is the scaling factor, used to control the order of magnitude of the dot product result; : is to perform normalization processing on the dot product result; The Critic network is used to input the same state tensor to the Critic network based on GCN to obtain the state value corresponding to the current state; when performing reward evaluation, according to the difference between the predicted value and the actual monitored value of the vertical profile of each chemical component concentration of PM2.5 output, combined with the reward function, calculate the reward value at the current moment; use the state value output by the Critic network and the actually collected reward to calculate the error between the two to obtain the advantage function; when updating the Actor network, use the advantage function to calculate the gradient of the current Actor network.

[0007] Preferably, as an implementable solution; the deep reinforcement learning model further includes a convergence output module; the convergence output module is used to continuously perform policy iteration by using the gradient descent method until the deep reinforcement learning model converges, and then obtain the corresponding Actor-Critic network parameters and output to obtain the optimal target deep reinforcement learning model.

[0008] Preferably, as an implementable solution; the multimodal terminal includes a ground-based lidar, a remote sensing satellite, and a ground monitoring station.

[0009] Preferably, as an implementable solution; the multimodal source data includes lidar optical parameters, satellite remote sensing data, and ground component concentration data.

[0010] The present invention provides a method for inverting the vertical profile of PM2.5 chemical components driven by multimodal reinforcement learning, which uses the above-mentioned system for inverting the vertical profile of PM2.5 chemical components driven by multimodal reinforcement learning to perform processing, including the following operation steps: The multi-modal data input layer obtains multi-modal source data collected by the multi-modal terminal, and performs preprocessing and normalization on the multi-modal source data; then performs an initialization network parameter processing operation on the Actor-Critic network parameters; The data fusion and preprocessing module performs spatio-temporal alignment processing and denoising processing; during spatio-temporal alignment processing, first perform spatial alignment, use Kriging interpolation to match the satellite remote sensing data to the lidar vertical grid, and then use the dynamic time warping (DTW) algorithm to measure the similarity of time series with different lengths and align the time series of the ground monitoring station; during denoising processing, apply wavelet threshold denoising to the lidar optical parameters to obtain the denoised data; The deep reinforcement learning model calculates the reward value at the current moment according to the difference between the predicted value and the actual monitoring value of the vertical profile of each chemical component concentration of PM2.5 output by the Action action space, in combination with the reward function; Then perform policy evaluation on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, calculate the gradient of the Actor network based on the policy gradient method, and continuously perform policy iteration using the gradient descent method until the deep reinforcement learning model converges, and then obtain the corresponding Actor-Critic network parameters and output the optimal target deep reinforcement learning model.

[0011] Preferably, as an implementable solution; the execution of the initialization network parameter processing operation specifically includes: configuring a multi-layer Transformer encoder for the Actor network; Configure the structure, initial weights and hyperparameters of the Critic network based on GCN.

[0012] Preferably, as an implementable solution; the deep reinforcement learning model calculates the reward value at the current moment according to the difference between the predicted value and the actual monitoring value of the vertical profile of each chemical component concentration of PM2.5 output by the Action action space, in combination with the reward function; then perform policy evaluation on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, calculate the gradient of the Actor network based on the policy gradient method, specifically including: The current state tensor of the normalized multi-modal data obtained; Convert the multi-modal input data and output a normalized action tensor; The Actor network performs forward propagation, uses a multi-layer Transformer encoder to capture and integrate long-range dependencies across height layers through the self-attention mechanism, and outputs the predicted values of the vertical profiles of each chemical component concentration of PM2.5; Perform forward propagation through the Critic network: Input the same state tensor into the GCN-based Critic network to obtain the state value corresponding to the current state; When performing reward evaluation, calculate the reward value at the current moment according to the difference between the predicted vertical profile of the PM2.5 chemical component concentrations output and the actual monitored values, in combination with the reward function; Calculate the advantage function by calculating the error between the state value output by the Critic network and the actually collected reward; When updating the Actor network, use the advantage function to calculate the gradient of the current Actor network.

[0013] Preferably, as an implementable solution; the reward function is: ; where α and β are weight coefficients, and Temporal Inconsistency is to punish the mutation of the prediction results of adjacent time steps.

[0014] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages: Analyzing the above multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion system and method provided by the present invention, it can be seen that in specific applications, first obtain data within the vertical 0-6 km space range through ground-based lidar (optical parameters), remote sensing satellites, and ground monitoring stations (PM2.5 component concentrations). This data covers satellite images and near-surface component concentrations, providing multi-modal source data input for the model; and perform preprocessing and normalization processing on the multi-modal source data; then perform the initialization network parameter processing operation of the Actor-Critic network parameters; The data fusion and preprocessing module performs spatio-temporal alignment processing and denoising processing; during spatio-temporal alignment processing, first perform spatial alignment, use Kriging interpolation to match the satellite remote sensing data to the lidar vertical grid, and then use the dynamic time warping (DTW) algorithm to measure the similarity of time series with different lengths and align the time series of the ground monitoring stations; during denoising processing, apply wavelet threshold denoising to the lidar optical parameters to obtain the denoised data; spatio-temporal alignment of the data through Kriging interpolation and the dynamic time warping (DTW) algorithm solves the problem of data inconsistency in time and space; at the same time, applying wavelet threshold denoising to the lidar optical parameters effectively reduces noise interference, improves data quality, and provides a technical foundation guarantee for subsequent data operations; The difference between the predicted values of the vertical profiles of the concentrations of each chemical component of PM2.5 output by the deep reinforcement learning model according to the Action action space and the actual monitored values is combined with the reward function to calculate the reward value at the current moment; then the current policy is evaluated to obtain the evaluation result; then, according to the evaluation result of the policy evaluation, the gradient of the Actor network is calculated based on the policy gradient method, and the policy iteration is continuously performed using the gradient descent method until the deep reinforcement learning model converges, and then the corresponding Actor-Critic network parameters are obtained and the optimal target deep reinforcement learning model is output. The policy gradient method optimizes the Actor network parameters through the "state-action-reward" loop, enabling the model to dynamically adapt to the time-varying characteristics of PM2.5 components such as diurnal changes and seasonal migrations; the above-mentioned deep reinforcement learning model uses an Actor-Critic network, combined with a reward function and a policy gradient method. Through the dynamic optimization of reinforcement learning, the model can adapt to changes in different times and spaces and improve the generalization ability. Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the overall architecture of a multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion system; Figure 2 It is a schematic diagram of the specific architecture principle of a multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion system; Figure 3 It is a schematic diagram of the deep reinforcement learning model in a multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion system; Figure 4 It is a schematic diagram of the main process of a multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion method; Figure 5 It is a schematic diagram of the comparison of error results at different height levels in an experimental data of a multi-modal reinforcement learning-driven PM2.5 chemical component vertical profile inversion method.

[0016] Reference Signs: Multi-modal data input layer 10; Data fusion and preprocessing module 20; Deep reinforcement learning model 30; State space 31; Action action space 32; Reward function 33; Actor network 34; Critic network 35. Detailed Embodiments

[0017] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.

[0019] Embodiment 1 As Figure 1 and Figure 2 shown, Embodiment 1 of the present invention provides a PM2.5 chemical component vertical profile inversion system driven by multi-modal reinforcement learning, including a multi-modal data input layer 10, a data fusion and preprocessing module 20, and a deep reinforcement learning model 30: The multi-modal data input layer 10 is used to obtain multi-modal source data collected by multi-modal terminals, perform preprocessing and normalization processing on the multi-modal source data; and then perform an initialization network parameter processing operation on the Actor-Critic network parameters; The data fusion and preprocessing module 20 is used to perform spatio-temporal alignment processing and denoising processing; during spatio-temporal alignment processing, first perform spatial alignment, use Kriging interpolation to match the satellite remote sensing data to the lidar vertical grid, and then use the dynamic time warping (DTW) algorithm to measure the similarity of time series with different lengths and align the time series of ground monitoring stations; during denoising processing, apply wavelet threshold denoising to the lidar optical parameters to obtain denoised data; The deep reinforcement learning model 30 is used to calculate the reward value at the current moment based on the difference between the predicted value of the vertical profile of the concentration of each chemical component of PM2.5 output by the Action action space and the actual monitoring value, in combination with the reward function; then perform policy evaluation on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, calculate the gradient of the Actor network based on the policy gradient method, and continuously perform policy iteration using the gradient descent method until the deep reinforcement learning model converges, and then obtain the corresponding Actor-Critic network parameters and output the optimal target deep reinforcement learning model.

[0020] The core architecture of the PM2.5 chemical component vertical profile inversion system driven by multi-modal reinforcement learning of the present invention includes the following modules (see Figure 2 and Figure 3 ) The above system includes a deep reinforcement learning model; the deep reinforcement learning model 30 includes a state space 31, an Action action space 32, a reward function 33, an Actor network 34, and a Critic network 35; The state space 31 (State Space) is used to perform data conversion on multi-modal input data (lidar optical parameters, satellite AOD, ground component concentration) and output a normalized state tensor; The Action action space 32 is used to convert the multi-modal input data into a normalized action tensor, and realize the vertical profile prediction values of the concentrations of various chemical components of PM2.5; The Actor network 34 is used to adopt a multi-layer Transformer encoder to capture the long range across altitude layers in the lidar vertical grid through the self-attention mechanism; the capture function of the long range across altitude layers of the self-attention mechanism is: ; The Critic network 35 is used to input the same state tensor to the Critic network based on GCN to obtain the state value corresponding to the current state; when performing reward evaluation, according to the difference between the predicted vertical profile values of the concentrations of various chemical components of PM2.5 output and the actual monitoring values, combined with the reward function, calculate the reward value at the current moment; use the state value output by the Critic network and the actually collected reward to calculate the error between the two to obtain the advantage function; when updating the Actor network, use the advantage function to calculate the gradient of the current Actor network.

[0021] Multi-modal data input layer 10: Ground-based lidar: Provide vertical profile data of the 532 nm aerosol backscattering coefficient, extinction coefficient and depolarization ratio (time resolution 1 hour, 60 vertical layers). Satellite remote sensing data: Integrate MODIS aerosol optical depth (AOD) and CALIPSO vertical profile data to make up for the high-altitude blind area. Ground monitoring stations: Obtain real-time monitoring values of the concentrations of near-surface PM2.5 chemical components (SO 4 ² - 、NO 3 - 、NH 4 + 、OM, BC).

[0022] Data fusion and preprocessing module 20: Spatiotemporal alignment algorithm: Use Kriging interpolation to match satellite data to the lidar vertical grid, and use dynamic time warping (DTW) to align the time series of ground monitoring stations. Noise suppression: Apply wavelet threshold denoising (Formula 1) to lidar data to eliminate instrument noise and weather interference.

[0023] ; Formula (1); Among them, is the wavelet basis function, is the threshold function.

[0024] The above-mentioned deep reinforcement learning model 30 (belonging to the framework of the Actor-Critic network): State Space 31: The normalized tensor of multi-modal input data (lidar optical parameters, satellite AOD, ground component concentrations).

[0025] Action Space 32: The predicted vertical profiles of the concentrations of each chemical component of PM2.5.

[0026] Reward Function 33: ; Equation 2; where α and β are weight coefficients, and Temporal Inconsistency is to penalize the mutation of the prediction results at adjacent time steps.

[0027] Actor Network 34: Adopts a multi-layer Transformer encoder to capture long-range dependencies across height layers through the self-attention mechanism (Equation 3): ; Equation 3; Q: Query, the query matrix, represents the input query vector and is responsible for locating the information that needs to be focused on currently. K: Key, the key matrix, is used to measure the association strength between different positions. V: Value, the content matrix, is used to store the actual semantic or spatial information to be transmitted.

[0028] : The attention mechanism focuses on local information, thereby suppressing useless information and improving the efficiency and accuracy of the model.

[0029] : Performs a dot product on the Query matrix and the Key matrix to calculate the similarity between the Query and the Key.

[0030] : The scaling factor is used to control the order of magnitude of the dot product result and prevent gradient vanishing or explosion. That is, through scaling, it ensures that the gradient of the softmax output will not be too small / large, improving stability.

[0031] : Normalizes the dot product result, that is, normalizes the similarity.

[0032] Critic Network 35: Evaluates the state value based on the graph convolutional network (GCN) and dynamically adjusts the Actor strategy.

[0033] Online learning and optimization module: Real-time data stream processing: Deploy edge computing nodes to support the pipeline operations of data reception, preprocessing, and model inference. Progressive update strategy: Adopts the proximal policy optimization (PPO) algorithm (Equation 4) to update the model parameters every 10 minutes to avoid policy oscillation: ; Equation 4; where is the advantage function estimated under the old policy, is the importance weight, representing the probability ratio of the same action under the old and new policies; is to take the ratio constrained within range of the clipping function.

[0034] The above system fills the high-altitude blind area through satellite data, and jointly constrains the near-surface component concentration with lidar and ground monitoring station data, improving the spatial coverage integrity by 40%. The reinforcement learning framework enables the model to respond in real time during pollution events, reducing the inversion delay from hours to minutes. At the same time, by combining edge computing with the PPO algorithm, the optimization speed is increased by 50% compared with the traditional NSGA-II.

[0035] Preferably, as an implementable solution; the deep reinforcement learning model further includes a convergence output module; the convergence output module is used to continuously perform policy iteration using the gradient descent method until the deep reinforcement learning model converges, and then obtain the corresponding Actor-Critic network parameters and output the optimal target deep reinforcement learning model.

[0036] Preferably, as an implementable solution; the multimodal terminal includes a ground-based lidar, a remote sensing satellite, and a ground monitoring station.

[0037] Preferably, as an implementable solution; the multimodal source data includes lidar optical parameters, satellite remote sensing data, and ground component concentration data.

[0038] It should be noted that the lidar optical parameters include vertical profile data of aerosol backscattering coefficient, extinction coefficient, and depolarization ratio; the satellite remote sensing data includes MODIS AOD data and CALIPSO vertical profile data; where the MODIS AOD data is the MODIS aerosol optical depth. The ground component concentration data is the ground PM2.5 chemical component concentration. The chemical components are mainly SO 4 ² - 、NO 3 - 、NH 4 + 、OM, BC, etc.

[0039] Embodiment 2 See Figure 4 This embodiment 2 of the present invention provides a method for inverting the vertical profile of PM2.5 chemical components driven by multimodal reinforcement learning, including the following operation steps: Step S10: The multi-modal data input layer acquires the multi-modal source data collected by the multi-modal terminal, performs preprocessing and normalization on the multi-modal source data; then performs the initialization network parameter processing operation for the Actor-Critic network parameters; The multi-modal source data needs to be separately subjected to data cleaning and outlier supplementation. For example, for ground-based lidar data, wavelet decomposition (Daubechies-4 basis) is performed on the original backscattering coefficient (σ_bsc), and the hard threshold λ is set to 3 times the noise standard deviation; afterwards, unified normalization processing is applied to various types of data to convert data with different dimensions into normalized tensors to ensure that the values are on the same scale during subsequent network (Actor and Critic) processing; perform the initialization network parameter processing operation: configure the structures, initial weights, and hyperparameters (learning rate, weight decay, gradient threshold, etc.) of the Actor network (multi-layer Transformer encoder) and the Critic network (graph convolutional network, GCN).

[0040] Step S20: The data fusion and preprocessing module performs spatio-temporal alignment processing and denoising processing; during spatio-temporal alignment processing, first perform spatial alignment, and use Kriging interpolation to match the satellite remote sensing data to the lidar vertical grid, and then use the dynamic time warping (DTW) algorithm to measure the similarity of time series with different lengths and align the time series of the ground monitoring station; during denoising processing, apply wavelet threshold denoising to the lidar optical parameters to obtain the denoised data; In the data fusion and preprocessing module, first for spatial alignment, use the Kriging interpolation method, based on regionalized variables and with the variogram as the basic tool, to perform linear unbiased and optimal estimation of unknown sample points. It is necessary to calculate the variogram of the satellite data in the vertical direction, determine the spatial correlation structure, and fit the variogram parameters by the maximum likelihood estimation method; divide the vertical layer according to the detection range of the radar to generate grid coordinate points (0 - 6 km, 60 layers), with a spatial resolution of 1 km × 1 km, and match the satellite data to the lidar vertical grid. Use the dynamic time warping (DTW) to measure the similarity of time series with different lengths and align the time series of the ground monitoring station, with the time error controlled within 5 minutes. Then apply wavelet threshold denoising to the lidar data. The denoising method is shown in formula (1) to eliminate instrument noise and weather interference.

[0041] Step S30: The deep reinforcement learning model calculates the reward value at the current moment according to the difference between the predicted vertical profile values of the PM2.5 chemical component concentrations output by the Action action space and the actual monitoring values, in combination with the reward function; Then, perform policy evaluation on the current policy to obtain an evaluation result; then, based on the evaluation result of the policy evaluation, calculate the gradient of the Actor network using the policy gradient method, and continuously perform policy iteration using the gradient descent method until the deep reinforcement learning model converges, and then obtain the corresponding Actor-Critic network parameters and output to obtain the optimal target deep reinforcement learning model.

[0042] That is, the normalized and multi-modal data obtained in step S1 / step S2 form the current state tensor (State), which represents the comprehensive information of the current environment at each moment. The Actor network performs forward propagation. Using a multi-layer Transformer encoder, it captures and integrates the long-range dependencies across altitude layers in the lidar vertical grid through the self-attention mechanism, and outputs the prediction of the vertical profiles of the concentrations of each chemical component of PM2.5 (i.e., the action, Action action space). The inter-layer information interaction of the Transformer ensures the full modeling of the dependencies between vertical structures (see formula 3).

[0043] And perform the forward propagation of the Critic network: input the same state tensor into the GCN-based Critic network to obtain the value estimate corresponding to the current state, which is used for subsequent TD (Temporal Difference) error calculation and policy evaluation.

[0044] When performing reward evaluation, calculate the reward at the current moment according to the difference between the action output and the actual monitored value, in combination with the reward function (formula 2).

[0045] Weight coefficients such as α and β are introduced into the reward function, and at the same time, the Temporal Inconsistency term is used to punish the drastic changes in the prediction results of adjacent time steps to ensure prediction smoothness.

[0046] Use the state value output by the Critic network and the actually collected reward to calculate the TD error (or advantage function) to measure the quality of the current policy. Based on the policy gradient method, use the advantage as a weighting term to calculate the gradient of the Actor network.

[0047] The goal is to maximize the expected reward. Therefore, backpropagate the loss between the predicted action and the actual high-altitude vertical profile, and at the same time consider the Temporal Inconsistency penalty term.

[0048] When updating the Actor network, it is necessary to use the advantage function to calculate the policy gradient.

[0049] To avoid the update of the Critic network affecting the update of the Actor network, the stop-gradient operator can be used to prevent gradient information from flowing to the Critic network. Before formally using the model, offline pre-training can be carried out, that is, the network is pre-trained using historical data. To constrain the policy network from deviating from the behavioral policy in the offline data, a regularization term (such as KL divergence) can be introduced during policy update to limit the change of the policy. Then, online fine-tuning is performed. The real-time data stream is input into the edge node, and the PPO algorithm updates the Actor-Critic network parameters every 10 minutes. If the reward value drops by more than 10% for three consecutive iterations, the model is rolled back to the previous stable version. Then, after the deep reinforcement learning model converges, the target deep reinforcement learning model should be output.

[0050] The Actor-Critic algorithm combines policy optimization (Actor) and value evaluation (Critic). Specifically: Actor network (policy network): responsible for generating policies. Given a state, it outputs the probability distribution of actions. It is updated through policy gradients to maximize the expected return. Critic network (value network): responsible for evaluating the value of the current policy. Given a state, it outputs the state value V, which is used to estimate the temporal difference (TD) error, thereby guiding the update of the policy.

[0051] Preferably, as an implementable solution; the operation of processing the initialized network parameters specifically includes: configuring the multi-layer Transformer encoder of the Actor network; Configuring the structure, initial weights, and hyperparameters of the Critic network based on GCN.

[0052] Preferably, as an implementable solution; the deep reinforcement learning model calculates the reward value at the current moment according to the difference between the predicted vertical profiles of the concentrations of each chemical component of PM2.5 output by the Action action space and the actual monitored values, in combination with the reward function; then performs policy evaluation on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, calculates the gradient of the Actor network based on the policy gradient method, specifically including: Step S31: The current state tensor of the normalized multi-modal data obtained in Step S1 / Step S2; Step S32: Converting the multi-modal input data to output a normalized action tensor; Step S33: The Actor network performs forward propagation. Using the multi-layer Transformer encoder, it captures and integrates the long-range dependencies across height layers through the self-attention mechanism, and outputs the predicted values of the vertical profiles of the concentrations of each chemical component of PM2.5; Step S34: Perform forward propagation through the Critic network: Input the same state tensor into the GCN-based Critic network to obtain the state value corresponding to the current state, which is used for subsequent TD (Temporal Difference) error calculation and policy evaluation.

[0053] Step S35: When performing reward evaluation, calculate the reward value at the current moment according to the difference between the predicted vertical profile values of each chemical component concentration of PM2.5 and the actual monitored values, in combination with the reward function. Weight coefficients such as α and β are introduced into the reward function. At the same time, the Temporal Inconsistency term is used to penalize the drastic changes in the prediction results of adjacent time steps to ensure prediction smoothness.

[0054] Calculate the advantage function by calculating the error between the state value output by the Critic network and the actually collected reward. Step S36: When updating the Actor network, use the advantage function to calculate the gradient of the current Actor network.

[0055] Preferably, as an implementable solution; the reward function 30 (Reward Function): ; where α and β are weight coefficients, and Temporal Inconsistency is to penalize the mutation of the prediction results of adjacent time steps. RMSE is the root mean square error, where represents the predicted value (i.e., the predicted vertical profile value), where represents the observed value (i.e., the actual monitored value); It should be noted that the researchers conducted the following experimental tests based on the above embodiments of the present application: 1 - Execute the data preprocessing process: Lidar data denoising: Perform wavelet decomposition (Daubechies-4 basis) on the original backscattering coefficient (σ_bsc), and set the hard threshold λ to 3 times the noise standard deviation.

[0056] Satellite data interpolation: Use Kriging interpolation to map MODIS AOD data to the lidar vertical grid (0 - 6 km, 60 layers), with a spatial resolution of 1 km × 1 km.

[0057] Time alignment: Align the ground monitoring station data and the lidar data to a unified timestamp through dynamic time warping (DTW) (time error < 5 minutes).

[0058] 2 - First, perform the offline model training step: Offline pre - training: Input: One - year historical data (time resolution: 1 hour), including pollution events (such as sandstorms) and non - polluted periods.

[0059] Objective: Minimize RMSE and KL divergence (Equation 5) to ensure that the prediction distribution conforms to physical laws: (5) Online fine - tuning: The real - time data stream is input into the edge node, and the PPO algorithm updates the Actor - Critic network parameters every 10 minutes.

[0060] If the reward value drops by more than 10% for three consecutive iterations, trigger the model to roll back to the previous stable version.

[0061] 3 - Output the test results 1. Experimental design Selected dataset: Sandstorm events (March 15 - 20, 2022) and conventional monitoring data in the Beijing - Tianjin - Hebei region.

[0062] Comparison baseline: The original patent CNN - ATT - BiLSTM - NSGA model, traditional chemical transport model (WRF - Chem).

[0063] 2. Performance metrics Accuracy: CORR (correlation coefficient), RMSE (μg / m³).

[0064] Real - time performance: Inversion delay (time from data input to result output).

[0065] Result analysis Ground - layer accuracy (see Table 1):

[0066] Upper - layer coverage ( Figure 5 ) : The fusion of satellite data above 6 km reduces the inversion error of OM concentration by 30% (from 4.2 μg / m³ to 2.9 μg / m³). At the same time Figure 5 also shows the comparison of error results at different altitude layers; Figure 5 also shows: = Comparison effect of upper - layer OM concentration inversion before and after satellite data fusion (verified by aerial survey).

[0067] Real - time response (see Table 2): During sandstorm events, the model updates parameters within 10 minutes, the inversion delay < 2 minutes, and the false - alarm rate is reduced by 40% compared to the baseline.

[0068] Table 2: Comparison of real - time response

[0069] In summary, a PM2.5 chemical composition vertical profile inversion system and method driven by multi-modal reinforcement learning proposed in the embodiments of the present invention is a real-time inversion system and method for the vertical profiles of PM2.5 chemical components (including sulfate, nitrate, ammonium salt, organic matter, and black carbon) based on multi-modal data fusion and deep reinforcement learning. By integrating ground-based lidar, satellite remote sensing, and ground monitoring station data and combining with a reinforcement learning framework, the system realizes dynamic monitoring of the vertical distribution of PM2.5 chemical components with high spatio-temporal resolution and high precision, and is applicable to emergency response to sudden pollution events and long-term air quality assessment.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; those of ordinary skill in the art can modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A PM2.5 chemical component vertical profile inversion system driven by multimodal reinforcement learning, characterized in that: Including multimodal data input layer, data fusion and preprocessing module, deep reinforcement learning model: The multimodal data input layer is used to obtain the multimodal source data collected by the multimodal terminal, and perform preprocessing and normalization processing on the multimodal source data; then perform the initialization network parameter processing operation of the Actor-Critic network parameters; The data fusion and preprocessing module is used to perform spatiotemporal alignment processing and denoising processing; in the spatiotemporal alignment processing, firstly, spatial alignment is performed, and Kriging interpolation is used to match the satellite remote sensing data to the vertical grid of the laser radar, and then the dynamic time warping algorithm is used to measure the similarity of time series of different lengths to align the time series of the ground monitoring station; in the denoising processing, wavelet threshold denoising is applied to the laser radar optical parameters to obtain denoised data; The deep reinforcement learning model is used to calculate the reward value at the current moment based on the difference between the vertical profile prediction value of each chemical component concentration of PM2.5 output in the Action action space and the actual monitoring value, combined with the reward function; Then, a policy evaluation is performed on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, the gradient of the Actor network is calculated based on the policy gradient method, and the policy iteration is continuously performed using the gradient descent method until the deep reinforcement learning model converges, and the corresponding Actor-Critic network parameters are obtained and the optimal target deep reinforcement learning model is output.

2. According to claim 1, a multimodal reinforcement learning driven PM2.5 chemical component vertical profile inversion system is characterized in that: Including a deep reinforcement learning model; the deep reinforcement learning model includes a state space, an action space, a reward function, an Actor network and a Critic network; The state space is used to perform data conversion on multimodal input data and output a normalized state tensor; The Action action space is used to transform the multimodal input data and output a normalized action tensor to achieve the vertical profile prediction value of the concentration of each chemical component of PM2.5; The Actor network is used to use a multi-layer Transformer encoder to capture the long range across height layers in the vertical grid of the lidar through a self-attention mechanism; the capture function of the long range across height layers of the self-attention mechanism is: ; Q: Query matrix, indicating the query matrix; K: Key matrix, representing the key matrix; V: Value matrix, which represents the content matrix; : is the attention mechanism; : Perform dot product on the Query matrix and the Key matrix to calculate the similarity between the Query matrix and the Key matrix; : is a scaling factor used to control the order of magnitude of the dot product result; : To normalize the dot product result; The Critic network is used to input the same state tensor to the GCN-based Critic network to obtain the state value corresponding to the current state; when performing reward evaluation, the reward value at the current moment is calculated based on the difference between the output vertical profile prediction value of each chemical component concentration of PM2.5 and the actual monitoring value, combined with the reward function; The state value output by the Critic network and the actual reward collected are used to calculate the error between the two to obtain the advantage function; when updating the Actor network, the advantage function is used to calculate the gradient of the current Actor network.

3. A PM2.5 chemical component vertical profile inversion system driven by multimodal reinforcement learning according to claim 2, characterized in that: The deep reinforcement learning model also includes a convergence output module; the convergence output module is used to continuously perform strategy iteration using a gradient descent method until the deep reinforcement learning model converges, obtains corresponding Actor-Critic network parameters, and outputs an optimal target deep reinforcement learning model.

4. According to claim 1, a multimodal reinforcement learning driven PM2.5 chemical component vertical profile inversion system is characterized in that: The multimodal terminal includes a ground-based laser radar, a remote sensing satellite and a ground monitoring station.

5. According to claim 3, a multimodal reinforcement learning driven PM2.5 chemical component vertical profile inversion system is characterized in that: The multimodal source data include laser radar optical parameters, satellite remote sensing data, and ground component concentration data.

6. A multimodal reinforcement learning driven PM2.5 chemical component vertical profile inversion method, characterized in that: The processing is performed using the PM2.5 chemical component vertical profile inversion system driven by multimodal reinforcement learning as described in any one of claims 1 to 5, including the following steps: The multimodal data input layer obtains the multimodal source data collected by the multimodal terminal, and performs preprocessing and normalization on the multimodal source data; then performs the initialization network parameter processing operation of the Actor-Critic network parameters; The data fusion and preprocessing module performs spatiotemporal alignment processing and denoising processing; in the spatiotemporal alignment processing, firstly, spatial alignment is performed, and the satellite remote sensing data is matched to the vertical grid of the lidar by using Kriging interpolation, and then the dynamic time warping algorithm is used to measure the similarity of time series of different lengths and align the time series of the ground monitoring station; in the denoising processing, wavelet threshold denoising is applied to the optical parameters of the lidar to obtain the denoised data; The deep reinforcement learning model calculates the reward value at the current moment based on the difference between the vertical profile prediction value of each chemical component concentration of PM2.5 output in the action space and the actual monitored value, combined with the reward function; Then, a policy evaluation is performed on the current policy to obtain an evaluation result; then, according to the evaluation result of the policy evaluation, the gradient of the Actor network is calculated based on the policy gradient method, and the policy iteration is continuously performed using the gradient descent method until the deep reinforcement learning model converges, and the corresponding Actor-Critic network parameters are obtained and the optimal target deep reinforcement learning model is output.

7. The method for inverting vertical profiles of PM2.5 chemical components driven by multimodal reinforcement learning according to claim 6, characterized in that: The execution of the initialization network parameter processing operation specifically includes: configuring a multi-layer Transformer encoder of the Actor network; Configure the structure, initial weights, and hyperparameters of the GCN-based Critic network.

8. The method for inverting vertical profiles of PM2.5 chemical components driven by multimodal reinforcement learning according to claim 7, characterized in that: The deep reinforcement learning model calculates the reward value at the current moment based on the difference between the vertical profile prediction value of each chemical component concentration of PM2.5 output in the action space and the actual monitoring value, combined with the reward function; Then, the current strategy is evaluated to obtain the evaluation result; Then, according to the evaluation results of the policy evaluation, the gradient of the Actor network is calculated based on the policy gradient method, including: The current state tensor of the normalized multimodal data obtained; Transform the multimodal input data and output a normalized action tensor; The Actor network performs forward propagation, using a multi-layer Transformer encoder to capture and integrate long-range dependencies across altitude layers through a self-attention mechanism, and outputs the predicted values ​​of the vertical profiles of the PM2.5 chemical component concentrations. And perform forward propagation through the Critic network: input the same state tensor to the GCN-based Critic network to obtain the state value corresponding to the current state; When performing reward evaluation, the reward value at the current moment is calculated based on the difference between the output vertical profile prediction value of each chemical component concentration of PM2.5 and the actual monitoring value, combined with the reward function; Using the state value output by the Critic network and the actual reward collected, the error between the two is calculated to obtain the advantage function; When updating the Actor network, the advantage function is used to calculate the gradient of the current Actor network.

9. The method for inverting vertical profiles of PM2.5 chemical components driven by multimodal reinforcement learning according to claim 8, characterized in that: The reward function is: ; Among them, α and β are weight coefficients, and Temporal Inconsistency is used to penalize the sudden changes in the prediction results of adjacent time steps.

Citation Information

Patent Citations

  • Inversion method of aerogel vertical profile based on laser radar

    CN110441777A

  • Method for detecting aerosol mass concentration profile by using single-wavelength laser radar

    CN112269189A

  • Vehicle path problem solving method integrating deep neural network and reinforcement learning

    CN115545350A

  • Space-time process simulation method and system based on stacked space-time memory units

    CN115719036A

  • Electric energy metering detection information fault diagnosis method

    CN118133203A

Cited By

  • CFD-DEM coupling acceleration method

    CN120764425A

  • High-precision self-adaptive control method and system for floating sealing device of lithium extraction rotary kiln

    CN120819984A

  • Raindrop spectrum parameter profile inversion method and system

    CN121028256A

  • Denitration reactor temperature field self-balancing control method and system based on directional ConvLSTM

    CN121070085A

  • Fusion method of ground station monitoring and satellite remote sensing inversion PM2.5 data

    CN121117940A