Transform and MOPSO-based PM2.5 chemical component vertical profile inversion system and method
Through the combination of Transformer encoder and MOPSO algorithm, the problems of insufficient space-time modeling and high computational complexity in PM2.5 monitoring are solved, and efficient and accurate PM2.5 chemical component inversion is achieved, which is suitable for urban air quality management and pollution source tracking.
Patent Information
- Application Number
- CN202510696508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has insufficient space-time modeling, low data fusion efficiency, and high computational complexity in PM2.5 monitoring, resulting in large inversion errors. The particle swarm optimization algorithm has divergence risks in terms of hyperparameter sensitivity, which affects the optimization effect.
The Transformer encoder combined with multi-objective particle swarm optimization (MOPSO) is used to perform vertical profile inversion of PM2.5 chemical components. Through the preprocessing and fusion of foundation lidar and ERA5 meteorological data, combined with global attention mechanism and efficient optimization strategy, the consumption of computing resources is reduced and the ability of collaborative inversion of multimodal data is improved.
It realizes high-precision and low-cost vertical profile inversion of PM2.5 chemical components, reduces inversion error, improves computing efficiency and Pareto solution set diversity, and is suitable for urban air quality management and pollution source tracking.
Smart Images

Figure CN120254879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air pollution monitoring, and in particular, to a PM2.5 chemical component vertical profile inversion system and method based on Transformer and MOPSO. Background Art
[0002] With the development of air environment monitoring technology, the monitoring accuracy of PM2.5 has been continuously improved. However, due to factors such as insufficient spatio-temporal modeling, low data fusion efficiency, and high computational complexity, many challenges still remain. Traditional CNN-ATT-BiLSTM models and WRF-Chem chemical transport models rely on local convolution operations and are difficult to capture long-distance non-linear relationships between vertical height layers. Especially in complex environments such as sandstorms, the dynamic correlation between the diffusion process of high-level aerosols and the underlying pollution sources is difficult to effectively model, resulting in large inversion errors. In addition, existing data fusion methods often use simple splicing or weighted average methods and cannot dynamically allocate feature weights according to the correlation between different height layers, thus affecting the accuracy of the model.
[0003] In terms of computational complexity, multi-objective optimization algorithms such as NSGA-II usually require frequent calculation of non-dominated sorting, which significantly increases the calculation time and cannot meet the requirements of real-time monitoring. For example, in the optimization process with a population of 1000 particles, the time-consuming of a single iteration of non-dominated sorting can reach 0.6 seconds, resulting in the overall optimization time exceeding the tolerance threshold of the actual system. To solve this problem, the particle swarm optimization algorithm (PSO) has gradually become an alternative because it has lower computational complexity and faster convergence speed compared to the NSGA-II algorithm. However, the multi-objective particle swarm optimization algorithm (MOPSO) still faces certain challenges in application. Especially in the process of updating particle velocities, the sensitivity of hyperparameters is not fully considered, which may lead to the risk of divergence during the training process, resulting in errors in the update of the particle movement speed during the particle swarm update process, and further affecting the optimization effect. Therefore, how to improve the update mechanism of the particle swarm optimization algorithm and avoid the negative impact of hyperparameter fluctuations on the optimization process has become the key to improving the stability and accuracy of the algorithm.
[0004] In addition, traditional Pareto solution set screening methods rely on global non-dominated sorting, with a computational complexity of O(N²), which is less efficient in large-scale solution sets. The crowding distance calculation does not normalize the differences in objective dimensions, resulting in an imbalance in maintaining the diversity of the solution set and further affecting the effect of the optimization process. Therefore, how to solve these problems has become the key to improving the monitoring accuracy and efficiency of PM2.5. Summary of the Invention
[0005] The object of the present invention is to provide a PM2.5 chemical component vertical profile inversion system and method based on Transformer and MOPSO, which solves the above technical problems pointed out in the prior art.
[0006] The present invention provides a PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO, including a basic data acquisition module and an inversion module; The basic data acquisition module is used to collect and obtain ground-based lidar data X and ERA5 meteorological data; the ground-based lidar data X and ERA5 meteorological data are preprocessed and then normalized to obtain a standard data tensor; The inversion module is used to perform PM2.5 chemical component inversion based on the standard data tensor through an inversion model obtained by pre-training with a Transformer encoder combined with multi-objective particle swarm optimization processing, and output an inversion result; Correspondingly, the present invention also proposes a PM2.5 chemical component vertical profile inversion method based on Transformer and MOPSO, including the following operating steps: Collect and obtain ground-based lidar data X and ERA5 meteorological data; the ground-based lidar data X and ERA5 meteorological data are preprocessed and then normalized to obtain a standard data tensor; Perform PM2.5 chemical component inversion based on the standard data tensor through an inversion model obtained by pre-training with a Transformer encoder combined with multi-objective particle swarm optimization processing, and output an inversion result.
[0007] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages: Analyzing the above PM2.5 chemical component vertical profile inversion system and method based on Transformer and MOPSO provided by the present invention, in specific applications, first, a ground-based lidar data with 60 vertical height layers is acquired by using a 532nm wavelength lidar at a time resolution of 1 hour, and ERA5 meteorological data including parameters such as integrated temperature, humidity, wind speed, and boundary layer height is collected at a time resolution of 1 hour and a spatial resolution of 0.25°×0.25°, providing a rich data basis for subsequent inversion processing; after preprocessing the ground-based lidar data and ERA5 meteorological data for denoising, data filling, and alignment, the preprocessed data (i.e., the above standard data tensor) is used for PM2.5 chemical component inversion through the trained inversion model, and the inversion result is output. By fusing lidar and meteorological data and combining the global attention mechanism and efficient optimization strategy, problems such as insufficient spatio-temporal modeling and high computational costs of traditional models are solved; specifically, the above trained inversion model models spatio-temporal dependencies through the Transformer encoder and the global attention mechanism to achieve spatio-temporal dependency modeling across height layers. During this period, through the dynamic data fusion strategy, the collaborative inversion ability of multi-modal data is improved, and furthermore, by using multi-objective particle swarm optimization to optimize the multi-objective algorithm framework, the consumption of computing resources is reduced and the diversity of the Pareto solution set is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 FIG. is a schematic diagram of the overall architecture of a PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO; Figure 2 FIG. is a schematic diagram of the main process simulation of a PM2.5 chemical component vertical profile inversion method based on Transformer and MOPSO; Figure 3 FIG. is a schematic diagram of the overall process simulation of a PM2.5 chemical component vertical profile inversion method based on Transformer and MOPSO; Figure 4 FIG. is a heat map of the self-attention weights of Transformer in a PM2.5 chemical component vertical profile inversion method based on Transformer and MOPSO; Figure 5 FIG. is a comparison of the MOPSO Pareto front distribution and a simulation diagram of the diversity of the NSGA-II solution set in a PM2.5 chemical component vertical profile inversion method based on Transformer and MOPSO.
[0009] Reference numerals: basic data acquisition module 10, inversion module 20, inversion model training module 21, initial inversion module 211, optimization and update module 212. Detailed implementation manners
[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0011] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.
[0012] Embodiment 1 As Figure 1 shown, Embodiment 1 of the present invention provides a PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO, including a basic data acquisition module 10 and an inversion module 20; The basic data acquisition module 10 is used to collect and obtain ground-based lidar data X and ERA5 meteorological data; perform normalization processing on the ground-based lidar data X and ERA5 meteorological data after preprocessing to obtain a standard data tensor; The inversion module 20 is used to perform PM2.5 chemical component inversion based on the standard data tensor through an inversion model obtained by pre-training with a Transformer encoder combined with multi-objective particle swarm optimization processing, and output an inversion result; The inversion module 20 includes an inversion model training module 21; The inversion model training module 21 includes an initial inversion module 211 and an optimization and update module 212; The initial inversion module 211 is used to collect and obtain a plurality of historical standard data tensors; perform spatio-temporal modeling on the historical standard data tensors through a pre-trained Transformer encoder and perform prediction output to obtain an initial inversion result ; The optimization and update module 212 is used to optimize the model parameters through multi-objective particle swarm optimization processing based on the initial inversion result and the historical PM2.5 chemical component concentration to obtain optimized model parameters; update the Transformer encoder based on the optimized model parameters to obtain an updated Transformer encoder.
[0013] The initial inversion module 211 is specifically configured to convert the historical standard data tensor through the text embedding layer in the Transformer encoder to obtain a high-dimensional vector; inject spatio-temporal information into the high-dimensional vector to obtain an embedding matrix integrating spatio-temporal positions. ; Based on the embedding matrix integrating spatio-temporal positions process and analyze through the multi-head self-attention mechanism to obtain a normalized embedding matrix ; Based on the normalized embedding matrix perform cross-modal attention fusion processing to output a fused feature ; Based on the fused feature extract high-order features through a feed-forward network (FFN) to obtain an enhanced feature ; Based on the enhanced feature map and output through a fully connected layer to obtain a normalized concentration matrix ; Perform inverse normalization processing on the normalized concentration matrix to obtain an initial inversion result .
[0014] The initial inversion module 211 is further configured to allocate a learnable time embedding vector to the high-dimensional vector per hour to obtain a time embedding matrix ; perform dynamic position encoding processing on the high-dimensional vector of each vertical height layer to obtain a spatial embedding matrix ; Add the time embedding matrix and the spatial embedding matrix element-wise and output to obtain an embedding matrix integrating spatio-temporal positions .
[0015] The initial inversion module 211 is further configured to output a query matrix Q, a key matrix K, and a value matrix V through linear transformation based on the embedding matrix integrating spatio-temporal positions ; Calculate scaled dot-product attention based on the query matrix Q, the key matrix K, and the value matrix V ; ; The calculation method of the scaled dot-product attention is: ; where is the query matrix, is the key matrix, is the value matrix, is the dimension of the key vector; is the activation function; Based on the scaled dot product attention and the embedding matrix that fuses spatio-temporal positions perform a residual connection to obtain the matrix after the residual connection ; Normalize the matrix after the residual connection to obtain the normalized embedding matrix .
[0016] The initial inversion module 211 is also used to perform feature splitting on the normalized embedding matrix to obtain the feature vector of the ground-based lidar and the feature vector of the ERA5 meteorological data ; Based on the feature vector of the ground-based lidar and the feature vector of the ERA5 meteorological data calculate the cross-modal weight ; The calculation method of the cross-modal weight is as follows: ; wherein, is the transpose of the weight matrix, is the feature vector of the lidar data, is the feature vector of the ERA5 meteorological data, is the sum of the exponential functions for position j; Based on the cross-modal weight weightedly fuse the feature vector of the ground-based lidar and the feature vector of the ERA5 meteorological data to obtain the fused feature .
[0017] The optimization and update module 212 is specifically used to obtain an initial particle swarm; each particle in the initial particle swarm represents a set of model hyperparameters; the model hyperparameters include the learning rate, the number of attention heads, the number of Transformer layers, and the batch size; Input the historical standard data tensor into the pre-trained Transformer encoder obtained based on the input of each particle, and output the initial inversion result ; Calculate the Pearson correlation coefficient CORR and the root mean square error RMSE based on the initial inversion result and the historical PM2.5 chemical component concentration ; The calculation method of the Pearson correlation coefficient CORR is as follows: ; The calculation method of the root mean square error RMSE is as follows: ; where is the historical PM2.5 chemical component concentration of the -th data point, is the average value of all historical PM2.5 chemical component concentrations, is the initial inversion result of the -th data point, is the average value of all initial inversion results, is the number of samples; CORR is the correlation coefficient, and RMSE is the root mean square error; Based on the Pearson correlation coefficient CORR and the root mean square error RMSE, a comprehensive score is calculated; it is judged whether the comprehensive score is greater than or equal to the comprehensive score threshold; if so, the particle corresponding to the comprehensive score is output as the candidate optimized model parameter; if not, the particle swarm is updated to obtain a new particle swarm, and the new particle swarm is returned to the above operation until the candidate optimized model parameter is output; Based on all the candidate optimized model parameters, combined with the Pareto front screening process, the optimized model parameter is output.
[0018] The optimization update module 212 is further configured to record the position corresponding to the comprehensive score of each particle at each iteration to obtain the position corresponding to the single-particle historical comprehensive score; select the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical comprehensive score to obtain the position corresponding to the single-particle historical optimal comprehensive score; Select the position corresponding to the single-particle historical optimal comprehensive score with the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all particles as the position corresponding to the global optimal historical comprehensive score; Record the moving speed of each particle at time t; calculate the new moving speed of the particle at time t + 1 based on the moving speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; The new moving speed is calculated as follows: ; where is the inertia weight, , are the learning factors , and are random numbers uniformly distributed between [0, 1], is the position corresponding to the single-particle historical optimal comprehensive score, is the position corresponding to the global optimal historical comprehensive score, is the moving speed of the particle at time t, is the moving speed of the particle at time t + 1, is the position corresponding to the comprehensive score of the particle at time t; Based on the new moving speeds of the respective particles and the positions of the particles at time t, new particles are calculated; based on all the new particles, a new particle swarm is obtained.
[0019] The optimization and update module 212 is further configured to record, at each iteration, the position corresponding to the comprehensive score of each particle, to obtain the position corresponding to the single-particle historical comprehensive score; select the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical comprehensive score to obtain the position corresponding to the single-particle historical optimal comprehensive score; select the position corresponding to the single-particle historical optimal comprehensive score with the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all the particles as the position corresponding to the global optimal historical comprehensive score; Record the Pearson correlation coefficient difference △CORR and the root mean square error difference △RMSE at each time t; calculate an adjustment coefficient z based on the Pearson correlation coefficient difference △CORR and the root mean square error difference △RMSE; The calculation method of the adjustment coefficient z is: ; In the formula, is the range normalization function (mapping the input to the interval [0, 1]); Record the moving speed of each particle at time t; calculate the new moving speed of the particle at time t + 1 based on the adjustment coefficient z, the moving speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; The new moving speed is calculated as follows: ; where is the inertia weight, , are learning factors , and and are random numbers uniformly distributed between [0, 1], is the position corresponding to the single-particle historical optimal comprehensive score, is the position corresponding to the global optimal historical comprehensive score, is the moving speed of the particle at time t, is the moving speed of the particle at time t + 1, is the position corresponding to the comprehensive score of the particle at time t; Based on the new moving speed of each particle and the position of the particle at time t to calculate new particles; new particle swarms are obtained based on all the new particles.
[0020] The optimization and update module 212 is further configured to map each of the to-be-selected optimized model parameters to obtain two-dimensional coordinate points; construct a kd-tree structure based on the two-dimensional coordinate points; The abscissa of the two-dimensional coordinate point is the Pearson correlation coefficient CORR of the to-be-selected optimized model parameter; the ordinate of the two-dimensional coordinate point is the root mean square error RMSE of the to-be-selected optimized model parameter; Traverse each candidate solution in the kd-tree structure, and obtain a plurality of neighborhood candidate solutions corresponding to the candidate solution in the kd-tree structure by means of k-nearest neighbor search for the candidate solution; Perform a dominance judgment on the candidate solution and the neighborhood candidate solutions corresponding to the candidate solution to obtain the solutions on the first layer of the Pareto front; The method of the dominance judgment is that when the Pearson correlation coefficient CORR of the candidate solution is greater than the Pearson correlation coefficient CORR of any one of the neighborhood candidate solutions of the candidate solution, and the root mean square error RMSE of the candidate solution is less than the root mean square error RMSE of any one of the neighborhood candidate solutions of the candidate solution, it is determined that the candidate solution is the solution on the first layer of the Pareto front; Screen out all the solutions on the first layer of the Pareto front from the candidate solutions, and repeat the above operations until all candidate solutions are classified to obtain a set of solutions on multiple layers of the Pareto front; Sort the solutions on each layer of the Pareto front in descending order according to the value of the Pearson correlation coefficient CORR to obtain a sequence set of the solutions on each layer of the Pareto front; calculate the crowding distance for each solution on each layer of the Pareto front based on the sequence set of the solutions on the Pareto front ; The crowding distance The calculation method is as follows: ; In the formula, is the crowding distance of the i-th solution on the Pareto front of the current layer; i is the i-th solution on the sequence set of the solutions on the Pareto front of the current layer; is the Pearson correlation coefficient of the (i + 1)-th solution on the sequence set of the solutions on the Pareto front of the current layer; is the Pearson correlation coefficient of the (i - 1)-th solution on the sequence set of the solutions on the Pareto front of the current layer; is the maximum value of the Pearson correlation coefficients of all the solutions on the sequence set of the solutions on the Pareto front of the current layer; is the minimum Pearson correlation coefficient of all Pareto front solutions in the sequence set of solutions on the Pareto front of the current layer; is the root mean square error of the (i + 1)-th Pareto front solution in the sequence set of solutions on the Pareto front of the current layer; is the root mean square error of the (i - 1)-th Pareto front solution in the sequence set of solutions on the Pareto front of the current layer; is the maximum root mean square error of all Pareto front solutions in the sequence set of solutions on the Pareto front of the current layer; is the minimum root mean square error of all Pareto front solutions in the sequence set of solutions on the Pareto front of the current layer; Based on the crowding degree Select multiple optimized model parameters in sequence according to the levels of the solutions on the Pareto front.
[0021] Embodiment 2 Such as Figure 2 And Figure 3 As shown, Embodiment 2 of the present invention provides a method for inverting the vertical profile of PM2.5 chemical components based on Transformer and MOPSO, including the following operation steps: Step S10: Collect the ground-based lidar data X and ERA5 meteorological data of 60 vertical height layers with a time resolution of 1 hour; preprocess the ground-based lidar data to obtain the preprocessed ground-based lidar data ; at the same time, preprocess the ERA5 meteorological data to obtain the preprocessed ERA5 meteorological data; perform normalization processing on the preprocessed ground-based lidar data and the preprocessed ERA5 meteorological data to obtain a standard data tensor (the standard data tensor is the normalized ground-based lidar data and the normalized thickness ERA5 meteorological data tensor, presented as 60 layers × 15 features, as the input information of the Transformer encoder in the subsequent operation process); It should be noted that in the above embodiment of the present application, the ground-based lidar data of 60 vertical height layers (each layer is 100 meters, that is, the ground-based lidar data of a vertical height of 6000 meters) is collected with a time resolution of 1 hour by using a lidar with a wavelength of 532 nm, including the aerosol backscattering coefficient , extinction coefficient and the vertical profile data of the depolarization ratio δ; further, preprocess the ground-based lidar data, that is, use wavelet threshold denoising processing. In the specific operation process, first decompose the ground-based lidar data into 3 layers by using the Daubechies-4 wavelet basis to extract the high-frequency noise components; then introduce a hard threshold λ The high-frequency noise components are filtered out, and then, after filtering out the high-frequency noise components, reconstruction is performed to output the denoised ground lidar data (i.e., the preprocessed ground lidar data as described above); The calculation method for filtering out the high-frequency noise components is as follows: ; is the k-th wavelet basis function, X is the ground lidar data, λ is the hard threshold parameter, is the hard threshold function, is the preprocessed ground lidar data; Further, the ERA5 meteorological data in the embodiments of the present application includes parameters such as integrated temperature, humidity, wind speed, boundary layer height, etc., with a time resolution of 1 hour and a spatial resolution of 0.25°×0.25°; By preprocessing the ERA5 meteorological data, the preprocessed ERA5 meteorological data is obtained, which is obtained by filling and aligning the ERA5 meteorological data and outputting the preprocessed ERA5 meteorological data; specifically: using the GAN framework, first inputting random noise z, generating missing meteorological parameters (the missing meteorological parameters are the predicted meteorological parameters, i.e., the predicted temperature, humidity, wind speed, etc. that need to be filled) through a fully connected layer (512→256→128); further, analyzing and outputting the authenticity probability of the ERA5 meteorological data by using three layers of convolution and a fully connected layer, and then calculating the loss function of the authenticity probability and the missing meteorological parameters through the Wasserstein distance. When the loss function is the lowest or reaches the loss function threshold, the filled ERA5 meteorological data (i.e., the preprocessed ERA5 meteorological data) is output; The specific form of the loss function of the above Wasserstein distance is: ; Where is the probability that the input data belongs to the real data, is a sample similar to the real data output by the generator, is the real data distribution, is the noise distribution, X and Z are both sample values, is the loss function (or the Wasserstein distance); Further, the above preprocessed ERA5 meteorological data is interpolated to the same spatio-temporal resolution as the above preprocessed ground lidar data (i.e., a time resolution of 1 hour and 60 vertical height layers) to ensure consistent data dimensions; Step S20: Perform PM2.5 chemical component inversion based on the standard data tensor using an inversion model obtained by pre-training with a Transformer encoder combined with multi-objective particle swarm optimization, and output the inversion result; It should be noted that the above inversion results refer to the vertical profiles of the concentrations of PM2.5 chemical components (sulfate, nitrate, ammonium salt, organic matter, black carbon) at each altitude layer; In the above embodiments of the present application, by using the preprocessed and normalized data, the spatio-temporal dependence is modeled through a Transformer encoder, and the inversion parameters are optimized by multi-objective particle swarm optimization (MOPSO) to output the vertical profiles of PM2.5 chemical components. By integrating multi-source heterogeneous data, global spatio-temporal modeling, and an efficient optimization strategy, high-precision and low-computation-cost vertical pollution distribution monitoring is achieved, which is applicable to the refined management of urban air quality and the tracking of pollution sources; In the above embodiments of the present application, first, ground-based lidar data with 60 vertical altitude layers (each layer is 100 meters, that is, ground-based lidar data with a vertical altitude of 6000 meters) is collected using a 532nm wavelength lidar at a time resolution of 1 hour, and ERA5 meteorological data including parameters such as integrated temperature, humidity, wind speed, and boundary layer height is collected at a time resolution of 1 hour and a spatial resolution of 0.25°×0.25°, providing a rich data basis for subsequent inversion processing; after preprocessing the ground-based lidar data and ERA5 meteorological data for denoising, data filling, and alignment, the preprocessed data (i.e., the above standard data tensor) is used for PM2.5 chemical component inversion through a trained inversion model, and the inversion result is output. By integrating lidar and meteorological data, combining the global attention mechanism and an efficient optimization strategy, problems such as insufficient spatio-temporal modeling and high computational cost of traditional models are solved; specifically, the above trained inversion model models spatio-temporal dependence through a Transformer encoder and the global attention mechanism to achieve spatio-temporal dependence modeling across altitude layers. During this period, through a dynamic data fusion strategy, the collaborative inversion ability of multi-modal data is improved (see steps S2131 - S2133 in detail), and moreover, by using multi-objective particle swarm optimization to optimize the multi-objective algorithm framework, the consumption of computing resources is reduced and the diversity of the Pareto solution set is improved.
[0022] Specifically, in step S20, the training of the inversion model obtained by training using a Transformer encoder combined with multi-objective particle swarm optimization includes the following operating steps: Step S21: Collect and obtain multiple historical standard data tensors; perform spatio-temporal modeling and prediction output based on the historical standard data tensors through a pre-trained Transformer encoder to obtain an initial inversion result ; The historical standard data tensor is obtained by collecting historical ground-based lidar data and historical ERA5 meteorological data and then through preprocessing and normalization processing; the historical standard data tensor includes historical ground-based lidar data, historical ERA5 meteorological data, and historical PM2.5 chemical component concentrations. (In the embodiment of the present application, the final result inversion result of the above step S20 is the PM2.5 chemical component concentration obtained by inverting the currently collected and preprocessed and normalized standard data tensor); The training of the above-mentioned pre-trained Transformer encoder is specifically as follows: input the historical standard data tensor into a pre-established Transformer encoder to output a model predicted concentration distribution; calculate the total loss function based on the model predicted concentration distribution and the prior concentration distribution corresponding to the historical standard data tensor; when the total loss function is less than or equal to the loss function threshold, output the pre-trained Transformer encoder. The calculation method of the loss function is: ; is the total loss function, is the weighted mean square error, is the KL divergence, is the model predicted concentration distribution, is the prior concentration distribution; Step S22: Based on the initial inversion result and the historical PM2.5 chemical component concentrations Optimize the model parameters through multi-objective particle swarm optimization to obtain optimized model parameters; update the Transformer encoder based on the optimized model parameters to obtain an updated Transformer encoder (the updated Transformer encoder is the above-mentioned inversion model trained by combining the Transformer encoder with multi-objective particle swarm optimization).
[0023] It should be noted that in the above embodiment of the present application, global spatio-temporal modeling is performed through the Transformer encoder. Specifically, the long-range dependencies between vertical height layers are captured through the Transformer encoder, and the inversion error is reduced by 12% compared to CNN-ATT-BiLSTM; the Transformer encoder can focus on modeling any inter-layer relationships, thereby reducing the loss of inter-layer information transmission and improving the accuracy. Among them, the introduction of inter-layer relative distance coding can enhance the model's ability to distinguish the relationships between adjacent layers and non-adjacent layers, enhance the modeling of cross-layer relationships, improve the accuracy, and thus reduce the inversion error.
[0024] Specifically, in step S21, spatio-temporal modeling is performed on the historical standard data tensor through a pre-trained Transformer encoder, and a prediction output is obtained to get an initial inversion result , including the following operation steps: Step S211: The historical standard data tensor is transformed through the text embedding layer in the Transformer encoder to obtain a high-dimensional vector; spatio-temporal information is injected into the high-dimensional vector to obtain an embedding matrix integrating spatio-temporal positions ; Step S212: Based on the embedding matrix integrating spatio-temporal positions , it is processed and analyzed through a multi-head self-attention mechanism to obtain a normalized embedding matrix ; Step S213: Based on the normalized embedding matrix , through cross-modal attention fusion processing, a fusion feature is output ; Based on the fusion feature , high-order features are extracted through a feed-forward network (FFN) to obtain enhanced features ; Step S214: Based on the enhanced features , it is mapped and output through a fully connected layer to obtain a normalized concentration matrix ; Through inverse normalization processing on the normalized concentration matrix , an initial inversion result is obtained ; It should be noted that in the above embodiments of the present application, the historical data tensor of 60 layers × 15 features after standardization is first mapped to a high-dimensional space through the embedding layer, and spatio-temporal information (such as time stamps, vertical layer numbers, spatial coordinates) is explicitly injected to form an embedding matrix integrating spatio-temporal features, and structured data (lidar features, meteorological parameters) is converted into dense vectors to eliminate the difference in feature dimensions and enhance the non-linear representation ability. At the same time, through learnable position encoding (such as time step encoding, vertical layer position encoding) or directly splicing spatio-temporal labels (number of hours, height layer serial number), the model can perceive the time series and vertical hierarchical structure of the data, laying a foundation for subsequent spatio-temporal dependence modeling; the high-dimensional embedding retains the physical meaning of the original data, and at the same time the spatio-temporal encoding enables the model to distinguish the spatio-temporal correlation of data at different height layers and different times; further, the multi-head self-attention mechanism is used to analyze the embedding matrix to capture the global dependence relationship across height layers and time steps, and the training is stabilized through normalization. Among them, each "head" independently calculates the attention weights in different sub-spaces, such as Figure 4As shown, by focusing on complex patterns such as the dynamic interaction between layers (e.g., the impact of bottom - layer aerosols on high - layer diffusion) and the correlation between meteorological conditions and pollutant concentrations, layer normalization (LayerNorm) is performed on the attention output to alleviate the vanishing / exploding gradients and accelerate model convergence; the Transformer encoder can be used to focus on modeling any inter - layer relationship, thereby reducing the loss of information transfer between layers and improving accuracy; further, by designing a cross - modal attention mechanism, the potential correlation between lidar data (aerosol optical properties) and meteorological data (temperature, humidity, wind speed) is dynamically aligned. For example, through the query - key - value mechanism, meteorological parameters are used as queries to retrieve key information in lidar features to generate fused features; high - order features are extracted through a feed - forward network (FFN) to obtain enhanced features, and the high - order features are mapped to the normalized concentration space corresponding to the target chemical components (5 types such as sulfate), and a normalized matrix of 60 layers × 5 components is output; according to the mean and standard deviation statistics in the training stage, the normalized concentration is restored to the actual physical dimension to generate the initial inversion result (the concentration of PM2.5 components in each layer).
[0025] Specifically, in step S211, temporal - spatial information is injected into the high - dimensional vector to obtain an embedding matrix that fuses temporal - spatial positions , including the following operation steps: Step S2111: Assign learnable temporal embedding vectors to the high - dimensional vectors for each hour to obtain a temporal embedding matrix ; Step S2112: Process the high - dimensional vectors for each vertical height layer through dynamic positional encoding to obtain a spatial embedding matrix ; It should be noted that since the Transformer model processes in parallel, positional encoding is required to capture the order information of elements in the sequence. Dynamic positional encoding uses sine - function encoding and cosine - function encoding to represent different positions.
[0026] Step S2113: Add the temporal embedding matrix and the spatial embedding matrix element - by - element, and output the embedding matrix that fuses temporal - spatial positions .
[0027] It should be noted that in the above - mentioned embodiment of the present application, the cross - modal attention mechanism is used to improve the multi - source data fusion efficiency, and the feature weight allocation speed is increased by 30%.
[0028] Specifically, in step S212, based on the embedding matrix that fuses temporal - spatial positions process and analyze it through a multi - head self - attention mechanism to obtain a normalized embedding matrix , including the following operation steps: Step S2121: Based on the embedding matrix that fuses spatio-temporal positions Through linear transformation, output the query matrix Q, the key matrix K, and the value matrix V; It should be noted that in the above embodiments of the present application, through linear transformation, is divided into 8 attention heads (head dimension 32), generating the query matrix Q, the key matrix K, and the value matrix V (each 60×8×32); Step S2122: Calculate the scaled dot-product attention based on the query matrix Q, the key matrix K, and the value matrix V ; The calculation method of the scaled dot-product attention is: ; where Q is the query matrix, K is the key matrix, V is the value matrix, is the dimension of the key vector; is the activation function; It should be noted that in the above embodiments of the present application, by multiple attention heads simultaneously focusing on different parts of the input sequence, capturing complex dependencies in the sequence. Each attention head independently calculates the query, key, and value representations of the input sequence and concatenates the results; Step S2123: Based on the scaled dot-product attention and the embedding matrix that fuses spatio-temporal positions Perform a residual connection to obtain the matrix after the residual connection ; The matrix after the residual connection The calculation method is: ; Step S2124: Normalize the matrix after the residual connection to obtain the normalized embedding matrix ; It should be noted that in the above embodiments of the present application, by normalizing each layer of features of the matrix after the residual connection to make its mean 0 and variance 1, output the normalized embedding matrix , stabilizing the training process and accelerating model convergence.
[0029] Specifically, in step S213, based on the normalized embedding matrix Through cross-modal attention fusion processing, output the fused feature , including the following operation steps: Step S2131: Split the features of the normalized embedding matrix to obtain the feature vector of the ground lidar and the eigenvectors of ERA5 meteorological data ; Step S2132: Based on the eigenvectors of the ground-based lidar and the eigenvectors of ERA5 meteorological data calculate the cross-modal weights ; The calculation method of the cross-modal weights is as follows: ; wherein, is the transpose of the weight matrix, is the eigenvector of the lidar data, is the eigenvector of ERA5 meteorological data, is the summation of the exponential function at position j; Step S2133: Based on the cross-modal weights weightedly fuse the eigenvectors of the ground-based lidar and the eigenvectors of ERA5 meteorological data to obtain the fused features ; The calculation method of the fused features is as follows: . It should be noted that in the above embodiments of the present application, spatio-temporal information is gradually fused through a multi-layer encoder. For example, each layer includes a self-attention module and a feed-forward neural network (FFN). The training stability can be improved through residual connections and layer normalization (LayerNorm). Among them, the shallow encoder captures local spatio-temporal patterns (such as object edges, short-term movements), and the deep encoder models global semantic relationships (such as object interactions, long-term temporal dependencies).
[0030] Specifically, in step S22, based on the initial inversion result and the historical PM2.5 chemical component concentrations optimize the model parameters through multi-objective particle swarm optimization to obtain the optimized model parameters, including the following operation steps: Step S221: Obtain the initial particle swarm; each particle in the initial particle swarm represents a set of model hyperparameters; the model hyperparameters include the learning rate, the number of attention heads, the number of Transformer layers, and the batch size; Step S222: Input the historical standard data tensor into the pre-trained Transformer encoder obtained based on each particle's input, and output the initial inversion result ; calculate the Pearson correlation coefficient CORR and the root mean square error RMSE based on the initial inversion result and the historical PM2.5 chemical component concentrations ; The calculation method of the Pearson correlation coefficient CORR is as follows: ; The calculation method of the root mean square error RMSE is as follows: ; where is the historical PM2.5 chemical component concentration of the th data point, is the average value of all historical PM2.5 chemical component concentrations, is the initial inversion result of the th data point, is the average value of all initial inversion results, N is the number of samples; CORR is the correlation coefficient, and RMSE is the root mean square error; Step S223: Calculate a comprehensive score based on the Pearson correlation coefficient CORR and the root mean square error RMSE; Determine whether the comprehensive score is greater than or equal to the comprehensive score threshold; If so, output the particle corresponding to the comprehensive score as the candidate optimized model parameter; If not, update the particle swarm to obtain a new particle swarm, and return the new particle swarm to the above operation (that is, use the new particle swarm as the initial particle swarm in Step S221 and re - perform the iterative process) until the candidate optimized model parameter is output; Step S224: Based on all the candidate optimized model parameters, perform screening processing in combination with the Pareto front to output the optimized model parameter.
[0031] It should be noted that the MOPSO algorithm in the embodiment of the present application reduces the calculation time by 35% compared with NSGA - II, and the diversity of the Pareto solution set is increased by 20%. As Figure 5 shown, the particle swarm optimization algorithm uses the sharing of information by individuals in the group to make the movement of the entire group evolve from disorder to order in the problem - solving space, thereby obtaining the optimal solution. Pareto is a solution set. How to select the most satisfactory result from it according to a certain rational criterion will be the key to solving this type of multi - objective optimization problem. Combining these two methods in this solution can adapt to complex objective spaces and dynamic environments and improve the convergence speed. While considering the diversity of the solution set, the convergence accuracy can also be guaranteed.
[0032] Specifically, in Step S223, updating the particle swarm to obtain a new particle swarm includes the following operation steps: Step S2231: Record the position corresponding to the comprehensive score of each particle during each iteration to obtain the position corresponding to the single - particle historical comprehensive score; Select the position corresponding to the single - particle historical comprehensive score with the highest single - particle historical comprehensive score to obtain the position corresponding to the single - particle historical optimal comprehensive score; Step S2232: Select the position corresponding to the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all particles; the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical optimal comprehensive score is the position corresponding to the global optimal historical comprehensive score; Step S2233: Record the moving speed of each particle at time t; calculate the new moving speed of the particle at time t + 1 based on the moving speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; The above-mentioned new moving speed is calculated as follows: ; where is the inertia weight, , are learning factors , and are random numbers uniformly distributed between [0, 1], is the position corresponding to the single-particle historical optimal comprehensive score, is the position corresponding to the global optimal historical comprehensive score, is the moving speed of the particle at time t, is the moving speed of the particle at time t + 1 (i.e., the above-mentioned new moving speed), is the position corresponding to the comprehensive score of the particle at time t; Step S2234: Calculate new particles based on the above-mentioned new moving speed of each particle and the position of the particle at time t (the new particles are the positions of the new particles, that is, the positions where the particles corresponding to the updated model hyperparameters are mapped to the two-dimensional plane); obtain a new particle swarm based on all the new particles; It should be noted that the above-mentioned embodiments of the present application optimize the multi-objective algorithm framework by updating the particle swarm, reduce the consumption of computing resources, and improve the diversity of the Pareto solution set.
[0033] In the specific implementation process of the above-mentioned embodiments of the present application, the technical personnel found that the Transformer model is highly sensitive to hyperparameters such as the learning rate and the number of layers, and small adjustments may cause drastic fluctuations in the verification performance (such as the training divergence caused by too large a learning rate); this will cause the calculation of the new moving speed in the above-mentioned step S2233 , resulting in the new moving speed Fluctuations occur, thereby reducing the robustness of the subsequent new particle swarm. For example, the initial parameters are: learning rate lr = 0.001, number of Transformer layers layers = 4. At this time, the model CORR = 0.85 and RMSE = 0.12. When the learning rate is adjusted to lr = 0.1, with other parameters unchanged, it results in a severe oscillation of the loss function during training, the CORR drops sharply to 0.40, the RMSE soars to 0.50, the model fails to converge, the inversion result completely fails, and the new movement speed changes greatly, which is obviously unreasonable.
[0034] Specifically, in step S223, the particle swarm is updated to obtain a new particle swarm, including the following operation steps: Step S2231’: Record the position corresponding to the comprehensive score of each particle at each iteration to obtain the position corresponding to the single-particle historical comprehensive score; select the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical comprehensive score to obtain the position corresponding to the single-particle historical optimal comprehensive score; Step S2232’: Select the position corresponding to the single-particle historical optimal comprehensive score with the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all particles as the position corresponding to the global optimal historical comprehensive score; Step S2233’: Record the Pearson correlation coefficient difference △CORR and root mean square error difference △RMSE at each moment t; calculate the adjustment coefficient z based on the Pearson correlation coefficient difference △CORR and root mean square error difference △RMSE; The calculation method of the adjustment coefficient z is: ; In the formula, is the range normalization function (mapping the input to the [0,1] interval); Step S2234’: Record the movement speed of each particle at time t; calculate the new movement speed of the particle at time t + 1 based on the adjustment coefficient z, the movement speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; The new movement speed The calculation method is: ; where is the inertia weight, , are the learning factors , and are random numbers uniformly distributed between [0,1], is the position corresponding to the single-particle historical optimal comprehensive score, is the position corresponding to the globally optimal historical comprehensive score, is the moving speed of the particle at time t, is the moving speed of the particle at time t + 1 (i.e., the above-mentioned new moving speed), is the position corresponding to the comprehensive score of the particle at time t; Step S2235': Based on the new moving speeds of the particles and the positions of the particles at time t, calculate to obtain new particles (the new particles are the positions of the new particles, that is, the positions where the particles corresponding to the updated model hyperparameters are mapped to the two-dimensional plane); based on all the new particles, obtain a new particle swarm; It should be noted that the above embodiments of the present application optimize the multi-objective algorithm framework by updating the particle swarm, reduce the consumption of computing resources, and improve the diversity of the Pareto solution set; Based on the technical solutions adopted in the above embodiments of the present application, for example, Example 1, for the problem of Transformer hyperparameter sensitivity (training divergence caused by sudden change in learning rate), a certain particle has a large learning rate setting (such as suddenly increasing from 0.001 to 0.1) during the iteration process, resulting in severe fluctuations in CORR and RMSE: The iteration history (time t = 1 to t = 4) is shown in Table 1 below: Calculate the adjustment coefficient z: 1. Cumulative absolute value of fluctuations: ; 2. Average fluctuation: ; 3. Range normalization: ; 4. Adjustment coefficient : ; Based on the above example, when the sudden change in the learning rate causes a sudden drop in CORR and a sudden increase in RMSE, z = 0.37 significantly reduces the speed update amplitude, inhibits the particle from further moving to the high-risk area, and avoids training divergence; In summary, the technical solution adopted in the embodiments of the present application adds an adjustment coefficient z to the calculation of the original particle update speed, obtains a new calculation method for the particle update speed. When the hyperparameter adjustment causes severe fluctuations, z automatically reduces the speed update amplitude, inhibits the risk of training divergence (such as the problem of too large learning rate), so that the new particle moving speed obtained from the update speed of the particle suppresses fluctuations and is more robust.
[0035] During the execution of step S224 in the embodiment of the present application above, optimized model parameters are obtained through Pareto front screening. Performing non-dominated sorting on particles in each round of update (especially in the case of a large-scale population or multiple objectives) will bring a high computational burden, thereby reducing the running efficiency of the entire algorithm.
[0036] Specifically, in step S224, based on all the candidate optimized model parameters and combined with Pareto front screening, optimized model parameters are output, including the following operating steps: Step S2241: Map each of the candidate optimized model parameters to obtain two-dimensional coordinate points; construct a kd-tree structure based on the two-dimensional coordinate points; The abscissa of the two-dimensional coordinate point is the Pearson correlation coefficient CORR of the candidate optimized model parameter; the ordinate of the two-dimensional coordinate point is the root mean square error RMSE of the candidate optimized model parameter. It should be noted that in the above embodiment of the present application, each candidate optimized model parameter has two objective index performances, namely the Pearson correlation coefficient CORR and the root mean square error RMSE; in the above embodiment of the present application, by using these two index performances as the abscissa and ordinate of the two-dimensional coordinate point, two-dimensional coordinate points of each candidate optimized model parameter are obtained, and then each two-dimensional coordinate point is stored in the kd-tree structure to obtain the kd-tree structure for subsequent rapid local nearest neighbor query; in the embodiment of the present application, by mapping each candidate solution to the two-dimensional coordinate and using the KD tree to complete the spatial index, the traversal range during subsequent non-dominated comparison and crowding distance calculation can be greatly reduced, improving the efficiency; Step S2242: Traverse each candidate solution (the candidate solution refers to the above two-dimensional coordinate point, that is, each candidate optimized model parameter) in the kd-tree structure, and obtain multiple neighborhood candidate solutions corresponding to the candidate solution in the kd-tree structure by means of k-nearest neighbor search; Step S2243: Through dominance judgment on the candidate solution and the neighborhood candidate solutions corresponding to the candidate solution, obtain the solutions on the first layer of the Pareto front; The method of dominance judgment is that when the Pearson correlation coefficient CORR of the candidate solution is greater than the Pearson correlation coefficient CORR of any one of the neighborhood candidate solutions of the candidate solution, and the root mean square error RMSE of the candidate solution is less than the root mean square error RMSE of any one of the neighborhood candidate solutions of the candidate solution, it is determined that the candidate solution is the solution on the first layer of the Pareto front; Step S2244: Screen out all the solutions on the first layer of the Pareto front from the candidate solutions, and repeat the above operations until all candidate solutions are classified to obtain a set of solutions on multiple layers of the Pareto front; Step S2245: Sort the solutions of the Pareto frontiers of each layer in descending order according to the value of the Pearson correlation coefficient CORR to obtain a sequence set of the solutions of the Pareto frontiers of each layer; calculate the crowding distance for each solution of the Pareto frontiers of each layer based on the sequence set of the solutions of the Pareto frontiers ; The calculation method of the crowding distance is as follows: ; In the formula, is the crowding distance of the i-th solution of the Pareto frontier of the current layer; i is the i-th solution of the Pareto frontier of the sequence set of the solutions of the Pareto frontier of the current layer; is the Pearson correlation coefficient of the (i + 1)-th solution of the Pareto frontier of the sequence set of the solutions of the Pareto frontier of the current layer; is the Pearson correlation coefficient of the (i - 1)-th solution of the Pareto frontier of the sequence set of the solutions of the Pareto frontier of the current layer; is the maximum value of the Pearson correlation coefficients of all solutions of the Pareto frontier in the sequence set of the solutions of the Pareto frontier of the current layer; is the minimum value of the Pearson correlation coefficients of all solutions of the Pareto frontier in the sequence set of the solutions of the Pareto frontier of the current layer; is the root mean square error of the (i + 1)-th solution of the Pareto frontier of the sequence set of the solutions of the Pareto frontier of the current layer; is the root mean square error of the (i - 1)-th solution of the Pareto frontier of the sequence set of the solutions of the Pareto frontier of the current layer; is the maximum value of the root mean square errors of all solutions of the Pareto frontier in the sequence set of the solutions of the Pareto frontier of the current layer; is the minimum value of the root mean square errors of all solutions of the Pareto frontier in the sequence set of the solutions of the Pareto frontier of the current layer; Step S2246: Based on the crowding degree select multiple optimized model parameters in sequence according to the levels of the solutions of the Pareto frontier.
[0037] It should be noted that in the above embodiments of the present application, the larger the crowding degree is, the more dispersed its distribution is in the kd-tree structure (or the more dispersed the distribution of its corresponding two-dimensional coordinate points and the two-dimensional coordinate points of its neighborhood candidate solutions is), the more representative it is, and further, it is more likely to be selected as the optimized model parameter for subsequent processing; Based on this, by selecting in sequence according to the levels of the solutions of the Pareto frontier, the optimized model parameters are obtained, that is, the solutions of the Pareto solution set of the first layer are preferentially selected, and then the solutions of the Pareto solution set of the second layer are selected, and so on.
[0038] It should be noted that the above embodiments of the present application combine Pareto front screening with kd-tree spatial indexing to optimize the computational efficiency of non-dominated sorting in traditional multi-objective optimization. Traditional Pareto front screening (such as NSGA-II) usually judges dominance by traversing all candidate solutions, and the computational complexity is O(N²). However, in this solution, candidate solutions are mapped to two-dimensional coordinate points and a kd-tree is constructed. K-nearest neighbor search is used to convert global comparison into local neighborhood comparison (the computational complexity is reduced to O(NlogN)), which can significantly reduce the computational burden. Further, in the embodiments of the present application, step S2242 is adopted to dynamically determine neighborhood candidate solutions by k-nearest neighbor search instead of a fixed neighborhood range. This dynamic neighborhood strategy adapts to the distribution density of candidate solutions, avoids missing key solutions in sparse regions or repeated calculations in dense regions, and improves the integrity of the Pareto front. Further, the present application introduces normalization processing (scaling by the maximum value and the minimum value) through the crowding distance formula in step S2245 to optimize the problem that traditional crowding distance fails in high-dimensional or large-dimensional difference situations. For example, if the dimensional differences between CORR and RMSE are significant (such as CORR ∈ [0, 1], RMSE ∈ [0, 100]), traditional crowding distance will be dominated by RMSE. This solution makes the weights of the two objectives balanced in the calculation through normalization, enhancing the robustness of the algorithm to diversity maintenance.
[0039] According to the technical solutions adopted in the above embodiments of the present application, after implementation, the researchers obtained the following experimental records: System Deployment Hardware: NVIDIA A100 GPU server, supporting FP16 mixed-precision training.
[0040] Software: PyTorch framework integrated with Hugging Face Transformer library, and the MOPSO algorithm is implemented based on the DEAP library.
[0041] Test Results 1. Experimental Design Dataset: Lidar and meteorological data in the Beijing-Tianjin-Hebei region in 2022, including sandstorm (March) and haze (December) events.
[0042] 2. Comparison Baseline: CNN-ATT-BiLSTM-NSGA model and WRF-Chem chemical transport model in the prior art.
[0043] The performance indicators are shown in Table 2 below:
[0044] 3. Result Analysis Precision Advantage: The inversion error of the high-altitude OM concentration during sandstorm events is reduced by 18% by the Transformer encoder.
[0045] Efficiency Improvement: Each iteration of the MOPSO algorithm takes 0.4 seconds, which is 33% faster than NSGA-II (0.6 seconds).
[0046] In summary, a PM2.5 chemical component vertical profile inversion system and method proposed in the embodiments of the present invention first uses a 532nm wavelength lidar to collect ground-based lidar data with 60 vertical height layers at a time resolution of 1 hour and ERA5 meteorological data including parameters such as integrated temperature, humidity, wind speed, and boundary layer height collected at a time resolution of 1 hour and a spatial resolution of 0.25°×0.25°, providing a rich data basis for subsequent inversion processing; after preprocessing the ground-based lidar data and ERA5 meteorological data for denoising, data filling, and alignment, the preprocessed data (i.e., the above standard data tensor) is used to invert the PM2.5 chemical components through a trained inversion model, and the inversion result is output. By fusing lidar and meteorological data and combining the global attention mechanism and the efficient optimization strategy, problems such as insufficient spatio-temporal modeling and high computational costs of traditional models are solved; specifically, the above-trained inversion model models spatio-temporal dependencies through the Transformer encoder and the global attention mechanism to achieve spatio-temporal dependency modeling across height layers. During this period, through the dynamic data fusion strategy, the collaborative inversion ability of multi-modal data is improved, and furthermore, by using the multi-objective particle swarm optimization process to optimize the multi-objective algorithm framework, the consumption of computing resources is reduced and the diversity of the Pareto solution set is improved; Furthermore, in view of the problem that the Transformer model is highly sensitive to hyperparameters such as the learning rate and the number of layers during the multi-objective particle swarm optimization algorithm process, and small adjustments may lead to drastic fluctuations in the verification performance, and thus lead to fluctuations in the update of the new movement speed of the particles, the embodiments of the present application introduce an adjustment coefficient z to obtain a new calculation method for the particle update speed. When the hyperparameter adjustment causes drastic fluctuations, z automatically reduces the speed update amplitude, suppresses the risk of training divergence, makes the update speed of the particles obtain a new particle movement speed that suppresses fluctuations, and is more robust; At the same time, in view of the problem of high computing power pressure in the large population processing during the Pareto front screening process, the Pareto front screening is combined with the kd-tree spatial index to optimize the computational efficiency of non-dominated sorting in traditional multi-objective optimization.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; those of ordinary skill in the art can modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO, characterized in that It includes a basic data acquisition module and an inversion module; The basic data acquisition module is used to collect and obtain ground-based lidar data X and ERA5 meteorological data; the ground-based lidar data X and ERA5 meteorological data are preprocessed and then normalized to obtain a standard data tensor; The inversion module is used to perform PM2.5 chemical component inversion based on the standard data tensor through an inversion model obtained by pre-training with a Transformer encoder combined with multi-objective particle swarm optimization, and output an inversion result.
2. The PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 1, wherein, The inversion module includes an inversion model training module; The inversion model training module includes an initial inversion module and an optimization and update module; The initial inversion module is used to collect and obtain multiple historical standard data tensors; based on the historical standard data tensors, spatio-temporal modeling is performed through a pre-trained Transformer encoder and prediction output is carried out to obtain an initial inversion result ; The optimization and update module is used to based on the initial inversion result and the historical PM2.5 chemical component concentration Optimize the model parameters through multi-objective particle swarm optimization to obtain the optimized model parameters; update the Transformer encoder based on the optimized model parameters to obtain the updated Transformer encoder.
3. The PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 2, wherein The initial inversion module is specifically configured to convert the historical standard data tensor through the text embedding layer in the Transformer encoder to obtain a high-dimensional vector; inject spatio-temporal information into the high-dimensional vector to obtain an embedding matrix that fuses spatio-temporal positions ; The embedding matrix based on the fused spatio-temporal position The normalized embedding matrix is obtained by processing and analyzing through the multi-head self-attention mechanism ; Based on the normalized embedding matrix Through cross-modal attention fusion processing, the fused features are output ; Based on the fused features High-order features are extracted through a feed-forward network to obtain enhanced features ; Based on the enhanced features The normalized concentration matrix is obtained by mapping the output through a fully connected layer ; For the normalized concentration matrix Through inverse normalization processing, an initial inversion result is obtained .
4. A PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 3, characterized in that, The initial inversion module is further configured to assign learnable time embedding vectors to the high-dimensional vectors for each hour to obtain a time embedding matrix ; The high-dimensional vectors of each vertical height layer are processed by dynamic positional encoding to obtain a spatial embedding matrix ; Embed the time into the matrix with the spatial embedding matrix element-wise added to output an embedding matrix that fuses the spatio-temporal positions .
5. The PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 4, characterized in that The initial inversion module is further configured to output a query matrix Q, a key matrix K, and a value matrix V through linear transformation based on the embedding matrix integrating spatio-temporal positions. Calculating scaled dot - product attention based on the query matrix Q, the key matrix K, and the value matrix V ; Based on the scaled dot - product attention and the embedding matrix that fuses spatio - temporal positions perform residual connection to obtain the matrix after residual connection ; For the matrix after the residual connection perform normalization to obtain the normalized embedding matrix .
6. The PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 5, characterized in that, The initial inversion module is further configured to perform feature splitting on the normalized embedding matrix to obtain the feature vector of the ground-based lidar and the feature vector of the ERA5 meteorological data ; Feature vectors based on the ground lidar and the feature vectors of ERA5 meteorological data Calculate the cross-modal weights ; Based on the cross-modal weights weight the feature vector of the ground lidar and the feature vector of ERA5 meteorological data to perform weighted fusion and obtain a fused feature .
7. A PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 6, characterized in that, The optimization and update module is specifically used to obtain an initial particle swarm; each particle in the initial particle swarm represents a set of model hyperparameters; the model hyperparameters include a learning rate, the number of attention heads, the number of Transformer layers, and a batch size; Input the historical standard data tensor into the pre-trained Transformer encoder obtained based on each particle input, and output the initial inversion result. Based on the initial inversion result and the historical PM2.5 chemical component concentration Calculate the Pearson correlation coefficient CORR and the root mean square error RMSE. Calculate a comprehensive score based on the Pearson correlation coefficient CORR and the root mean square error RMSE; Judge whether the comprehensive score is greater than or equal to a comprehensive score threshold; if so, output the particle corresponding to the comprehensive score as the candidate optimized model parameter; if not, update the particle swarm to obtain a new particle swarm, and return the new particle swarm to the above operation until the candidate optimized model parameter is output; Based on all the candidate optimized model parameters, perform a Pareto front screening process to output the optimized model parameters.
8. A PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 7, characterized in that, The optimization and update module is also used to record the position corresponding to the comprehensive score of each particle at each iteration to obtain the position corresponding to the single-particle historical comprehensive score; select the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical comprehensive score to obtain the position corresponding to the single-particle historical optimal comprehensive score; Select the position corresponding to the single-particle historical optimal comprehensive score with the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all particles as the position corresponding to the global optimal historical comprehensive score; Record the moving speed of each particle at time t; calculate the new moving speed of the particle at time t+1 based on the moving speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; Based on the new moving speed of each particle and the position of the particle at time t, new particles are calculated; a new particle swarm is obtained based on all the new particles.
9. A PM2.5 chemical component vertical profile inversion system based on Transformer and MOPSO according to claim 8, characterized in that, The optimization and update module is also used to record the position corresponding to the comprehensive score of each particle at each iteration to obtain the position corresponding to the single-particle historical comprehensive score; select the position corresponding to the single-particle historical comprehensive score with the highest single-particle historical comprehensive score to obtain the position corresponding to the single-particle historical optimal comprehensive score; Select the position corresponding to the single-particle historical optimal comprehensive score with the highest single-particle historical optimal comprehensive score among the positions corresponding to the single-particle historical optimal comprehensive scores of all particles as the position corresponding to the global optimal historical comprehensive score; Record the Pearson correlation coefficient difference △CORR and the root mean square error difference △RMSE at each moment t; calculate an adjustment coefficient z based on the Pearson correlation coefficient difference △CORR and the root mean square error difference △RMSE; Record the moving speed of each particle at time t; calculate the new moving speed of the particle at time t+1 based on the adjustment coefficient z, the moving speed of the particle at time t, the position corresponding to the single-particle historical optimal comprehensive score, and the position corresponding to the global optimal historical comprehensive score ; Based on the new moving speed of each particle and the position of the particle at time t, new particles are calculated; a new particle swarm is obtained based on all the new particles.
10. A method for inverting the vertical profile of PM2.5 chemical components based on Transformer and MOPSO, characterized in that, It includes the following operation steps: Collect and obtain ground-based lidar data X and ERA5 meteorological data; preprocess the ground-based lidar data X and ERA5 meteorological data and then normalize them to obtain a standard data tensor; Based on the inversion model obtained by pre-training through the combination of a Transformer encoder and multi-objective particle swarm optimization processing using the standard data tensor, PM2.5 chemical component inversion is performed, and the inversion result is output.
Citation Information
Patent Citations
Radar radiation source signal feature selection method based on membrane particle swarm multi-target algorithm
CN107590436A
Method for inverting soil moisture based on optimized BP neural network microwave remote sensing
CN113255874A
Troposphere atmospheric temperature and humidity profile inversion method based on multi-objective genetic algorithm
CN115406911A
Sar-based soil water content inversion method
CN117826112A
Meteorological satellite data processing method based on deep learning
CN119861429A
Cited By
Cross-border area PM2.5 three-dimensional reconstruction method and device
CN122134950A