Method for estimating ozone concentration data at the 100-meter level based on the LG-Transformer model
By using the LG-Transformer model in the estimation of large-area ozone concentrations, combined with grid maps and deep learning technology, the problem of low accuracy of large-area ozone concentration estimation in the existing technology is solved, and high-precision ozone concentration estimation and pollution source positioning are achieved.
Patent Information
- Application Number
- CN202411864533.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-17
AI Technical Summary
It is difficult for the prior art to effectively carry out accurate estimation of large-area ozone concentrations, especially when the data distribution is sparse or complex, the estimation accuracy of inverse distance weighted interpolation and Kriging interpolation methods is low and is not suitable for large-area areas.
The estimation method of ozone concentration at 100 meters based on the LG-Transformer model is adopted. By deeply fusion of the Transformer model and the LightGBM connection layer, the LG-Transformer model is constructed, and the data is resampled and information fusion is combined with grid maps to realize the relationship between spatiotemporal information and ozone concentration and model training.
It significantly improves the precision and practical application value of ozone concentration estimation, can estimate ozone concentration at a resolution of 100 meters, improves the positioning accuracy of pollution sources and the refinement of the spatial distribution of ozone concentration, and enhances the understanding of ozone concentration distribution in large areas.
Smart Images

Figure CN119783886B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ozone concentration model estimation, and particularly to a method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model. Background Art
[0002] In the prior art, to obtain ozone concentration data, an ozone concentration measuring station needs to be set up at a certain location point, and the ozone concentration in the air can be quickly and accurately measured through the ozone concentration measuring station. However, the ozone concentration measuring station is only applicable to the monitoring of ozone concentration at a specific location point, and the ozone concentration data measured by the ozone concentration measuring station can only represent the ozone concentration at that location point. For the monitoring of ozone concentration in a large area or a large region, the monitoring method of the ozone concentration measuring station often cannot be achieved. Currently, for the estimation of ozone concentration in a large area, the data of each ozone concentration measuring station are often used for spatial interpolation processing (the purpose is to convert discrete point data into a continuous data surface method), such as using the inverse distance weighted interpolation method (the inverse distance weighted interpolation method assumes that points closer in distance have a greater influence on the predicted value than points farther away), and the predicted estimation is carried out by assigning weights to known points and calculating the weighted average value. The weight is the inverse function of the distance between the known point and the target point, and the smaller the distance, the greater the weight; this method is simple and intuitive, with a fast calculation speed, and is applicable to the case of uniform data distribution; however, the inverse distance weighted interpolation method is based on the assumption that the spatial variation in all directions is consistent. In practical applications, it performs poorly in the case of sparse or complex data distribution, with low estimation accuracy, and is not applicable to the estimation of ozone concentration over a large area. For example, spatial interpolation processing also includes the Kriging interpolation method (a spatial interpolation method based on geostatistics, assuming that geographical variables have spatial autocorrelation). Kriging interpolation not only considers the values of known points but also considers spatial variability (described by the semivariogram function); when calculating, it is necessary to first calculate the semivariogram function to estimate spatial correlation, and then select a suitable model (such as a spherical model, an exponential model, a Gaussian model, etc.) for fitting according to the calculated semivariogram values, and calculate the weights through the semivariogram model with the goal of minimizing the prediction error. Relatively speaking, the Kriging interpolation method can more accurately describe spatial variability, but the Kriging interpolation method is very sensitive to the selection of the semivariogram model and has low estimation accuracy, and is also not applicable to the estimation of ozone concentration over a large area.
[0003] Atmospheric chemistry models can mainly simulate and study the physical transport and chemical reaction processes of the atmosphere. By comparing and verifying with ground, remote sensing and other observation values, the simulation of ozone (O 3) The photochemical reaction process of it and its precursors is analyzed to identify the sources of pollutants and estimate the distribution of pollutant concentrations. The core of a conventional atmospheric model mainly consists of three parts: physical transport, pollutant emissions, and diffusion and chemical transformation. At the beginning of the model, pollutant emissions are usually input as a list. Then, under the driving of the meteorological field and the calculation of the chemical reaction mechanism, the real-time concentration of atmospheric pollutants is simulated. Depending on the input parameters, the pollutant concentration over a specific time period can be integrated to obtain the average pollutant concentration during that period. In recent years, machine learning models have developed rapidly. How to deeply integrate ground station monitoring data, multi-source geographical element data, and satellite remote sensing data for learning, iterative optimization, and further use in estimating near-surface ozone concentrations is a technical challenge that needs to be studied and solved for estimating ozone concentrations over large areas. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for estimating ozone concentration data at the 100-meter scale based on the LG-Transformer model. The LG-Transformer model is constructed by deeply integrating the Transformer model and the LightGBM connection layer. The data is resampled corresponding to the grids of the grid map, and the spatial information, time information, and feature information of the data in the grids are fused and encoded to capture the relationship between spatio-temporal information, data feature information, and ozone concentration and train the model, enabling effective estimation of ozone concentrations in the study area according to the grids and further obtaining the ozone concentration distribution of the entire study area.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A method for estimating ozone concentration data at the 100-meter scale based on the LG-Transformer model, the method comprising:
[0007] S1. Construct a grid map and collect road network data, ERA5 data, and remote sensing data and resample them into the corresponding grids of the grid map;
[0008] S2. Construct the LG-Transformer model that includes a Transformer model and a LightGBM connection layer, and use the sample dataset to train the model. The sample dataset associates the corresponding sample data and ozone concentration labels according to the grid. The sample data includes road network data, ERA5 data, and remote sensing data. The Transformer model in the LG-Transformer model fuses and encodes the spatial information, temporal information, and feature information of the data in the grid. The Transformer model sets the time step to perform correlation encoding between each time step using the multi-head attention mechanism, then performs residual connection and layer normalization processing, and then uses the ReLu function for non-linear transformation. The Transformer model uses the interactive attention mechanism for decoding and outputs a sequence feature matrix.
[0009] S3. The LightGBM connection layer calculates the feature importance of the sequence feature matrix and selects the features with high importance as the splitting points of the decision tree. At the same time, it calculates the residual between the model prediction value and the true value as the model constraint. The LightGBM connection layer weights and sums the prediction results of all decision trees and generates an output value, and completes the final regression of the output value through the fully connected layer.
[0010] S4. Collect the road network data, ERA5 data, and remote sensing data of the study area, resample them to the corresponding grids of the grid map, and then input them into the trained LG-Transformer model and output the ozone estimation concentration data according to the grid.
[0011] To better implement the present invention, in step S1, the grid in the grid map has a pixel size of 100m * 100m. The ERA5 data is sourced from two datasets, Era5-land and Era5 single level. The ERA5 data includes 2m dew point temperature, 2m temperature, surface latent heat flux, surface solar radiation, surface thermal radiation, boundary layer height, and 100m wind speed. The ERA5 data is resampled to the grids of the grid map using the bilinear interpolation method. The remote sensing data includes normalized difference vegetation index data, aerosol optical depth data, night radiation data, tropospheric formaldehyde data, tropospheric nitrogen oxide data, and atmospheric ozone column concentration data. The remote sensing data is resampled to the grids of the grid map using the bilinear interpolation method.
[0012] Preferably, in step S2, the Transformer model uses the spherical embedding layer to extract the spatial information of the data in the grid. The method is as follows:
[0013] The spatial information of the data in the grid is two-dimensional longitude and latitude information, which is mapped to a three-dimensional sphere. The expression of the spherical coordinate vector (x, y, z) is as follows:
[0014] x = cos(φ) × cos(φ), y = cos(φ) × sin(λ), z = sin(φ), where φ and λ represent latitude and longitude in two-dimensional longitude and latitude information respectively;
[0015] Feature fusion for extracting spatial correlation is performed through a neural network, and the fusion expression is as follows:
[0016] fusion(x, y, z) = ReLU((x, y, z) × W 1 +b 1 ) × W 2 +b 2 , where fusion(x, y, z) represents the result after feature fusion, and W 1 , W 2 represent weight matrices respectively, and b 1 , b 2 are bias terms respectively.
[0017] Preferably, in step S2, the Transformer model extracts the time information of the data in the grid and performs forward and backward position encoding in sequence according to the set time steps. The position encoding expression is as follows:
[0018]
[0019] where P t,2i represents the encoded data at even positions, P t,2i+1 represents the encoded data at odd positions, t represents the time step index number, i represents the i-th data item in the grid data, and d model represents the feature dimension.
[0020] Preferably, the multi-head attention mechanism of the Transformer model performs operations related to correlation between each time step, which are divided into operations of query Q, key K, and value V. The expressions are as follows:
[0021] where Q = X input ·W Q , K = X input ·W K , V = X input ·W v , W Q , W k represent projection matrices respectively, K T represents the transpose of key K, d k represents the dimension of key K, and X input represents the input for the correlation operation; the multi-head attention mechanism performs and iterates the correlation operation h times independently for each head.
[0022] Preferably, the residual connection and layer normalization processing method of the Transformer model are as follows: input X for the correlation operation input Perform a residual connection with the output of the correlation operation of the multi-head attention mechanism iterated h times, and then perform layer normalization processing; the non-linear transformation expression of the ReLu function is as follows:
[0023] X ffn = ReLu(X attn ·W 1 + b 1 )·W 2 + b 2 , where X attn represents the output after layer normalization processing, and X ffn represents the output after non-linear transformation.
[0024] Preferably, in step S3, the final regression expression is as follows:
[0025] y final = F final (x)·W out + b out , where y final represents the ozone estimated concentration, F final (x) represents the output value generated by the LightGBM connection layer after weighted summation of the prediction results of all decision trees, W out represents the weight layer, and b out represents the bias layer.
[0026] Preferably, the sequence feature matrix is flattened according to all time step information before being input into the LightGBM connection layer; the LightGBM connection layer divides continuous features into discrete bins through an optimized histogram algorithm.
[0027] Preferably, in step S1, the road network data is road network vector data, and the road network data is subjected to kernel density analysis to output road network density raster data with a pixel size of 100m * 100m.
[0028] Preferably, in step S2, the ozone concentration label in the sample dataset is derived from the actual ozone concentration monitored by the measuring station.
[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0030] (1) The present invention obtains the LG-Transformer model by deeply fusing the Tansformer model and the LightGBM connection layer. The data is resampled according to the grids of the grid map, and the spatial information, time information, and feature information of the data in the grid are fused and encoded, realizing the capture and model training of the relationship between spatiotemporal information, data feature information and ozone concentration, and being able to estimate the ozone concentration at a resolution of 100 meters. The final estimation result reaches a spatial resolution of 100m×100m, which significantly improves the precision and practical application value of the estimation result. The present invention can effectively estimate the ozone concentration in the study area according to the grid, and then obtain the ozone concentration distribution in the entire study area.
[0031] (2) The present invention uses the Transformer model to obtain the time information of the data in the grid and encodes the position in sequence, which can effectively process the correlation between the previous and next sequence data and enhance the time correlation in the ozone concentration estimation; the LightGBM connection layer of the present invention improves the processing capabilities of high-dimensional discrete and continuous features and improves the estimation accuracy; the present invention can capture slight changes in ozone concentration, making the positioning of pollution sources more accurate and the spatial distribution of ozone concentration more refined; it significantly improves the spatial precision of the ozone concentration distribution results, and can provide more targeted support for urban micro-scale atmospheric environmental governance, health risk assessment and policy decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of the method of the present invention;
[0033] Figure 2 This is a fitting relationship diagram of the real value and the predicted value of a city selected as the study area in the embodiment;
[0034] Figure 3 This is a schematic diagram of the ozone concentration distribution at the hundred-meter level in a certain city selected as the study area in the embodiment;
[0035] Figure 4 This is a graph showing the ozone concentration results for a certain area of the study area in the example. DETAILED DESCRIPTION
[0036] The present invention is further described in detail below in conjunction with embodiments:
[0037] Example
[0038] like Figure 1 As shown, a method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model includes:
[0039] S1. Construct a grid map and collect road network data, ERA5 data, and remote sensing data, and resample them to the corresponding grids of the grid map.
[0040] In some embodiments, the grid in the grid map has a pixel size of 100m * 100m. The ERA5 data is sourced from two datasets, Era5-land and Era5 single level. The ERA5 data includes 2m dew point temperature (abbreviated as d2m), 2m temperature (abbreviated as t2m), surface latent heat flux (abbreviated as slhf), surface solar radiation (abbreviated as ssr), surface thermal radiation (abbreviated as str), boundary layer height (abbreviated as b1h), and 100m wind speed. In this embodiment, the 100m wind speed includes u100 and v100. u100 refers to the component of the wind speed at a height of 100 meters in the east-west direction (i.e., the meridional direction), and v100 refers to the component of the wind speed at a height of 100 meters in the north-south direction (i.e., the zonal direction). The ERA5 data is resampled to the grids of the grid map using the bilinear interpolation method. The bilinear interpolation method is used to resample the ERA5 data from the initial resolution to the grids of the grid map with a resolution of 100m * 100m. The remote sensing data includes normalized difference vegetation index data, aerosol optical depth data, night-time radiation data, tropospheric formaldehyde data, tropospheric nitrogen oxide data, and atmospheric ozone column concentration data. The remote sensing data can be sourced from multiple sources. For example, the normalized difference vegetation index data (abbreviated as NDVI) and aerosol optical depth data (abbreviated as AOD) are sourced from M0DIS (the full English name is Moderate Resolution Imaging Spectroradiometer, i.e., the remote sensing data in the Moderate Resolution Imaging Spectroradiometer), the night-time radiation data (abbreviated as NR) is sourced from VIIRS (the full English name is Visible Infrared Imaging Radiometer Suite), and the tropospheric formaldehyde data, tropospheric nitrogen oxide data, and atmospheric ozone column concentration data are sourced from TropOMI (the full English name is Tropospheric Monitoring Instrument). The remote sensing data is resampled to the grids of the grid map using the bilinear interpolation method (resampled from the initial resolution to the grids of the grid map with a resolution of 100m * 100m. In this embodiment, the remote sensing data is collected on a monthly scale). The road network data is road network vector data (in this embodiment, it can be sourced from 0pen Street Map, including all road network vector data such as highways, arterial roads, main roads, and secondary roads). The road network data is subjected to kernel density analysis to output road network density raster data with a pixel size of 100m * 100m.
[0041] In this embodiment, a certain city is selected as the research area. The research area is divided into grids of 100m * 100m in the grid map, and road network data, ERA5 data, and remote sensing data are collected and matched to the corresponding grids in the grid map according to the corresponding relationship in time and space (the grid contains road network data, ERA5 data, and remote sensing data, and also contains time information and space information, and the space information is longitude and latitude information).
[0042] S2. Construct an LG-Transformer model that includes a Tansformer model and a LightGBM connection layer, and use the sample data set for model training. The LG-Transformer model of the present invention is formed by the deep fusion of the Tansformer model and the LightGBM connection layer. The sample data set associates the corresponding sample data and ozone concentration labels according to the grid (the ozone concentration labels in the sample data set come from the actual monitored true ozone concentration at the stations), and the sample data includes road network data, ERA5 data, and remote sensing data; the Tansformer model of the LG-Transformer model performs fusion encoding on the spatial information, time information, and feature information of the data in the grid. The Tansformer model sets the time step to perform correlation encoding between each time step using the multi-head attention mechanism, then performs residual connection and layer normalization processing, and then uses the ReLu function for non-linear transformation.
[0043] In some embodiments, the Tansformer model uses a spherical embedding layer to extract the spatial information of the data in the grid. The method is as follows:
[0044] The spatial information of the data in the grid is two-dimensional longitude and latitude information, and the two-dimensional longitude and latitude information is mapped to a three-dimensional sphere (the two-dimensional longitude and latitude information is mapped to a three-dimensional sphere and non-linearly fused to learn complex spatial relationships). The expression of the spherical coordinate vector (x, y, z) is as follows:
[0045] x = cos(φ) × cos(φ), y = cos(φ) × sin(λ), z = sin(φ), where φ and λ respectively represent the latitude and longitude in the two-dimensional longitude and latitude information;
[0046] Feature fusion for extracting spatial correlation is performed through a neural network. The fusion expression is as follows:
[0047] fusion(x, y, z) = ReLU((x, y, z) × W 1 +b 1 ) × W 2 +b 2 , fusion(x, y, z) represents the result after feature fusion (the output geographical embedding vector), W 1 、W 2represent the weight matrix and b respectively 1 and b 2 are the bias terms respectively.
[0048] In some embodiments, the Transformer model extracts the time information of the data in the grid and performs forward and backward position encoding in sequence according to the set time step (in this embodiment, the month is selected as the time step), and the position encoding expression is as follows:
[0049]
[0050] where P t,2i represents the encoded data of the even position, P t,2i+1 represents the encoded data of the odd position, t represents the time step index number, i represents the ith data item in the grid data, and d model represents the feature dimension; to ensure that the encodings of different positions are not too similar and provide good time step discrimination, 10000 is selected as the base; through position encoding, the uniqueness of the position is introduced in the high-dimensional space. The input feature d input (the input feature is the feature dimension of each time step, and the feature dimension is the dimension of the data item in the grid data) is mapped to the feature dimension d model of the Transformer model.
[0051] In some embodiments, the multi-head attention mechanism of the Transformer model performs operations related to each time step, which are divided into operations of query Q, key K, and value V, and the expression is as follows:
[0052] where Q = X input ·W Q , K = X input ·W K , V = X input ·W v , W Q and W k represent the projection matrices respectively, K T represents the transpose of the key K, d k represents the dimension of the key K, and X input represents the input of the correlation operation; the multi-head attention mechanism performs independently for each head and iterates the correlation operation h times respectively. For example, if there are k heads, the data output expression of the multi-head attention mechanism is as follows:
[0053] MultiHead(X) = Concat(head 1 ,..., head k )·W 0 , where head k represents the output of the kth head's correlation operation, and W0 Output represents the projection matrix, Concat represents the CONCAT function, and MultiHead(X) represents the output of the multi-head attention mechanism after h iterations of the correlation operation.
[0054] In some embodiments, the residual connection and layer normalization processing methods of the Transformer model are as follows: the input X of the correlation operation input is subjected to a residual connection with the output MultiHHead(X) of the multi-head attention mechanism after h iterations of the correlation operation, and then layer normalization processing is performed; the non-linear transformation expression of the ReLu function is as follows:
[0055] X ffn = ReLu(X attn ·W 1 + b 1 )·W 2 + b 2 , where X attn represents the output after layer normalization processing, and X ffn represents the output after non-linear transformation. Preferably, a normalization connection is performed on the output after layer normalization processing and the output after non-linear transformation.
[0056] The Transformer model uses an interactive attention mechanism for decoding and outputs a sequence feature matrix. The interactive attention mechanism is used to capture the relationship between the input sequence and the target sequence, and the expression is as follows:
[0057] where Q u represents the input sequence, and K u , V u represent the key and value respectively (usually the hidden state of the last layer in the encoding stage), represents the dimension of the key K u , and K u T represents the transpose of the key K u .
[0058] S3. The LightGBM connection layer calculates the feature importance of the sequence feature matrix and selects the features with high importance as the splitting points of the decision tree. At the same time, the residual between the model prediction value and the true value is calculated as the model constraint; the LightGBM connection layer weights and sums the prediction results of all decision trees and generates an output value, and the output value is finally regressed through the fully connected layer. Preferably, the sequence feature matrix is flattened according to all time step information before being input into the LightGBM connection layer; the LightGBM connection layer divides the continuous features into discrete bins through an optimized histogram algorithm, which can accelerate the calculation and reduce the memory occupation.
[0059] In some embodiments, the residual expression between the predicted value and the true value of the computational model is as follows:
[0060] where y j represents the true value at the j-th position, and F t-1 (x j ) represents the predicted value of the model for the j-th sample after the (t - 1)-th iteration.
[0061] The loss function adopts the mean squared error, and the expression is as follows: where represents the total difference between the true value and the predicted value, y j , respectively represent the true value and the predicted value at the j-th position, and n represents the total number of samples j. According to the input features and the residual r i a new decision tree is constructed; first, for each feature, multiple splitting points are tried, the mean squared error (MSE) of the residuals of the left and right child nodes after splitting is calculated, and the splitting point that can maximize the error reduction is selected; then, starting from the root node, the features are gradually split until the maximum depth or the number of leaf nodes is reached. Then the model is updated by adding the output result of the new tree to the current model:
[0062] F t (x) = F t-1 (x) + η·T(x), where η is the learning rate and T(x) is the prediction result of the new tree; in each round of training, the LightGBM connection layer calculates the residual (i.e., the negative gradient) as the new target variable according to the loss function and fits it; finally, the prediction results of all decision trees are weighted and summed to generate the output value F final (x).
[0063] In some embodiments, the final regression expression is as follows:
[0064] y final = F final (x)·W out + b out , where y final represents the estimated ozone concentration, F final (x) represents the output value generated by the LightGBM connection layer after weighting and summing the prediction results of all decision trees, W out represents the weight layer, and b out represents the bias layer.
[0065] S4. Resample the road network data, ERA5 data, and remote sensing data of the study area to the corresponding grids of the grid map, then input them into the trained LG-Transformer model and output the ozone estimation concentration data corresponding to the grids. In this embodiment, a certain city is selected as the study area. The study area is divided into 100m * 100m grids in the grid map. Finally, the road network data, ERA5 data, and remote sensing data of the study area are collected and input into the trained LG-Transformer model to obtain the ozone estimation concentration of each grid in the study area of the grid map, and then according to as Figure 3 shown, a schematic diagram of the ozone estimation concentration distribution at the 100-meter level can be obtained. The model of this embodiment is verified using the five-fold cross-validation index. As Figure 2 shown, the ozone concentration estimation of the LG-Transformer model of the present invention has good centralization, and the deviation from the true value is small. The estimation result fits well with the true value, and the accuracy of the ozone concentration estimation result is high; its coefficient of determination R 2 reaches 0.936 (close to 1, indicating relatively accurate estimation), RMSE reaches 7.461 μg / m 3 , MAE is 5.7000 μg / m 3 , the Pearson correlation coefficient r is 0.9759 (close to 1, indicating good correlation), and MAPE reaches 6.53% (small deviation). Through the fusion of ozone-related multi-source feature data and the innovation of the model structure, the present invention greatly improves the accuracy of the estimation of near-surface ozone concentration data. In this embodiment, a certain city is selected as the study area. The study area is divided into 100m * 100m grids in the grid map to obtain Figure 3 the schematic diagram of the ozone estimation concentration distribution at the 100-meter level shown; as Figure 4 shown, the full coverage data of near-surface ozone concentration with a resolution of 100m * 100m realized in the present invention can provide more detailed ozone concentration information within the city, especially in small-scale areas such as blocks and road intersections; this high-resolution monitoring can capture the minute changes in ozone concentration, making the positioning of pollution sources more accurate and the spatial distribution of ozone concentration more refined. The present invention significantly improves the spatial fineness of the ozone concentration distribution result, and can provide more targeted support for urban micro-scale air environment governance, health risk assessment, and policy decision-making.
[0066] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model, characterized by: The methods include: S1, construct a grid map, collect road network data, ERA5 data and remote sensing data and resample them to the corresponding grids of the grid map; S2. Construct an LG-Transformer model including the Tansformer model and the LightGBM connection layer and use the sample data set to train the model. The sample data set is based on the sample data and ozone concentration labels corresponding to the grid association. The sample data includes road network data, ERA5 data and remote sensing data. The Tansformer model of the LG-Transformer model integrates and encodes the spatial information, temporal information and feature information of the data in the grid. The Tansformer model sets the time step and uses a multi-head attention mechanism to encode the correlation between each time step, followed by residual connection and layer normalization, and then uses the ReLu function for nonlinear transformation. The Tansformer model uses an interactive attention mechanism for decoding and outputs a sequence feature matrix. S3, LightGBM connection layer calculates the feature importance of the sequence feature matrix and selects the features with high importance as the splitting points of the decision tree, and calculates the residual between the model prediction value and the true value as the model constraint; LightGBM connection layer weights the prediction results of all decision trees and generates an output value, and completes the final regression through the fully connected layer; S4. Collect the road network data, ERA5 data and remote sensing data of the study area and resample them to the corresponding grids of the grid map. Then input the trained LG-Transformer model and output the ozone estimation concentration data according to the grid correspondence.
2. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S1, the grid in the grid map has a pixel size of 100m*100m, the ERA5 data comes from two data sets, Era5-land and Era5 s ingle 1evel, the ERA5 data includes 2m dew point temperature, 2m temperature, surface latent heat flux, surface solar radiation, surface thermal radiation, boundary layer height, 100m wind speed, and the ERA5 data is resampled to the grid of the grid map using bilinear interpolation; the remote sensing data includes normalized vegetation index data, aerosol optical thickness data, nighttime radiation data, tropospheric formaldehyde data, tropospheric nitrogen oxides data and atmospheric ozone column concentration data, and the remote sensing data is resampled to the grid of the grid map using bilinear interpolation.
3. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S2, the Tansformer model uses the spherical embedding layer to extract spatial information from the data in the grid as follows: The spatial information of the data in the grid is converted into two-dimensional longitude and latitude information, and the two-dimensional longitude and latitude information is mapped to a three-dimensional sphere. The expression of the spherical coordinate vector (x, y, z) is as follows: x=cos(φ)×cos(φ), y=cos(φ)×sin(λ), z=sin(φ), where φ, λ 分 Indicates the latitude and longitude in the two-dimensional longitude and latitude information; The feature fusion of spatial correlation is extracted through neural network, and the fusion expression is as follows: fusion(x, y, z) = ReLU((x, y, z)×W1+b1)×W2+b2, fusion(x, y, z) represents the result of feature fusion, W1 and W2 represent weight matrices, b1 and b2 are bias terms.
4. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S2, the Tansformer model extracts the time information of the data in the grid and performs forward and backward position encoding in sequence according to the set time step. The position encoding expression is as follows: Among them, Pt t,2i Indicates the coded data of even number of bits, P t,2i+1 represents the odd-numbered coded data, t represents the time step index number, i represents the data item in the grid data, d model Represents the feature dimension.
5. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: The multi-head attention mechanism of the Tansformer model is divided into query Q, key K and value V operations for the correlation operations between each time step. The expression is as follows: Where Q = X input ·W Q , K = X input ·W K , V = X input ·w v , W Q , W k They represent the projection matrix, K T represents the transpose of key K, d k represents the dimension of key K, X input Represents the input of the correlation operation; the multi-head attention mechanism performs the correlation operation independently on each head and iterates the correlation operation h times.
6. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 5, characterized in that: The residual connection and layer normalization processing method of the Tansformer model is as follows: Input the correlation operation into X input The output of the iterative correlation operation h times of the multi-head attention mechanism is residually connected, and then the layer is normalized; the nonlinear transformation expression of the ReLu function is as follows: X ffn =ReLu(X attn ·W1+b1)·W2+b2, where X attn Represents the output of the normalized layer, X ffn Represents the output after nonlinear transformation.
7. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S3, the final regression expression is as follows: y final =F final (x)·W out +b out , where y final represents the estimated ozone concentration, F final (x) indicates that the LightGBM connection layer generates an output value by weighted summing up the prediction results of all decision trees, W out represents the weight layer, b out Represents the bias layer.
8. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: The sequence feature matrix is flattened according to all time step information before entering the LightGBM connection layer; the LightGBM connection layer divides the continuous features into discrete buckets through an optimized histogram algorithm.
9. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S1, the road network data is road network vector data, and the road network data is subjected to kernel density analysis to output road network density raster data with a pixel size of 100m*100m.
10. The method for estimating ozone concentration data at the hundred-meter level based on the LG-Transformer model according to claim 1, characterized in that: In step S2, the ozone concentration label in the sample data set is derived from the real ozone concentration actually monitored by the measuring station.
Citation Information
Patent Citations
System for monitoring ingredient change of insulating oil of transformer
CN102539283A
Fine-grained costume retrieval method based on CNN-Transform double-flow network
CN115410067A