Ozone concentration estimation method based on space-time attention mechanism and multi-scale convolution

By constructing a spatiotemporal hybrid model STMO3Net, combining the spatiotemporal attention mechanism and multi-scale convolution technology, the accuracy and coverage problems of ozone concentration estimation in the existing technology are solved, and efficient and accurate near-surface ozone concentration estimation is achieved, improving the stability and calculation efficiency of the model in complex environments.

CN120277429APending Publication Date: 2025-07-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380615.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high-precision, full coverage near-surface ozone concentration estimation, especially under complex terrain and meteorological conditions, and the calculation complexity is high, so it is impossible to effectively capture the complex nonlinear relationship and space-time dynamic characteristics between ozone concentration and multiple factors.

Method used

The space-time hybrid model STMO3Net is constructed, combining the space-time attention mechanism and multi-scale convolution technology, and through spatiotemporal data matching and feature fusion, it can achieve efficient and accurate ozone concentration estimation.

Benefits of technology

It significantly improves the accuracy of ozone concentration estimation, reduces the computational complexity, generates ozone concentration distribution maps that are seamlessly covered in the entire region, enhances the robustness and stability of the model in complex environments, and provides high-precision and high-resolution near-surface ozone concentration data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277429A_ABST
    Figure CN120277429A_ABST
Patent Text Reader

Abstract

The invention relates to an ozone concentration estimation method based on a space-time attention mechanism and multi-scale convolution, and belongs to the technical field of atmospheric pollution monitoring. According to the method, a time module (including multi-head attention and a time convolutional network) and a space module (including a residual network, position coordinate attention, a multi-scale asymmetric convolutional network and channel attention) are processed in parallel through a space-time hybrid model STMO3Net, and a high-precision ozone concentration estimated value is output after space-time features are fused. By utilizing a 5km gridding spatio-temporal data matching technology, ground monitoring, satellite remote sensing and meteorological data are integrated, ozone concentration distribution estimation of global coverage of a region is realized, and the model precision R2 is greater than or equal to 0.921. The method solves the problems of complex calculation of a traditional numerical model and insufficient space-time modeling of a data-driven model, and has the advantages of high resolution, high robustness and low calculation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of air pollution monitoring, and relates to an ozone concentration estimation method based on a spatio-temporal attention mechanism and multi-scale convolution Background Art

[0002] Ozone (O3) plays an important role in blocking ultraviolet radiation in the stratosphere. However, high concentrations of O3 near the surface can cause serious harm to human health, crop growth, and building materials. Research shows that long-term exposure to high-concentration O3 environments can trigger respiratory diseases, reduce crop yields, and accelerate the aging of building materials. Therefore, accurately monitoring and estimating the spatio-temporal distribution of near-surface O3 concentration is of great significance for environmental pollution prevention and control and public health management

[0003] Currently, the monitoring of O3 concentration mainly relies on ground monitoring networks and satellite remote sensing technology. The ground monitoring network provides high-precision data through fixed stations, but its coverage is limited and the station distribution is uneven, making it difficult to reflect regional exposure differences; satellite remote sensing technology can obtain spatially continuous atmospheric O3 distribution data, but the detection accuracy of near-surface (about 2 meters height) concentration is insufficient, especially with significant errors under complex terrain and meteorological conditions. In addition, the existing technical system has the following key defects

[0004] Numerical models based on chemical reactions and atmospheric transport (such as the Weather Research and Forecasting Chemical model WRF-Chem, the Community Multiscale Air Quality model CMAQ) need to process a large number of parameters, with high computational complexity and difficulty in achieving large-scale and long-term predictions

[0005] Statistical models (such as multiple linear regression, geographically weighted regression) cannot effectively capture the complex non-linear relationship between O3 concentration and multiple factors

[0006] Traditional machine learning models (such as random forest, XGBoost) rely on point data and ignore the correlation between spatial information around stations and historical time series

[0007] Although deep learning models (such as LSTM, CNN) can process spatio-temporal data, a single network structure is difficult to fully mine multi-scale spatio-temporal features, and has limited modeling ability for long-distance dependencies and spatial heterogeneity

[0008] Existing methods have not fully integrated the spatio-temporal dynamic characteristics of ozone concentration. For example, the combined effects of factors such as short-term fluctuations in meteorological conditions, spatial distribution of pollution sources, and seasonal periodic changes have not been effectively modeled, resulting in spatial discontinuities or abnormal deviations in estimation results in areas with sparse stations

[0009] In view of the above problems, there is an urgent need for a new ozone concentration estimation method that can deeply integrate spatio-temporal features, adaptively capture multi-scale dependencies, and take into account computational efficiency. Through the construction of a spatio-temporal hybrid model (STMO3Net), this invention combines spatio-temporal attention mechanism and multi-scale convolution technology to break through the bottleneck of insufficient utilization of spatio-temporal information by traditional models, and realizes high-precision and full-coverage estimation of near-surface O3 concentration. Summary of the Invention

[0010] In view of this, the purpose of the present invention is to provide an ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution. According to the problems of limited coverage of the ground monitoring network and uneven distribution of stations, by making full use of the spatial and temporal information of the monitoring stations, accurate estimation of daily ground O3 concentration with full spatial coverage in China is realized.

[0011] By using the spatio-temporal hybrid model (STMO3Net) to replace high-cost and high-complexity numerical models (such as WRF-Chem, CMAQ), efficient and accurate ozone concentration estimation is realized.

[0012] Enhance the adaptability of the model to spatial heterogeneity and time dependence, so that it shows stronger robustness and stability in complex environments, and thus successfully captures sudden pollution events.

[0013] Through the spatio-temporal hybrid model (STMO3Net), the estimation method of combining points and surfaces and time series is realized, overcoming the limitations of traditional point-to-point estimation methods.

[0014] To achieve the above purpose, the present invention provides the following technical solutions:

[0015] An ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution, comprising the following steps:

[0016] S1: Obtain ozone concentration monitoring data, satellite remote sensing data and meteorological data, and perform data preprocessing;

[0017] S2: Construct the preprocessed data into a sample matrix (b, c, h, w) including the spatial dimension and a sample tensor (b, t, c) including the time dimension to complete spatio-temporal data matching;

[0018] S3: Input the spatio-temporal data into the spatio-temporal hybrid model STMO3Net for processing, and the model includes a parallel time module and a spatial module;

[0019] S4: Fuse the time features output by the time module and the spatial features output by the spatial module to generate an estimated value of near-surface ozone concentration;

[0020] Wherein the time module includes an embedding layer, a multi-head self-attention mechanism, and a gated neural network, and the spatial module includes a residual network, a position coordinate attention mechanism, a multi-scale asymmetric convolution network, and a context anchor attention.

[0021] Furthermore, the processing process of the time module specifically includes:

[0022] S31: Convert time series data into high-dimensional learnable vectors through linear mapping;

[0023] S32: Inject temporal position features using a position encoder;

[0024] S33: Calculate the global dependencies of the time series through a multi-head self-attention mechanism;

[0025] S34: Process the time convolutional network TCN branch and the feed-forward neural network branch in parallel in the gated neural network, and the TCN adopts a dilated causal convolution and a residual connection structure.

[0026] Furthermore, in the S34, the time convolutional network TCN includes:

[0027] At least two causal convolution layers with different dilation rates;

[0028] After each convolution, weight normalization, LeakyReLU activation, and Dropout regularization are connected;

[0029] A residual structure with cross-layer connections, and its output is concatenated with the multi-head self-attention result for feature splicing.

[0030] Furthermore, the processing process of the spatial module includes:

[0031] S35: Extract local spatial features through a residual network Residual_Block;

[0032] S36: Use the position coordinate attention mechanism PCA to perform spatial pooling in the vertical and horizontal directions respectively to generate direction-sensitive spatial attention weights;

[0033] S37: Perform feature dimensionality reduction through depthwise separable convolution DSC;

[0034] S38: Use a multi-scale asymmetric convolution MSAC module to capture spatial features with different receptive fields.

[0035] S39: Combine the context anchor attention CAA with the multi-scale asymmetric convolution MSAC to update the global channel weights.

[0036] Furthermore, the position coordinate attention mechanism is specifically implemented as:

[0037] Global average pooling is performed along the height dimension to generate a vertical feature vector;

[0038] Global average pooling is performed along the width dimension to generate a horizontal feature vector;

[0039] The vertical and horizontal feature vectors are concatenated and then passed through a convolutional layer to generate a spatial attention map;

[0040] The final spatial weighting coefficient is generated through the Sigmoid activation function.

[0041] Furthermore, the multi-scale asymmetric convolution module includes:

[0042] Four convolutional kernel structures of 1×1, 3×1, 5×1, and 7×1 in parallel;

[0043] The feature maps output from each branch are fused after being weighted by channel attention;

[0044] The channel attention adopts the context anchor attention mechanism to generate channel weights through adaptive average pooling.

[0045] Furthermore, in S2, the spatio-temporal data matching specifically includes:

[0046] The monitoring station data is extended to a 5km×5km grid spatial sample;

[0047] A time sample sequence including continuous t-day time steps is constructed;

[0048] The data with different resolutions is unified to the 5km spatial grid through bilinear interpolation.

[0049] Furthermore, the training process of the time convolutional network TCN adopts:

[0050] A hybrid loss function combining mean squared error and smooth L1 loss;

[0051] Adaptive moment estimation optimizer combined with cosine annealing learning rate adjustment;

[0052] The gated attention mechanism is adopted in the spatio-temporal feature fusion stage to dynamically adjust the feature weights.

[0053] Furthermore, the data preprocessing specifically includes:

[0054] Outlier removal and multiple imputation of missing values are performed on the monitoring data;

[0055] The meteorological data is resampled to a 5km spatial resolution;

[0056] Radiometric correction and atmospheric correction are performed on the satellite remote sensing data.

[0057] Further, in S4, the estimated near-surface ozone concentration is ozone concentration grid data covering the entire region, with a spatial resolution of 5 km × 5 km, and the coefficient of determination R 2 ≥ 0.921 is output through the model verification module as the accuracy index.

[0058] The beneficial effects of the present invention are as follows:

[0059] (1) Through the collaborative design of the spatio-temporal attention mechanism and multi-scale convolution, the accuracy of ozone concentration estimation is significantly improved (R 2 ≥ 0.921), effectively solving the problem of insufficient modeling of complex non-linear relationships and long spatio-temporal dependencies by traditional models.

[0060] The multi-head self-attention mechanism and gated TCN network in the time module capture multi-scale time features, and the position coordinate attention (PCA) and multi-scale asymmetric convolution (MSAC) in the space module extract global and local spatial information.

[0061] (2) The spatio-temporal hybrid model (STMO3Net) is used to replace the traditional numerical model (such as WRF-Chem). Through lightweight depthwise separable convolution (DSC) and parallel spatio-temporal processing, the computational complexity is significantly reduced, and efficient estimation of large-scale regions is achieved.

[0062] The model abandons the calculation of high-complexity physical equations and realizes end-to-end inference based on a data-driven framework, with the computational efficiency increased by more than 30%.

[0063] (3) Breaking through the limitation of uneven distribution of ground monitoring stations, combining the grid-based spatio-temporal matching (5 km × 5 km) of satellite remote sensing and meteorological data, generating an ozone concentration distribution map with seamless coverage of the entire region, and filling the estimation gap in areas without monitoring stations.

[0064] The bilinear interpolation and grid-based modeling techniques in spatio-temporal data matching ensure the spatial alignment of multi-source data.

[0065] (4) The time module captures short-term concentration fluctuations in sudden pollution events through dilated causal convolution (TCN) and gating mechanisms; the space module dynamically weights key regions through position coordinate attention (PCA) to improve the estimation stability in scenarios with complex terrain and source heterogeneity.

[0066] (5) The multi-scale asymmetric convolution (MSAC) combines 1×1, 3×1, 5×1, and 7×1 convolution kernels to extract spatial features with different receptive fields, reducing background noise interference; the context anchor point attention (CAA) screens key channel information in multi-scale features, enhancing the model's ability to identify the spatial distribution characteristics of pollution sources.

[0067] (6) Provide high-precision and high-resolution near-surface ozone concentration data, providing reliable data support for air pollution prevention and control, public health policy formulation, and global warming research, and promoting quantitative analysis and decision-making optimization in the field of environmental science.

[0068] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0070] Figure 1 is the overall architecture diagram of STMO3Net;

[0071] Figure 2 is the TCN network architecture diagram;

[0072] Figure 3 is the Residual_Block network architecture;

[0073] Figure 4 is the PCA network architecture;

[0074] Figure 5 is the MSAC network architecture diagram;

[0075] Figure 6 is the CAA network architecture diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0077] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0078] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0079] (1) Obtain resources: monitoring data of ozone concentration at monitoring stations, satellite remote sensing data (nitrogen dioxide and ozone column concentration in the troposphere), and meteorological data (era5 dataset) provided by the nasa website; preprocessing: outlier removal, filling of missing values, data correction, and resampling (resampling features with different resolutions into a 5km data grid).

[0080] (2) Spatiotemporally match the preprocessed feature dataset in step (1). First, combined with the image thinking, use the idea of combining points and surfaces to construct the corresponding spatial samples (b, c, h, w) (where b represents the batch size, c represents the number of features of the sample, and h and w represent the height and width of the sample respectively); second, in order to utilize the historical time information, process the time samples into (b, t, c) (where t represents the time step).

[0081] (3) To accurately estimate the near-surface ozone concentration, the present invention proposes a spatio-temporal hybrid model STMO3Net that can make full use of the spatial information and historical time information around the ozone monitoring stations. The overall architecture of the STMO3Net model is as Figure 1 shown.

[0082] STMO3Net mainly consists of two modules: a time module for effectively capturing the time dependence of ozone; a spatial module for obtaining the geospatial relationship of ozone. Finally, the output module is used to fuse the time features and spatial features to obtain the estimated near-surface O3 concentration. This time and space parallel method effectively breaks the inherent limitations of a single network in exploring geographical spatio-temporal dependencies.

[0083] (a) Time module

[0084] The time module is designed and improved based on Transformer. Among them, the Transformer model is a new type of deep neural network based on the multi-head attention mechanism, which can adaptively balance the importance of each part in the sequence. AsFigure 1 As shown, the time module is developed based on the Transformer encoder and mainly consists of an embedding layer, a multi-head self-attention mechanism, and a gated neural network. The input embedding includes an embedding layer and a positional encoding. Here, the embedding layer is used to convert the input time series data into high-dimensional learnable vectors, so as to obtain more useful information and represent it as (1):

[0085] Embedding(x time )=Linear(x time )(1)

[0086] where x time ∈R t×c represents the input time series data, t and c respectively represent the time step and the number of input feature numbers. In the embedding layer, the linear layer Linear is used to convert x time into a high-dimensional learnable variable x1 time ∈R t×d , and d represents the vector dimension of the Linear mapping. To help the model understand and process the positional relationship in the input sequence, positional encoding is used to introduce the order information of the sequence into the model, which can be expressed as formula (2):

[0087]

[0088] where pos = 1,…,t, and k = 1,…,d. After obtaining the position vector position∈R t×d , it is added to in position order for positional encoding to obtain the time vector data

[0089] In Figure 1 the time module, multi-head attention enhances the model's expressive power by introducing multiple sets of projections of query (Q), key (K), and value (V). The present invention uses s attention heads, and three identical Linear are used to map into (Q∈R t×d ), key (K∈R t×d ), value (V∈R t×d ). Each head uses different projection weights, expressed as (3):

[0090] Q i =QW Q,i ,K i =KW K,i ,V i =VW V,i (3)

[0091] where i = 1,…,s, represents the i-th attention head, W Q,i , WK,i , W V,i is the weight matrix of the i-th head. Then, for each attention head i, calculate the dot-product attention scores based on the Softmax activation, Attention(Q i , K i , V i ), expressed as Equation (4):

[0092]

[0093] where d is the dimension of K. Then, concatenate the outputs of each attention head to form the final output of the multi-head attention:

[0094] MultiHead(Q, K, V) = Concat(head1,..., head s )W o (5)

[0095] head i = Attention(Q i , K i , V i ) (6)

[0096] where W o ∈ R t×d is the weight matrix generated by the last layer of Linear for fusing the outputs of all heads. After multiple linear transformations of the multi-head self-attention and then using the residual connection, we get

[0097] The third part of the time module improves the original feed-forward neural network to a gated network ( Figure 1 ). By fusing the feature information from two different branches, not only can the key features in the input data be captured, but it also helps to alleviate the problem of vanishing or exploding gradients. These two branches play different roles. One uses Linear to update the weights of the input data and combines ReLU to achieve non-linear mapping, updating to The other one introduces the TCN network to strengthen the deeper extraction of time feature information, thus obtaining Figure 2Shows the detailed TCN network. It expands the receptive field through dilated causal convolution (Dilated Causal Conv) to capture the dependencies between different sequences. At the same time, each TCN has a convolutional layer with weight normalization to ensure strong temporal feature extraction ability. The activation function LeakyReLU, Dropout regularization, and residual connections further enhance the training stability. By using loops and setting different sizes of dilation to utilize the TCN network, temporal feature information at different scales can be better extracted, and the specific implementation is as follows:

[0098]

[0099] where l represents the number of layers of the TCN network (l ≥ 1), represents the output of the l-th layer, D is the Dropout layer to prevent overfitting, and W l is the weight updated by WeightNorm of the l-th layer, represents the causal convolution, and b l is the bias term of the l-th layer. Finally, by fusing different feature maps from the two branches, is obtained, and in order to obtain more refined temporal feature information, the time encoder part of the present invention is iterated n times.

[0100] (b) Spatial module

[0101] In the spatial module, first, an embedding layer is used, and then a residual block developed based on the ResNet is used to further extract spatial information from the input data. As Figure 3 shown, this residual block is stacked by convolutional layers with different kernel sizes, normalization layers, ReLU activation functions, and residual connections. After passing through the residual convolutional block, (c1, H, and W represent the channel dimension, height, and width of the spatial vector, respectively).

[0102] At the same time, based on different coordinate points, Positional Coordinate Attention (PCA) can effectively capture the global geospatial information through the corresponding spatial attention mechanism and dynamically update the weight sizes of each position. As Figure 4 shown, PCA performs average pooling on the horizontal and vertical directions respectively to obtain two 1D vectors. The output of the m-th channel at height h can be expressed as:

[0103]

[0104] Similarly, the output of the m-th channel with width w can be written as:

[0105]

[0106] where x all refers to Using Concatenate and CNN in the spatial dimension to fuse and compress channels, and then encoding the spatial information in the vertical and horizontal directions through BN and ReLU, it can be expressed as:

[0107] f = δW o (C(v h , v w )) (10)

[0108] where C, W o , δ respectively represent the Concatenate operation, the weights updated by the convolutional layer, and the ReLU activation function. f ∈ R c1r×H×W is the intermediate feature map encoding spatial information in the horizontal and vertical directions. Here, r is the compression ratio used to control the channel size. The main advantage of introducing the compression ratio r is to enhance the network's attention regulation ability for each channel, thereby improving the model's expressive ability and performance. Then, f is divided into two independent tensors f h ∈ R c1r×H and f w ∈ R c1r×W along the spatial dimension, and then each passes through a convolutional layer to obtain the same number of channels as the input feature Figure 1 , and then uses Sigmoid for weighting, and the corresponding representation is as follows:

[0109] g h = σW1(f h ) (11)

[0110] g w = σW2(f w ) (12)

[0111] where W1 and W2 respectively represent the weights of different convolutional layers here; σ represents Sigmoid, and g h and g w are the weights in the vertical and horizontal directions respectively. Finally, the global position weights are updated, expressed as:

[0112]

[0113] where, y m (i, j) is the pixel value at coordinates (i, j). Then, after a series of non-linear transformations and adding residual connections,

[0114] Next, the Depthwise Separable Convolution (DSC) is used to extract spatial and channel feature information, resulting in a new feature map. Moreover, a multi-scale method is introduced to extract diverse channel information. However, a large symmetric convolution kernel introduces background noise, so an asymmetric convolution kernel is adopted to extract channel information, and the feature map is converted to The structure of the Multi-Scale Asymmetric Convolution (MSAC) is as Figure 5 shown.

[0115] First, by using asymmetric convolutions of different sizes, feature maps of different scales are obtained. Then, the shapes of multiple-scale feature maps are adjusted through the upsampling method to achieve the fusion of feature maps, resulting in a feature map with feature diversity. Meanwhile, the channel weights are updated using the Context Anchor Attention (CAA) to obtain and filter out unimportant features. The structure of the CAA module is as Figure 6 shown.

[0116] After average pooling and max pooling, the sigmoid function is used to update the weights of the channel dimension for the feature map obtained by scaling the features through the convolutional layer, as shown in Equation (14):

[0117] w = σW4(δW3(C(Avg_x,Max_x))) (14)

[0118] where Avg_x and Max_x represent the feature maps obtained after average pooling and max pooling, respectively. W3 and W4 are the weights of different convolutional layers.

[0119] (c) Output module

[0120] The output module fuses the feature information from time and space. Then, three fully connected layers are used to further enhance the extraction of feature information. At the same time, the ReLU activation function is used for non-linear mapping. Finally, to prevent overfitting, Dropout with a ratio of 0.1 is set.

[0121] Experimental data: monitoring data of ozone concentration at monitoring stations in 2023, satellite remote sensing data (tropospheric nitrogen dioxide and ozone column concentrations), and meteorological data provided by the NASA website (ERA5 data set (respectively boundary layer height (BLH, unit: m), relative humidity (R, unit: %), surface air pressure (SP, unit: Pa), surface solar radiation (SSRD, unit: J / m2), 2-meter surface temperature (T2M, unit: K), total cloud cover (TCC, unit: m), total precipitation (TP, unit: m), 10-meter east wind component (U10, unit: m / s), and 10-meter north wind component (V10, unit: m / s)))

[0122] Experimental parameters: batch size batch_size is 1024, cross entropy loss function is used, Adam optimizer is used on the dataset, learning rate is 0.001, and the number of iterations is 120 rounds.

[0123] Example 1: Spatiotemporal data preprocessing and grid matching

[0124] Workflow:

[0125] 1. Data acquisition: Obtain ozone hourly concentration data from monitoring sites, and simultaneously download NO2 column concentration and O3 column concentration data from the TROPOMI satellite and ERA5 meteorological data set (surface solar radiation, temperature, wind speed).

[0126] 2. Outlier processing: The 3σ principle is used to eliminate outliers in the monitoring data, and missing values ​​are interpolated using spatiotemporal weighted interpolation of adjacent stations.

[0127] 3. Grid processing:

[0128] The regional spatial grid was divided into 5km×5km, and the monitoring station data were interpolated to the grid center using the inverse distance weighted method;

[0129] Satellite data were resampled to 5 km resolution using bilinear interpolation;

[0130] The meteorological data were matched to the same grid using Kriging interpolation.

[0131] 4. Space-time tensor construction:

[0132] Spatial sample: With the monitoring site as the center, the ozone concentration, meteorological factors, and satellite data of the grid are matched in time and space to form a matrix of (b, c, h, w) = (number of batches, 11 features, 3, 3);

[0133] Time samples: Construct a tensor of (b, t, c) = (number of batches, 4 days of time steps, 11 features) for 4 consecutive days of time windows.

[0134] Example 2: Dynamic feature extraction of time module

[0135] Workflow:

[0136] 1. Embedding layer processing: Map the input time tensor (b, 4, 11) to d = 128 dimensions through a linear layer to obtain (b, 4, 128).

[0137] 2. Position encoding: Encode the 4-day time steps according to the sine / cosine function to generate a position vector of (b, 4, 128), and add it to the embedding result.

[0138] 3. Multi-head attention calculation:

[0139] Split into s = 8 attention heads, and each head calculates the Q, K, and V matrices;

[0140] Perform scaled dot-product attention: softmax(QKT / √128)V, and the output after splicing is passed through a linear layer to obtain (b, 4, 128).

[0141] 4. Gated TCN branch:

[0142] Dilated causal convolution: The first layer has a convolution kernel k = 3 and a dilation rate d = 1; the second layer has k = 3 and d = 2, and the output is (b, 4, 64);

[0143] Residual connection: The original input is aligned in dimension through a 1×1 convolution and added to the convolution output;

[0144] Activation and regularization: LeakyReLU(α = 0.1) + Dropout(p = 0.2).

[0145] 5. Feature fusion: Weightedly splice the multi-head attention output (b, 4, 128) and the TCN output (b, 4, 64) through a gated weight (sigmoid(gate)) to obtain the final time feature (b, 4, 192).

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution, characterized in that: It includes the following steps: S1: Obtain ozone concentration monitoring data, satellite remote sensing data and meteorological data, and perform data preprocessing; S2: Construct the preprocessed data into a sample matrix (b, c, h, w) including spatial dimensions and a sample tensor (b, t, c) including time dimensions to complete spatio-temporal data matching; S3: Input the spatio-temporal data into the spatio-temporal hybrid model STMO3Net for processing, and the model includes a parallel time module and spatial module; S4: Fuse the time features output by the time module and the spatial features output by the spatial module to generate an estimated value of the near-surface ozone concentration; Wherein the time module includes an embedding layer, a multi-head self-attention mechanism and a gated neural network, and the spatial module includes a residual network, a position coordinate attention mechanism, a multi-scale asymmetric convolutional network and a context anchor attention.

2. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, wherein: The processing process of the time module specifically includes: S31: Convert the time series data into a high-dimensional learnable vector through linear mapping; S32: Inject time series position features by using a position encoder; S33: Calculate the global dependence of the time series through the multi-head self-attention mechanism; S34: Process the time convolutional network TCN branch and the feed-forward neural network branch in parallel in the gated neural network, and the TCN adopts a dilated causal convolution and residual connection structure.

3. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 2, wherein: In S34, the time convolutional network TCN includes: At least two causal convolutional layers with different dilation rates; After each convolution, weight normalization, LeakyReLU activation and Dropout regularization are connected; A residual structure with cross-layer connection, and its output is concatenated with the multi-head self-attention result for feature splicing.

4. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, characterized in that: The processing process of the spatial module includes: S35: Extract local spatial features through the residual network Residual_Block; S36: Use the position coordinate attention mechanism PCA to perform spatial pooling in the vertical and horizontal directions respectively to generate direction-sensitive spatial attention weights; S37: Perform feature dimensionality reduction through depthwise separable convolution DSC; S38: Use the multi-scale asymmetric convolution MSAC module to capture spatial features with different receptive fields; S39: Combine the context anchor attention CAA with the multi-scale asymmetric convolution MSAC to update the global channel weights.

5. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 4, characterized in that: The position coordinate attention mechanism is specifically implemented as: Perform global average pooling along the height dimension to generate a vertical direction feature vector; Perform global average pooling along the width dimension to generate a horizontal direction feature vector; Concatenate the vertical and horizontal feature vectors and generate a spatial attention map through a convolutional layer; Generate the final spatial weighting coefficient through the Sigmoid activation function.

6. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 4, characterized in that: The multi-scale asymmetric convolution module includes: Four convolution kernel structures of 1×1, 3×1, 5×1, and 7×1 in parallel; The feature maps output by each branch are fused after being weighted by channel attention; The channel attention adopts the context anchor attention mechanism to generate channel weights through adaptive average pooling.

7. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, characterized in that: In S2, the spatio-temporal data matching specifically includes: Expand the monitoring station data into a 5km×5km grid-like spatial sample; Construct a time sample sequence including continuous t-day time steps; Unify data with different resolutions to a 5-km spatial grid through bilinear interpolation.

8. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, characterized in that: The training process of the Time Convolutional Network (TCN) uses: A hybrid loss function that combines mean squared error and smooth L1 loss; An Adaptive Moment Estimation (Adam) optimizer combined with cosine annealing learning rate adjustment; In the spatio-temporal feature fusion stage, a gated attention mechanism is used to dynamically adjust feature weights.

9. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, wherein: The specific data preprocessing includes: Removing outliers and performing multiple imputations for missing values in the monitoring data; Resampling meteorological data to a 5-km spatial resolution; Performing radiometric correction and atmospheric correction on satellite remote sensing data.

10. The ozone concentration estimation method based on spatio-temporal attention mechanism and multi-scale convolution according to claim 1, characterized in that: In S4, the estimated value of near-surface ozone concentration is ozone concentration grid data covering the entire region, with a spatial resolution of 5 km × 5 km, and the coefficient of determination R 2 ≥ 0.921 is output through the model verification module as the accuracy index.