Air quality prediction method fusing multi-site time series

Through the Transformer model structure, combined with global and multi-site coding modules, the timing characteristics and correlation characteristics of the air quality vector are extracted, which solves the problem that factor correlation information is not mined and long-term dependence is not captured in air quality prediction, and achieves higher precision air quality prediction.

CN120277383APending Publication Date: 2025-07-08TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510165652.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art cannot fully tap the comprehensive correlation impact information between various factors in the air quality system, ignore the mutual influence between different sites, and cannot effectively capture long-term dependency relationships, resulting in inaccurate air quality prediction.

Method used

The Transformer model structure is adopted, combining a single-site coding module, a multi-site coding module and a timing encoding module, and the timing and correlation characteristics of the air quality vector are extracted through the global attention mechanism and the multi-head self-attention mechanism to generate prediction results.

Benefits of technology

It improves the accuracy of air quality prediction, can better capture and understand the comprehensive impact information of the air system, overcomes long-term time dependence problems, and speeds up the model training and testing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277383A_ABST
    Figure CN120277383A_ABST
Patent Text Reader

Abstract

The invention relates to an air quality prediction method fusing a multi-site time sequence, and the method comprises the following steps: obtaining detection data synchronously collected by a plurality of detection sites in a target region at different moments, and carrying out the preprocessing of the detection data, and obtaining a first air quality vector; analyzing the first air quality vector by using an air quality prediction model to obtain a prediction result; the air quality prediction model comprises a plurality of single-site coding modules for extracting the time sequence characteristics of the first air quality vectors to obtain a plurality of second air quality vectors; the plurality of multi-site coding modules are used for extracting correlation characteristics among the second air quality vectors on each set time step to obtain a plurality of third air quality vectors; the plurality of time sequence coding modules are used for converting the third air quality vector into a time sequence of each detection site to obtain a fourth air quality vector; and the decoding module is used for generating a prediction result according to the fourth air quality vector. The air quality change can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air quality prediction, and in particular to an air quality prediction method that integrates multi-site time series. Background Art

[0002] Ambient air quality is a time series change system affected by various factors, such as the interaction of industrial gas emissions, climate, geographical location, and urban population density, etc., with certain periodicity and trend, and is closely related to social lifestyles and human life and health issues. The interweaving of various factors makes air quality prediction a complex prediction task. Within the scope of the same town, there are multiple detection stations, and the air quality data spans a relatively long time range. Predicting the corresponding air quality indicators is somewhat challenging. The Transformer network structure is a powerful model for time series prediction. Due to its advantages of strong parallel computing ability, long-term dependence processing ability, flexibility, no need to manually design features, and adaptability to multi-scale input, it has achieved great success in the fields of artificial intelligence dialogue, AI image generation, and general artificial intelligence. This makes it also an effective tool for processing complex sequence data tasks such as air quality prediction. The existing methods for air quality prediction using the Transformer network have the following problems:

[0003] 1) Existing technologies often cannot fully exploit the comprehensive correlation and influence information among various factors in the local air system. For example, there are complex interaction relationships between meteorological conditions such as temperature, humidity, wind speed, etc. and air quality, and these relationships may vary by region. Therefore, it is necessary to develop methods and models that can better capture and understand the comprehensive influence information of the local air system;

[0004] 2) Traditional prediction methods such as linear regression, simple time series models, and even deep networks such as RNN often cannot effectively capture such long-term dependence relationships. Therefore, it is necessary to develop advanced models that can handle long-term dependencies to improve the accuracy and reliability of air quality prediction.

[0005] 3) Existing technologies often regard air quality prediction as a single-site problem, ignoring the mutual influence between different sites. Therefore, it is necessary to develop models that can make full use of the cross-correlation of time and space to more accurately predict the air quality in different regions and provide more refined air quality management suggestions. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an air quality prediction method that integrates multi-site time series, which can accurately predict air quality changes.

[0007] The technical solution adopted by the present invention to solve its technical problems is: to provide an air quality prediction method that fuses time series of multiple sites, including the following steps:

[0008] Obtain the detection data synchronously collected by multiple detection sites in the target area at different times;

[0009] Preprocess the detection data to obtain the first air quality vector of each detection site;

[0010] Use the air quality prediction model to analyze the first air quality vector to obtain a prediction result; the air quality prediction model includes:

[0011] Multiple single-site encoding modules, which are in one-to-one correspondence with the detection sites, and are used to extract the temporal features of the first air quality vector corresponding to the detection sites to obtain multiple second air quality vectors;

[0012] Multiple multi-site encoding modules, which are in one-to-one correspondence with multiple set time steps, and are used to extract the correlation features between the second air quality vectors at the corresponding set time steps to obtain multiple third air quality vectors;

[0013] Several temporal encoding modules, which are in one-to-one correspondence with the detection sites, and are used to convert the third air quality vector into a time series corresponding to the detection site through a time fusion attention mechanism to obtain a fourth air quality vector;

[0014] A decoding module, which is used to generate a prediction result according to the fourth air quality vector.

[0015] Further, the preprocessing of the detection data to obtain the first air quality vector of each detection site includes:

[0016] Extract features from the detection data of each detection site respectively to obtain multiple fifth air quality vectors;

[0017] Remove the time information in all the detection data to obtain global information data, and extract features from the global information data to obtain a global feature information vector;

[0018] Based on the global feature information vector, use the global attention mechanism to adjust the weights of the features in the fifth air quality vector to obtain the first air quality vector.

[0019] Further, the extraction of features from the detection data of each detection site respectively is implemented through a linear layer.

[0020] Further, the feature extraction of the global information data is implemented by a linear layer or an LSTM network.

[0021] Further, based on the global feature information vector, adjusting the weights of the features in the fifth air quality vector by using a global attention mechanism includes:

[0022] Converting the dimension of the global feature information vector to the same dimension as the fifth air quality vector through a linear transformation;

[0023] Using the softmax function to solve for the weights of the features in the fifth air quality vector after the linear transformation and multiplying by a scaling factor to obtain a global attention score vector;

[0024] Adjusting the weights of the features in the fifth air quality vector according to the global attention score vector.

[0025] Further, the dimension of the global attention score vector is equal to the dimension of the fifth air quality vector.

[0026] Further, the extraction of the temporal features of the first air quality vector corresponding to the detection site includes:

[0027] Obtaining a position encoding vector;

[0028] Converting the dimension of the first air quality vector to the same dimension as the position encoding vector;

[0029] Generating a sixth air quality vector according to the position encoding vector and the position encoding vector;

[0030] Performing feature extraction on the sixth air quality vector based on a multi-head self-attention mechanism.

[0031] Further, converting the third air quality vector into time series of each detection site through a temporal fusion attention mechanism to obtain a fourth air quality vector includes:

[0032] Calculating the temporal attention weights of each set time step through a linear quadratic transformation;

[0033] For any one of the detection sites, obtaining the fourth air quality vector as the time series of the features corresponding to the any one of the detection sites in the third air quality vector and weighted based on the temporal attention weights.

[0034] Further, the generation of the prediction result according to the fourth air quality vector is implemented through a linear layer.

[0035] Beneficial effects

[0036] Due to the above technical solution, compared with the prior art, the present invention has the following advantages and positive effects:

[0037] (1) Based on the Transformer model structure, through the combination of the single-site encoding module corresponding to each detection site, the multi-site encoding module corresponding to each specific time step, and the temporal encoding module, the different time feature data in different sites can interact with each other, jointly exploring the hidden spatio-temporal relationship patterns, enabling the model to better abstract the complex air quality changes mathematically, realizing the cross-site and cross-time cross-information flow, further fully extracting the inherent spatio-temporal information of the data, considering the overall situation of air quality changes in a region, and improving the prediction accuracy.

[0038] (2) By combining the global attention mechanism and the multi-site encoding module, the present invention first uses the global attention mechanism to initially capture the correlation of different environmental climate indicators, and then uses the multi-site encoding module to further deeply explore the correlation between the site characteristics in different regions on this basis, being able to fully explore the complex interaction relationships between various factors in the local air system, such as temperature, humidity, wind speed, etc., and air quality. Therefore, it can better capture and understand the comprehensive influence information of the local air system;

[0039] (3) Through the self-attention mechanism, the present invention overcomes the long-term temporal dependence problem existing in the prior art, avoids the information loss problem caused by too long time series, and makes the model easier to be generalized and used;

[0040] (4) Through the multi-head attention method, the present invention can process the computing tasks in parallel, accelerating the training and testing of the model. In addition, the data collection and processing processes of this model are relatively simple and supported by corresponding modules, making the deployment and application of the model more convenient to implement. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the structural diagram of the air quality prediction model of the embodiment of the present invention;

[0042] Figure 2 is the structural diagram of the single-site encoder of the embodiment of the present invention;

[0043] Figure 3 is the schematic diagram of cross-space-time information exchange of the embodiment of the present invention;

[0044] Figure 4 is the structural diagram of the temporal encoder of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0046] An embodiment of the present invention relates to a multi-encoder air quality prediction system and prediction method based on the Transformer structure. The specific solution is as follows:

[0047] First, a global feature information vector is established according to the comprehensive information of the air quality system This vector summarizes the global air quality feature data information within the time period T through the hidden layer in the LSTM, and then generates corresponding global attention scores for the features of each site through the global attention mechanism to adjust the influence of the site features, generating the input of the single-site encoder

[0048] The single-site encoder is designed based on the Encoder structure of the Transformer. Through the multi-head self-attention mechanism, the feature information of a single site can interact at different time spans. Subsequently, the output of the single-site encoder acts on the multi-site encoder. Although the multi-site encoder only transmits information of the feature vectors of different sites at the same moment, since the single-site encoder itself has mined the time correlation, the multi-site encoder can actually further explore the different time feature correlations of different sites, that is, it completes the cross-time and cross-space correlation analysis.

[0049] Then, the output of the multi-site encoder acts on the time series encoder to obtain the final embedding vector h of site u u , which not only summarizes the comprehensive air quality information but also summarizes the information of other sites and other time points, and is an encoded vector with rich correlation meanings. Finally, the embedding encoded vector h u obtained through multiple encoder pipelines is decoded, and the prediction index at the corresponding time point, such as the PM2.5 value, is calculated through linear transformation.

[0050] The overall design architecture is as shown in the appendix Figure 1 as follows, and the specific solution details are as follows:

[0051] Problem definition: Let the set of air quality detection sites in a certain area be S, and the vector Denote the air quality feature vector of site u at time t, where u ∈ S, t ∈ [1, T], and T represents the current learning window size. According to the dataset and the air quality prediction model Multi-Encoder Former, the main air quality indicators (such as PM2.5) of site u at time T+1 are predicted

[0052]

[0053] Global air quality system impact: By establishing a global air quality attention mechanism, certain weights are assigned to each feature of the site to represent the importance of the feature for the prediction result, thereby improving the prediction accuracy. The global feature information vector is used Denote the overall situation of air quality, where m is the global notation and also represents the dimension of the vector. This vector is generated by encoding other data of each site except time information through a linear layer or an LSTM network. The global air quality attention mechanism passes the feature vector through a linear transformation and then calculates the weights of each feature through the softmax function, adaptively generating a scaling factor α for the air quality features of the site t , which is used to amplify or reduce the corresponding feature information, playing the role of screening important influencing factors. In the softmax function, by increasing the temperature control factor λ, the generated attention score distribution can be made more uniform, or λ can be reduced to make the distribution more sharp, increasing the weight of prominent features. The originally generated attention scores are between [0, 1]. By multiplying by the scaling coefficient n (i.e., the feature dimension), the attention scores are expanded to between [0, n], thus truly playing the role of scaling feature information. After obtaining the global attention scores, the influence of the air system on the monitoring target of a single site is fully utilized to generate a weighted feature vector Thereby highlighting the importance of certain features.

[0054] Single-site encoder: This module focuses on studying the relationship of features of a single site in chronological order. Compared with the complex entire air quality system, the air quality change trend of a single site is easier to analyze and study. We can first simply start from the correlation of features at different time series. As shown in the appendix Figure 2 The single-site encoder is improved based on the Encoder structure of Transformer. Similar to the way Transformer maps word tokens to embeddings, the weighted site feature vector obtained above is mapped to a feature embedding through a linear transformation Through this linear mapping, the n-dimensional feature is transformed into an E-dimensional embedding vector, obtaining richer semantic information. On this basis, in order to represent the chronological order of the feature vector, a position encoding vector p t ∈ RE This vector is composed of trigonometric function values with different frequencies and phases, enabling the encoding vectors between different sequential positions to be distinguishable from each other. The combinations of trigonometric functions with different frequencies are similar to the weight value combinations of each bit in binary encoding, ensuring the numerical meaning of the same interval in different-length sequences. After obtaining the embedding vector and position encoding, the input of the multi-head attention mechanism is generated According to the self-attention mechanism, three trainable parameter transformation matrices are given Combine and the corresponding transformation matrices to obtain three matrices Q I , K I , V I . Then, Self-Attention operations are performed on these three matrices. To extract the internal relationship patterns of features from multiple perspectives, the multi-head attention method is used here, that is, h1 groups of different sub-query, key, and value matrices are set. After passing through h1 groups of Self-Attention operations respectively, the multiple head outputs obtained are concatenated and fused. This method not only searches for the relevance of features from multiple semantic levels but also makes full use of computational parallelism, which is beneficial to the training and development of the model

[0055] The output obtained through multi-head attention is just a simple concatenation of multiple attention heads. To better represent this result, it is input into a feed-forward neural network so that the results generated by each sub-attention head can be further combined. The feed-forward neural network adopts a two-layer fully connected layer structure, with the activation function being ReLU, and some optimization methods such as normalization LayerNorm and random dropout Dropout are used for optimization to ensure the model training effect. In addition, Residual connections are adopted in both the multi-attention layer and the feed-forward network layer, which can directly transfer the original feature information to the subsequent layers, alleviating the problem of gradient disappearance. Thus, it helps the network learn more discriminative feature representations and improves the performance of the network; and makes the network easier to learn the identity mapping, thereby reducing the degrees of freedom of the network and reducing the risk of overfitting. Finally, the output of the single-site encoder for site u at time t is obtained

[0056] Multi-site encoder: After processing the temporal relationship of features of a single site, the processing target turns to the analysis of the potential correlation between features of multiple different sites at a specific time step t. Similar to the single-site encoder, the self-attention mechanism is continued to automatically mine the asymmetric and dynamic correlations between multiple sites. Given the output of the single-site encoder and the position encoding p u , the input of the multi-site encoder at time t is constructed The component of site u in Calculated. Through three trainable parameter transformation matrices Combined with the input The query, key, and value matrices Q M , K M , V M of the multi-site encoder can be obtained. The calculation method is the same as that of the above single-site encoding. In the same way as above, using the multi-head attention method, h2 groups of different sub-query, key, and value matrices are set. After obtaining the output of the multi-head attention mechanism, similar operations to the single-site encoder are performed to complete the multi-head information fusion and obtain the output of the multi-site encoder

[0057] Combined effect of the single-site encoder and the multi-site encoder: The single-site encoder exchanges and transmits information on the feature factors of a single site at different time points, so that the features at time t carry information related to the previous and subsequent time correlations; on this basis, the multi-site encoder conducts information exchange among different sites at the same time t, which can not only enable information to flow among different sites at the same time, but also enable it to flow among different sites at different times because the output of the single-site encoder at a specific time has information from other time points. The combination of the single-site encoder and the multi-site encoder realizes the cross-site and cross-time cross-information flow, and further fully extracts the inherent spatio-temporal information of the data. As shown in the appendix Figure 3 The information flow principle after the combination of the single-site encoder and the multi-site encoder is shown

[0058] Temporal encoder: After passing through the multi-site encoder Certain other site and time-related information is also attached to the original feature information. In order to more concisely represent the feature information of the entire learning time window, a time fusion attention mechanism is adopted here to fuse all T sequential embedding vectors of site u into one embedding vector, as shown in the appendix Figure 4 The specific method is to first calculate the temporal attention weights at each time point through a linear quadratic transformation W T ∈R E×E so as to obtain The temporal weighted sum vector h u ∈R E . h u is the final feature summary vector of site u after being processed by multiple encoders, which contains very rich time and space information and can be directly used for predicting the target

[0059] Model prediction: After obtaining h u , it is used to predict the air quality index of site u at time T + 1. Here, a simple linear layer f(·) can be used for prediction to obtain the predicted value

[0060] The present invention will be further described in detail below in conjunction with the model architecture and implementation use cases given in the accompanying drawings.

[0061] The hardware device required for the present invention is a PC, preferably with a high-performance GPU for convenient calculation; the software resources required are a Python programming environment and a Pytorch package, and other software packages are downloaded and installed as needed. The specific steps of the multi-encoder air quality prediction method provided by the present invention are as follows:

[0062] Step 1: Obtain air quality data and perform data processing:

[0063] First, it is necessary to obtain multi-site air quality data of a certain area through the network or on-site monitoring. This data generally needs to include time information, such as year, month, day, hour, etc.; corresponding air quality indicators, such as PM2.5, SO2, NO2, CO, etc.; climate data such as temperature, humidity, rainfall, wind speed, etc. After obtaining the data, non-numerical features are removed or converted into numerical values, missing values are filled, and the prediction target feature such as PM2.5 is selected. The above time features, air pollution indicators, and climate indicators are used as the input of the global air quality attention mechanism. The air pollution indicators and climate indicators are used as the global feature information vector. Then, the data set is divided in chronological order, with the first 75% as the training set, the middle 5% as the validation set, and the last 20% as the test set. Then, it is necessary to construct a data set AirDataset and a sampler AirSamper according to the data of multiple sites. Finally, the corresponding training, validation, and test Dataloaders are constructed so that the model can perform batch iterative processing on the data, which is convenient for training and testing the data.

[0064] Step 2: Hyperparameter setting:

[0065] This step gives the key parameter settings in the model of the present invention, which is convenient for technicians to implement. The learning window T is set to 30; numbers are taken in powers of 2, and the embedding vector dimension E in the model architecture ∈ {128, 256, 512}; the learning rate lr ranges from {10e -2 , 10e -3 , 10e -4 , 10e -5 , 10e -6}; in the single-site encoder and the multi-site encoder, the number of sub-headers of the multi-head attention mechanism is set to h1, h2 ∈ {2, 4, 8}. A large number of experiments are carried out on the validation set to determine the optimal hyperparameters. Finally, according to the model performance, the best parameters are set as E = 256, lr = 10e -5, h1 = 4, h2 = 2. To avoid overfitting problems, the dropout technique is applied in the network layer, where the dropout rate is set to 0.1. The number of epochs is set to 50, which can ensure that the model training converges.

[0066] Step 3, Process the encoder input:

[0067] According to the technical solution, a GlobalAttn module can be written to process the global feature information vector This module contains a linear projection layer to transform the dimension to be the same as and then input it into the softmax function, and finally multiply by n to obtain the feature attention score vector α t α t Perform a dot product operation with to obtain the weighted feature vector It also needs to go through a linear transformation layer to obtain and expand its dimension to E. Then a PositionalEncoding module is needed to generate the position encoding vector p t , and the implementation method of this module is basically the same as that in the usual Transformer model. Then add it to p t to obtain the input of the subsequent encoder

[0068] Step 4, Process the single-site encoder:

[0069] After obtaining , use it as the input of the single-site encoder IndividualEncoder module. IndividualEncoder first performs a LayerNorm operation on to unify the data scales of each feature to the same magnitude, which is convenient for exploring the relationships between data. Then, it is necessary to obtain the input matrices Q , K I , and V I of the self-attention layer respectively through the pre-randomly initialized parameter matrices I . Immediately afterwards, Q I , K I , and V I are respectively divided into h1 groups of sub-matrices Q Ii , K Ii , and V Ii along the row dimension, and then calculate the corresponding multi-head attention sub-matrices according to the definition of SelfAttention, and concatenate these sub-matrices into a large matrix by rows. Finally, this large matrix is combined with the original input Perform residual network connection, conduct LayerNorm operation, and connect the obtained output to a feedforward neural network with a two-layer structure. Then perform another residual network connection to obtain the output of the single-site encoder.

[0070] Step 4: Multi-site encoder processing:

[0071] The processing of the multi-site encoder is basically the same as that of the single-site encoder. The main difference is that it is necessary to transpose and exchange the first and second dimensions of the Q M , K M , V M matrices, that is, the site dimension and the time dimension, for information interaction of features of different sites, and finally obtain the output.

[0072] Step 5: Temporal encoder processing and decoding prediction:

[0073] The temporal encoder calculates the temporal attention weights through a quadratic parameter matrix W T , according to the formula . After dot multiplication with , the final encoded result h u is obtained. Finally, input h u into a linear prediction function to obtain the corresponding predicted target result.

[0074] To more conveniently implement the algorithm model of the present invention, the present invention performs modular abstraction processing on the above steps and integrates them into the module MultiEncoderFormer. This module sets the corresponding parameters and connects each sub-module, that is, the data is processed in a pipeline manner according to: data preprocessing -> global feature information attention module -> single-site encoder module -> multi-site encoder module -> temporal encoder module -> linear decoding module.

[0075] The present invention also constructs a model processing module SequentialModel for model training and testing. This module can set the training platform such as CPU / GPU, set the training loss threshold stop_loss and the number of training epochs epoch, so that the training can reach an appropriate level, that is, achieve a certain accuracy while avoiding consuming too much computing resources and time. This module uses the mean squared error MSE as the loss function, Adam as the optimizer, and the stochastic batch gradient descent method SGD as the model training method. SequentialModel can save the trained model parameters to a specified file directory, and can also load the model file from a specified path to directly predict the data.

Claims

1. An air quality prediction method that fuses time series from multiple sites, characterized in that Including the following steps: Obtain the detection data synchronously collected by multiple detection stations in the target area at different times; Preprocess the detection data to obtain the first air quality vector of each detection station; Analyze the first air quality vector by using an air quality prediction model to obtain a prediction result; The air quality prediction model includes: Multiple single-station encoding modules, which are in one-to-one correspondence with the detection stations, and are used to extract the temporal features of the first air quality vector corresponding to the detection stations to obtain multiple second air quality vectors; Multiple multi-station encoding modules, which are in one-to-one correspondence with multiple set time steps, and are used to extract the correlation features between the second air quality vectors at the corresponding set time steps to obtain multiple third air quality vectors; Several temporal encoding modules, which are in one-to-one correspondence with the detection stations, and are used to transform the third air quality vector into a time series corresponding to the detection station through a time fusion attention mechanism to obtain a fourth air quality vector; A decoding module, which is used to generate a prediction result according to the fourth air quality vector.

2. The method according to claim 1, wherein The preprocessing of the detection data to obtain the first air quality vector of each detection station includes: Respectively perform feature extraction on the detection data of each detection station to obtain multiple fifth air quality vectors; remove the time information in all the detection data to obtain global information data, and perform feature extraction on the global information data to obtain a global feature information vector; Based on the global feature information vector, use a global attention mechanism to adjust the weights of the features in the fifth air quality vector to obtain the first air quality vector.

3. The method according to claim 2, wherein The feature extraction of the detection data of each detection station is realized through a linear layer.

4. The method according to claim 2, wherein The feature extraction of the global information data is realized through a linear layer or an LSTM network.

5. The method according to claim 2, wherein The adjusting the weights of the features in the fifth air quality vector by using a global attention mechanism based on the global feature information vector includes: Convert the dimension of the global feature information vector to the same dimension as the fifth air quality vector through a linear transformation; Use the softmax function to solve for the weights of the features in the fifth air quality vector after the linear transformation, and multiply by a scaling factor to obtain a global attention score vector; Adjust the weights of the features in the fifth air quality vector according to the global attention score vector.

6. The method according to claim 5, wherein The dimension of the global attention score vector is equal to that of the fifth air quality vector.

7. The method according to claim 1, characterized in that The extracting the temporal features of the first air quality vector corresponding to the detection station includes: Obtain a position encoding vector; Convert the dimension of the first air quality vector to the same dimension as the position encoding vector; Generate a sixth air quality vector according to the position encoding vector and the position encoding vector; Perform feature extraction on the sixth air quality vector based on the multi-head self-attention mechanism.

8. The method according to claim 1, wherein Converting the third air quality vector into time series of each of the detection stations through a temporal fusion attention mechanism to obtain a fourth air quality vector includes: Calculating the temporal attention weights of each of the set time steps through a linear quadratic transformation; For any one of the detection stations, obtaining the fourth air quality vector as the time series of the features corresponding to the any one of the detection stations in the third air quality vector and weighted based on the temporal attention weights.

9. The method according to claim 3, characterized in that Generating the prediction result according to the fourth air quality vector is implemented through a linear layer.