Method, device, equipment, medium and product for predicting multi-modal air pollutants

By combining multimodal data processing and deep learning network models with self-attention and cross-attention mechanisms, the problem of insufficient data fusion in air pollutant prediction is solved, thereby improving prediction accuracy.

CN120783905BActive Publication Date: 2026-04-10FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing air pollutant prediction methods suffer from shallow data fusion and low prediction accuracy.

Method used

By acquiring multimodal data of the target area, filling in missing values ​​and/or removing outliers, and performing standardization, prediction is made using a deep learning network model. Feature vectors are extracted by combining self-attention and cross-attention mechanisms, and meteorological field, emission source and pollutant concentration data are fused.

Benefits of technology

It improves the accuracy of air pollutant prediction and solves the problem of insufficient data fusion by characterizing the impact of regional pollutant concentrations through multi-source data synergy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783905B_ABST
    Figure CN120783905B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal air pollutants prediction method, device, equipment, medium and product, it is related to environmental monitoring technical field, the method includes: to be handled multi-modal data is preprocessed and standardization processing, obtains preprocessed multi-modal data;Preprocessing multi-modal data includes the current time atmospheric pollutant concentration monitoring station data in target area, emission source data and meteorological field data;Current time atmospheric pollutant concentration monitoring station data and the last time atmospheric pollutant concentration monitoring station data are spliced along time dimension, and the data features after splicing are extracted, obtain the feature vector of site position-time-pollutant concentration;The extracted feature fusion vector is fused with the feature vector of site position-time-pollutant concentration, and multi-modal feature vector is obtained;Multi-modal feature vector is input into pre-trained deep learning network model, and prediction result is obtained.The application can improve pollutant prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of environmental monitoring, and in particular to a multi-modal air pollutant prediction method, device, equipment, medium and product. BACKGROUND

[0002] With the acceleration of industrialization, air pollution problems are becoming increasingly serious, so the governance of air pollution has attracted more and more attention from countries, which makes the solution of air pollution problems imminent.

[0003] At present, when the traditional numerical mode and single-point site machine learning prediction method are used to predict air pollutants, there are problems of shallow detection data fusion degree and low prediction accuracy. Therefore, an air pollutant prediction method capable of improving prediction accuracy is urgently needed. SUMMARY

[0004] In view of the above defects or deficiencies in the related art, the purpose of the present application is to provide a multi-modal air pollutant prediction method, device, equipment, medium and product, which can effectively improve the prediction accuracy of air pollutants.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a multi-modal air pollutant prediction method, comprising: obtaining to-be-processed multi-modal data of a target area; filling in missing values of the to-be-processed multi-modal data and / or removing missing values of the to-be-processed multi-modal data, and performing standardization processing on the to-be-processed multi-modal data after filling in missing values and / or removing abnormal values to obtain pre-processed multi-modal data; the pre-processed multi-modal data includes current time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data in the target area; splicing the atmospheric pollutant concentration monitoring station data at the last time and the current time atmospheric pollutant concentration monitoring station data along the time dimension, and extracting the data features after splicing to obtain a feature vector of station position-time-pollutant concentration; extracting a feature fusion vector of the emission source data and the meteorological field data, and fusing the extracted feature fusion vector with the feature vector of the station position-time-pollutant concentration to obtain a multi-modal feature vector; inputting the multi-modal feature vector into a pre-trained deep learning network model, and using the pre-trained deep learning network model to predict the pollutant concentration of the target area to obtain a prediction result.

[0007] Optionally, the splicing the atmospheric pollutant concentration monitoring station data at the previous time and the atmospheric pollutant concentration monitoring station data at the current time along the time dimension, and extracting a feature of the spliced data, to obtain a feature vector of a site position-time-pollutant concentration, comprises: splicing the atmospheric pollutant concentration monitoring station data at the previous time and the atmospheric pollutant concentration monitoring station data at the current time along the time dimension, and embedding a relative position code and a time code of longitude and latitude between each monitoring station into the spliced atmospheric pollutant concentration monitoring station data at the previous time and the atmospheric pollutant concentration monitoring station data at the current time, to obtain spatio-temporal joint data of different monitoring stations in a target region; extracting a data feature of the spatio-temporal joint data of the different monitoring stations based on a self-attention mechanism and a feedforward network, to obtain the feature vector of the site position-time-pollutant concentration.

[0008] Optionally, the method further comprises: extracting a feature fusion vector of the emission source data and the meteorological field data, and fusing the extracted feature fusion vector with the feature vector of the site position-time-pollutant concentration, to obtain a multi-modal feature vector, comprising: extracting a feature of the emission source data and the meteorological field data based on a preset convolutional neural network, to obtain a meteorological-emission coupling multi-scale spatio-temporal feature vector; fusing the meteorological-emission coupling multi-scale spatio-temporal feature vector with the feature vector of the site position-time-pollutant concentration based on a cross-attention mechanism, to obtain the multi-modal feature vector.

[0009] Optionally, the method further comprises: performing a reverse standardization processing on the prediction result, to obtain a predicted value of an actual pollutant concentration; or, calculating an uncertainty interval of the prediction result, and evaluating an accuracy of the prediction result based on the uncertainty interval, to obtain an evaluation result.

[0010] Optionally, the training method of the pre-trained deep learning network model comprises: obtaining historical atmospheric pollutant concentration monitoring station data, historical emission source data and historical meteorological field data in a target area; filling in missing values of the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data and / or removing abnormal values of the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data, and performing standardization processing on the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data with the filled-in missing values and / or the removed abnormal values to obtain historical standard atmospheric pollutant concentration monitoring station data, historical standard emission source data and historical standard meteorological field data; embedding longitude and latitude phase position codes and time codes between each monitoring station in the target area into the historical standard atmospheric pollutant concentration monitoring station data, and extracting data features of the embedded historical standard atmospheric pollutant concentration monitoring station data to obtain a feature vector of station position-time-historical pollutant concentration; extracting a feature fusion vector of the historical emission source data and the historical meteorological field data, and fusing the extracted feature fusion vector with the feature vector of station position-time-historical pollutant concentration to obtain a historical multi-modal feature vector; inputting the historical multi-modal feature vector into a deep learning network model, performing single-step training on the deep learning network model with a first preset step size to obtain a first-stage deep learning network model; inputting the historical multi-modal feature vector into the first-stage deep learning network model, and performing joint optimization with autoregressive prediction of a second preset step size to obtain the pre-trained deep learning network model.

[0011] Optionally, the filling and / or removing of the missing values of the to-be-processed multi-modal data, and the standardization processing of the to-be-processed multi-modal data after filling the missing values and / or removing the abnormal values, to obtain the pre-processed multi-modal data, comprise: obtaining the mean and standard deviation of historical atmospheric pollutant concentration monitoring station data of a target area, the mean and standard deviation of historical emission source data, and the mean and standard deviation of historical meteorological field data; filling the missing values of corresponding data of the to-be-processed multi-modal data based on the mean of the historical atmospheric pollutant concentration monitoring station data, the mean of the historical emission source data, and the mean of the historical meteorological field data; and / or removing the abnormal values of corresponding data of the to-be-processed multi-modal data based on the standard deviation of the historical atmospheric pollutant concentration monitoring station data, the standard deviation of the historical emission source data, and the standard deviation of the historical meteorological field data; and performing standardization processing on corresponding data of the to-be-processed multi-modal data after filling the missing values and / or removing the abnormal values based on the mean and standard deviation of the historical atmospheric pollutant concentration monitoring station data, the mean and standard deviation of the historical emission source data, and the mean and standard deviation of the historical meteorological field data, to obtain the current-time atmospheric pollutant concentration monitoring station data, the emission source data, and the meteorological field data.

[0012] In a second aspect, the present application provides a multi-modal air pollutant prediction device, which comprises:

[0013] An acquisition module is configured to acquire to-be-processed multi-modal data of a target area.

[0014] A preprocessing module is configured to fill or remove missing values and abnormal values of the to-be-processed multi-modal data, and perform standardization processing on the to-be-processed multi-modal data after filling the missing values or removing the abnormal values, to obtain pre-processed multi-modal data; the pre-processed multi-modal data comprises current-time atmospheric pollutant concentration monitoring station data, emission source data, and meteorological field data.

[0015] A first extraction module is configured to splice the atmospheric pollutant concentration monitoring station data of a previous time with the current-time atmospheric pollutant concentration monitoring station data along a time dimension, and extract features of the spliced data, to obtain a feature vector of a station position-time-pollutant concentration.

[0016] A second extraction module is configured to extract a feature fusion vector of the emission source data and the meteorological field data, and further fuse the extracted feature fusion vector with the feature vector of the station position-time-pollutant concentration, to obtain a multi-modal feature vector.

[0017] A prediction module is configured to input the multi-modal feature vector into a pre-trained deep learning network model, and use the pre-trained deep learning network model to predict the pollutant concentration of the target region to obtain a prediction result.

[0018] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multi-modal air pollutant prediction method according to any one of the above embodiments.

[0019] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the multi-modal air pollutant prediction method according to any one of the above embodiments.

[0020] In a fifth aspect, the present application provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the steps of the multi-modal air pollutant prediction method according to any one of the above embodiments.

[0021] According to the embodiments of the present application, the following technical effects are achieved:

[0022] The application provides a multi-modal air pollutant prediction method, device, equipment, medium and product. The missing values of to-be-processed multi-modal data of a target area are filled in or the abnormal values of the to-be-processed multi-modal data are removed, and the to-be-processed multi-modal data after filling in the missing values or removing the abnormal values is subjected to standardization processing to obtain preprocessed multi-modal data, the preprocessed multi-modal data including current time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data; the atmospheric pollutant concentration monitoring station data at a previous time and the atmospheric pollutant concentration monitoring station data at the current time are spliced along a time dimension, and data features after splicing are extracted to obtain a station position-time-pollutant concentration feature vector; a feature fusion vector of the emission source data and the meteorological field data is extracted, and the extracted feature fusion vector is fused with the station position-time-pollutant concentration feature vector to obtain a multi-modal feature vector; the multi-modal feature vector is input into a pre-trained deep learning network model, and the pre-trained deep learning network model is used to predict the pollutant concentration of the target area to obtain a prediction result. That is, the multi-modal data is constructed by using the current time atmospheric pollutant concentration monitoring station data, the emission source data and the meteorological field data in the target area. The meteorological data, the pollutant monitoring data and the emission source data can be integrated by constructing the multi-modal data, and the purpose of multi-source data cooperation is achieved. Specifically, on the basis of considering the interaction between the meteorological field data and the pollutant concentration field, the pollution information at a local scale and a regional scale is fused, the influence of regional transportation on the pollutant concentration in the target area can be effectively represented, and the problem of low data fusion degree and low prediction accuracy in the prior art is solved, and the prediction accuracy of the pollutant concentration can be improved to a certain extent. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0024] Figure 1 An application environment schematic diagram of a multi-modal air pollutant prediction method provided by an embodiment of the present application;

[0025] Figure 2 A flowchart schematic diagram of a multi-modal air pollutant prediction method provided by an embodiment of the present application;

[0026] Figure 3 A generation flowchart schematic diagram of a multi-modal feature vector provided by an embodiment of the present application;

[0027] Figure 4 A functional module schematic diagram of a multi-modal air pollutant prediction device provided by an embodiment of the present application is provided.

[0028] Figure 5 A structural schematic diagram of a computer device provided by an embodiment of the present application is provided. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the technical solutions claimed by the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0031] The multi-modal air pollutant prediction method provided by the embodiments of the present application can be applied to, for example Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be set up separately, or integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the to-be-processed multi-modal data to the server 104, and after receiving the to-be-processed multi-modal data, the server 104 fills in the missing values of the to-be-processed multi-modal data and / or removes the missing values of the to-be-processed multi-modal data, and standardizes the to-be-processed multi-modal data after filling in the missing values and / or removing the abnormal values, to obtain pre-processed multi-modal data; the pre-processed multi-modal data includes current time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data in the target area; the atmospheric pollutant concentration monitoring station data at the last time is spliced with the atmospheric pollutant concentration monitoring station data at the current time along the time dimension, and the data features after splicing are extracted to obtain a feature vector of site position-time-pollutant concentration; the feature fusion vector of the emission source data and the meteorological field data is extracted, and the extracted feature fusion vector is fused with the feature vector of the site position-time-pollutant concentration to obtain a multi-modal feature vector; the multi-modal feature vector is input into a pre-trained deep learning network model, and the pre-trained deep learning network model is used to predict the pollutant concentration of the target area to obtain a prediction result. The server 104 can feed back the obtained prediction result to the terminal 102. In addition, in some embodiments, the multi-modal air pollutant prediction method can also be implemented by the server 104 or the terminal 102 alone, such as the terminal 102 directly processing the to-be-processed multi-modal data to obtain the prediction result, or the server 104 obtaining the to-be-processed multi-modal data from the data storage system and processing the to-be-processed multi-modal data to obtain the prediction result.

[0032] Among them, the terminal 102 can be but not limited to various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart televisions, smart air conditioners, smart vehicle devices, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0033] In an exemplary embodiment, as Figure 2 shown, a multi-modal air pollutant prediction method is provided, which is executed by a computer device, specifically can be executed by a terminal or a server, etc. Computer device alone, or can be executed by a terminal and a server together, in the embodiment of the application, the method is applied to Figure 1The server 104 in the system 100 is taken as an example for illustration, including the following steps S201 to S205. Among them:

[0034] In step S201, the target area's to-be-processed multi-modal data is obtained.

[0035] In an example embodiment, the to-be-processed multi-modal data includes atmospheric pollutant concentration monitoring station data, emission source data, and meteorological field data; the atmospheric pollutant concentration monitoring station data includes concentration data of pollutants such as sulfur dioxide (SO2), nitrogen dioxide (NO2), inhalable particulate matter (PM10), fine particulate matter (PM2.5), carbon monoxide (CO), and ozone (O3) at different locations and different times in the target area. The emission source data includes the China Multi-resolution Emission Inventory (MEIC), the China High-resolution and High-quality Near-surface Air Pollutant Dataset (CHAP), and the atmospheric pollution monitoring and emission inventory provided by the Copernicus Atmosphere Monitoring Service (CAMS); among them, the China Multi-resolution Emission Inventory (MEIC) contains emission data of multiple pollutants at different spatial resolutions, covering pollution source emission information in multiple fields such as industry, transportation, and agriculture, providing detailed information for studying the source and emission intensity of pollutants in the study area. The China High-resolution and High-quality Near-surface Air Pollutant Dataset (CHAP) provides high-resolution and reliable near-surface air pollutant data, which is of great value for understanding the near-surface layer pollutant emission situation and helps to accurately analyze the impact of emission sources on the surrounding air quality. The atmospheric pollution monitoring and emission inventory provided by the Copernicus Atmosphere Monitoring Service (CAMS) integrates atmospheric pollution monitoring data and emission inventory information worldwide, supplementing the emission source data from a global perspective, making the emission source data of the target area more comprehensive, and being conducive to comprehensive evaluation of the connection between the emission of the target area and the global atmospheric environment. The meteorological field data is mainly derived from the data published by the European Centre for Medium-Range Weather Forecasts, including the Global Atmospheric Reanalysis Data (ERA5) and the Forecast Meteorological Dataset. The ERA5 data is a reanalysis result of historical meteorological data, containing temperature, humidity, air pressure, wind speed, wind direction, and other meteorological elements, providing long-term and stable data support for studying the impact of past meteorological conditions on air pollution. The forecast meteorological dataset provides future weather forecast information, such as future weather conditions and trends of meteorological elements, which is an important basis for predicting future air pollution conditions. Combined with the atmospheric pollutant concentration monitoring station data and the emission source data, the influence of meteorological conditions on the diffusion, transmission, and transformation of pollutants can be comprehensively analyzed.

[0036] In step S202, missing values of the to-be-processed multi-modal data are filled in and / or removed, and the to-be-processed multi-modal data after filling in and / or removing the missing values and outliers is standardized to obtain pre-processed multi-modal data.

[0037] In the example embodiment, the preprocessing of the multi-modal data includes current time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data; in combination with the above embodiment, the current time atmospheric pollutant concentration monitoring station data is obtained after filling in the missing values and / or removing the outliers of the atmospheric pollutant concentration monitoring station data; and / or the emission source data is obtained after filling in the missing values or removing the outliers of the emission source data; and / or the meteorological field data is obtained after filling in the missing values and / or removing the outliers of the meteorological field data.

[0038] Optionally, the step S202 can include: obtaining the mean and standard deviation of the historical atmospheric pollutant concentration monitoring station data of the target region, the mean and standard deviation of the historical emission source data, and the mean and standard deviation of the historical meteorological field data; filling in the missing values of the corresponding data of the to-be-processed multi-modal data based on the mean of the historical atmospheric pollutant concentration monitoring station data, the mean of the historical emission source data, and the mean of the historical meteorological field data; removing the outliers of the corresponding data of the to-be-processed multi-modal data based on the standard deviation of the historical atmospheric pollutant concentration monitoring station data, the standard deviation of the historical emission source data, and the standard deviation of the historical meteorological field data; and performing standardization processing on the corresponding data of the to-be-processed multi-modal data after filling in the missing values or removing the outliers based on the mean and standard deviation of the historical atmospheric pollutant concentration monitoring station data, the mean and standard deviation of the historical emission source data, and the mean and standard deviation of the historical meteorological field data, to obtain the current time atmospheric pollutant concentration monitoring station data, the emission source data and the meteorological field data.

[0039] By filling in the missing values or removing the outliers based on the mean and standard deviation of the historical emission source data, and by standardizing the data of each grid point based on the mean and standard deviation of the historical emission source data, the data of different dimensions are unified to the same scale to generate standardized preprocessing multi-modal data, which helps to improve the prediction accuracy of the pre-trained deep learning network model.

[0040] In step S203, the atmospheric pollutant concentration monitoring station data at the last time and the atmospheric pollutant concentration monitoring station data at the current time are spliced along the time dimension, and the data features of the spliced data are extracted to obtain a feature vector of the station position-time-pollutant concentration.

[0041] It can be understood that the atmospheric pollutant concentration monitoring station data at the last time and the atmospheric pollutant concentration monitoring station data at the current time are spliced along the time dimension, and the latitude and longitude relative positions between the monitoring stations and the time coding are embedded into the spliced atmospheric pollutant concentration monitoring station data at the last time and the atmospheric pollutant concentration monitoring station data at the current time to obtain joint data of different monitoring stations; the data features of the joint data of different monitoring stations are extracted based on the self-attention mechanism and the feedforward network to obtain a feature vector of the station position-time-pollutant concentration.

[0042] Further understood, as shown in Figure 3 , first, the current time atmospheric pollutant concentration monitoring station data of multiple monitoring stations (Station1, Station2, etc.) are output to the station data processing module, including PM 2.5 , PM 10 , NO2, CO, SO2, O3, etc. Second, the station data processing module splices the current time atmospheric pollutant concentration monitoring station data X t and the previous time atmospheric pollutant concentration monitoring station data X t-6 along the time dimension through "Concat", and then processes through the full connection layer (FC) and the embedding layer (Embedding) to obtain the query vector (q), the key vector (k), and the value vector (v). Through the self-attention (SelfAttention) mechanism, the data features of the joint data of different monitoring stations after splicing are captured, and the spatial representative hidden features, i.e., the feature vectors of the station position-time-pollutant concentration, are output.

[0043] By fusing the relative position encoding of the longitude and latitude of the monitoring stations, the position information of different monitoring stations in the latitude direction can be represented by the sine and cosine functions, which facilitates the conversion of the current time atmospheric pollutant concentration monitoring station data into numerical form, and further facilitates the capture of the spatial features of the pollutants. By introducing the temporal encoding (TemporalEncoding), the Day of Year (DOY) and the Hour of Day (HOD) can be embedded (Embedding), which facilitates the capture of the daily variation and seasonal cycle characteristics of the pollutants.

[0044] It should be noted that the position encoding of different monitoring stations is performed by using the following formula (1); the time embedding is performed by using the following formula (2); the splicing is performed by using the following formula (3); and the self-attention mechanism capture is performed by using the following formula (4).

[0045]

[0046] wherein, is the relative position encoding of the longitude and latitude; lat i,j represents the latitude value of the i-th monitoring station at the j-th time step; lat 0,0 is the starting reference latitude, lat range is the latitude range.

[0047] T emb = Embedding(t doy , t hod ) (2)

[0048] wherein, Temb , t doy represents the day of year (Day of Year), t hod represents the hour of day (Hourof Day).

[0049] X inp = concat(X t-6 , PE lat,lon , T emb ) (3)

[0050] where X inp is the concatenated data features; X inp ∈R (2D)×Ns , D is the dimension of pollutants at each monitoring site, and Ns is the number of monitoring sites. X t-6 is the relevant data (such as pollutant concentration, meteorological data, etc.) 6 hours away from the current time, PE lat ,lon is the longitude and latitude position code of different monitoring sites, and T emb is the time embedding vector. Through the concatenation (concat) operation, the time, spatial position and historical data features are integrated.

[0051] X sa = Self Attention(X inp ) (4)

[0052] In the process of self-attention mechanism capture, the site data processing module generates query vector (q), key vector (k) and value vector (v) using the following formula (5); calculates the attention weight matrix using the following formula (6); calculates the weighted value vector using the following formula (7); and calculates the spatially representative hidden feature H′ sa using the following formula (8).

[0053] Q, K, V = W q ·X inp , W k ·X inp , W v ·X inp (5)

[0054]

[0055] H sa = AV (7)

[0056] H′ sa = MLP(LayerNorm(V+H sa )) (8)

[0057] where A ∈ R (2D)×NsThe self-attention matrix is a mutual importance evaluation matrix between different monitoring sites. The self-attention matrix is multiplied by V to obtain the feature hidden variable of the monitoring sites after interaction. Then, a feedforward network (FeedForward Layer) is used to perform nonlinear transformation and feature extraction on the residual structure, so as to improve the expression ability of the model and prevent overfitting of the model.

[0058] The local-regional pollution transmission correlation matrix is constructed by the self-attention mechanism, and the pollution contribution degrees at different spatial scales (such as the proportion of the transportation of PM2.5 in the target city by the upstream 300-kilometer industrial emissions) are quantified. The whole chain deduction from the emission source analysis to the receptor concentration prediction is realized, and the problem of insufficient regional transportation representation of the traditional model is effectively solved. The mechanism reduces the prediction error by modeling through multidimensional data cooperation and interaction, and provides accurate data support for pollution tracing and regional joint prevention and control.

[0059] In step S204, the feature fusion vector of the emission source data and the meteorological field data is extracted, and the extracted feature fusion vector is fused with the feature vector of the site location-time-pollutant concentration to obtain a multi-modal feature vector.

[0060] Understandably, the meteorological-emission coupled multi-scale spatio-temporal feature vector is obtained by extracting features of the emission source data and the meteorological field data based on the preset convolutional neural network; and the multi-modal feature vector is obtained by fusing the meteorological-emission coupled multi-scale spatio-temporal feature vector with the feature vector of the site location-time-pollutant concentration based on the cross-attention mechanism.

[0061] In combination with the above embodiments and with reference to Figure 3 Further understand that the grid data processing module obtains the emission source data and the meteorological field data, such as the emission source data (Emissions) and the ERA5 meteorological data of the European Centre for Medium-Range Weather Forecasts (ECMWF), and performs a downsampling operation on the input grid data to reduce the data dimension or resolution of the grid data, thereby generating a query vector (q), a key vector (k), and a value vector (v) to obtain the meteorological-emission coupled multi-scale spatio-temporal feature vector. In the embodiments of the present application, the grid data processing module uses a convolutional neural network (CNN) with a ResNet (Residual Network) architecture to perform multi-scale spatio-temporal feature extraction on the grid-based meteorological field data and the emission source data. Then, the feature fusion module uses the cross-attention (Cross Attention) mechanism to fuse the feature vectors (q, k, v) output by the site data processing module and the grid data processing module, so that the two types of data interact and capture the correlation between the preprocessed multi-modal data. Finally, the hidden feature H′ saAs Query, introduce residual connection in cross attention mechanism (Residual Connection), add hidden features H' sa Further processing by self-attention mechanism, that is, adding original Query and attention extraction features, outputting multi-modal feature vector, thus preserving original information while enhancing model representation ability.

[0062] It should be noted that in the cross attention (Cross Attention) mechanism, the following formula (9) is used for multi-source data feature extraction and integration; the following formula (10) is used for interactive attention vector generation; the following formula (11) is used for interactive attention weight calculation; the following formula (12) is used for weighted value vector calculation and feature output; and the following formula (13) is used to output the final multi-modal feature vector after fusion processing.

[0063] H EM =ResNet(concat(Met inp , EMS inp , PE lon,lat )) (9)

[0064] Q, K, V = W q ·H′ sa , W k ·H EM , W v ·H EM (10)

[0065]

[0066] H MEa =A cross V (12)

[0067] H′ MEa =MLP(LayerNorm(Q+H MEa )) (13)

[0068] Wherein, A cross ∈R Ns×Ngrids is the cross-attention matrix (Cross-Attention Matrix), H EM is the hidden variable of meteorological field data and emission source data, N grids is the grid number of H EM , A cross characterizes the interaction between the hidden variable of the current time atmospheric pollutant concentration monitoring station data and H EM , A crossThe coupling feature hidden variable is obtained by multiplying V, and then a feedforward network (FeedForward Layer) is used to perform nonlinear transformation and feature extraction on the residual structure. Since the current atmospheric pollutant concentration monitoring station data mainly provides background information, Q and H are added to complete the residual connection when the residual connection (Residual Connection) is performed, so as to effectively fuse the features obtained by the cross-attention mechanism. MEa The residual connection is completed by adding, so as to effectively fuse the features obtained by the cross-attention mechanism.

[0069] By fusing high-temporal and high-spatial resolution meteorological field data (such as ERA5 global reanalysis data), multi-pollutant monitoring data (SO2, PM2.5, and other site measured values), and multi-scale emission source lists (MEIC / CHAP / CAMS), a multi-source data set containing physical fields, chemical concentration fields, and anthropogenic emission fields is constructed, solving the problem of single data in the prior art.

[0070] The cross-attention mechanism is used to represent the interaction between the meteorological field and the pollutant concentration field, which not only captures the driving effect of wind speed / temperature on pollutant diffusion and transformation (such as PM2.5 accumulation under static stable weather), but also retains the feedback effect of pollutants on meteorological parameters (such as aerosol radiation effect) through residual connection, breaking through the limitations of traditional one-way modeling and effectively improving the prediction accuracy.

[0071] In step S205, the multi-modal feature vector is input into the pre-trained deep learning network model, and the pre-trained deep learning network model is used to predict the pollutant concentration of the target area to obtain a prediction result.

[0072] In the embodiment of the application, the pre-trained deep learning network model is used for 72-hour rolling prediction, and a prediction result with a 6-hour interval is generated; the prediction result includes a prediction value, uncertainty information, time information, and / or trend information. The prediction value is a specific numerical prediction of the pollutant concentration; the uncertainty information includes an uncertainty interval and related reliability evaluation indicators; the time information is within the next 72 hours; the trend information is, for example, an upward trend, a downward trend, or a stable trend of the pollutant concentration within the next 72 hours.

[0073] It should be noted that the pre-trained deep learning network model can be a convolutional neural network (CNN) model. The training method of the pre-trained deep learning network model includes the following steps S1 to S6, wherein:

[0074] In step S1, historical atmospheric pollutant concentration monitoring station data, historical emission source data, and historical meteorological field data in a target area are obtained.

[0075] Step S2, filling in missing values of historical atmospheric pollutant concentration monitoring station data, historical emission source data and historical meteorological field data, and / or removing abnormal values of historical atmospheric pollutant concentration monitoring station data, historical emission source data and historical meteorological field data, and standardizing the historical atmospheric pollutant concentration monitoring station data, historical emission source data and historical meteorological field data after filling in the missing values or removing the abnormal values to obtain historical standard atmospheric pollutant concentration monitoring station data, historical standard emission source data and historical standard meteorological field data;

[0076] Step S3, embedding the longitude and latitude phase position coding and time coding between each monitoring station into the historical standard atmospheric pollutant concentration monitoring station data, and extracting the data features of the embedded historical standard atmospheric pollutant concentration monitoring station data to obtain the feature vector of station position-time-historical pollutant concentration;

[0077] Step S4, extracting the feature fusion vector of historical emission source data and historical meteorological field data, and further fusing the extracted feature fusion vector with the feature vector of station position-time-historical pollutant concentration to obtain a historical multi-modal feature vector;

[0078] Step S5, inputting the historical multi-modal feature vector into the deep learning network model, and using a first preset step to perform single-step training on the deep learning network model to obtain a first-stage deep learning network model;

[0079] Step S6, inputting the historical multi-modal feature vector into the first-stage deep learning network model, and performing joint optimization using a second preset step of autoregressive prediction to obtain a pre-trained deep learning network model.

[0080] In the training process of the deep learning network model, by defining a suitable loss function (such as a mean square error loss function to measure the difference between the predicted pollutant concentration and the actual concentration), an optimization algorithm (such as stochastic gradient descent) is used to continuously optimize the parameters of the deep learning network model, so that the loss function value is minimized, and the deep learning network model learns the mapping relationship between the preprocessed multi-modal data and the pollutant concentration, so as to improve the prediction accuracy.

[0081] It should be noted that there are often a large number of missing values and abnormal data in the measured data of pollutants, and directly using strict quality control and missing value filling strategies may introduce additional errors, thereby affecting the actual prediction accuracy of the first-stage deep learning network model. Therefore, in order to reduce the dependence of the first-stage deep learning network model on the input data, the input of the first-stage deep learning network model only contains X t-6 and X t two times, and based on the autoregressive strategy, a 6-hour resolution future 72-hour pollutant concentration prediction (a total of 12 time steps) is generated.

[0082] By implementing the steps S201 to S205, the missing values of the to-be-processed multi-modal data of the target region are filled in and / or the abnormal values of the to-be-processed multi-modal data are removed, and the to-be-processed multi-modal data after filling in the missing values and / or removing the abnormal values are standardized to obtain pre-processed multi-modal data, the pre-processed multi-modal data including current time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data; the atmospheric pollutant concentration monitoring station data at the last time and the current time are spliced along the time dimension, and the data features after splicing are extracted to obtain a station location-time-pollutant concentration feature vector; a feature fusion vector of the emission source data and the meteorological field data is extracted, and the extracted feature fusion vector is fused with the station location-time-pollutant concentration feature vector to obtain a multi-modal feature vector; the multi-modal feature vector is input into a pre-trained deep learning network model, and the pre-trained deep learning network model is used to predict the pollutant concentration of the target region to obtain a prediction result; that is, the current time atmospheric pollutant concentration monitoring station data, the emission source data and the meteorological field data in the target region are used to construct multi-modal data, and the multi-modal data can integrate meteorological data, pollutant monitoring data and emission source data to achieve the purpose of multi-source data collaboration; specifically, on the basis of considering the interaction between the meteorological field data and the pollutant concentration field, the pollution information at the local scale and the regional scale is fused, which can effectively represent the influence of regional transportation on the pollutant concentration in the target region, thereby solving the problems of insufficient data fusion and low prediction accuracy in the prior art, and improving the prediction accuracy of the pollutant concentration to a certain extent.

[0083] In addition, there are a large number of missing values and abnormal data in the actual pollutant data, and directly using strict quality control and missing value filling strategies may introduce additional errors, thereby affecting the actual prediction accuracy of the model. Therefore, in order to reduce the dependence of the model on the input data, the input of the prediction model only contains X t-6 and X t two times, and based on the autoregressive strategy, a 6-hour resolution future 72-hour pollutant concentration forecast (a total of 12 time steps) is generated.

[0084] In other embodiments, in order to facilitate viewing of the prediction result, after the above step S205, the method further includes: performing inverse standardization processing on the prediction result to obtain a predicted value of the actual pollutant concentration.

[0085] It can be understood that the specific method of anti-standardization corresponds to the method of standardization. In the above embodiment, the mean and standard deviation of historical atmospheric pollutant concentration monitoring site data, the mean and standard deviation of historical emission source data, and the mean and standard deviation of historical meteorological field data are used to standardize the multi-modal data to be processed. In this embodiment, the mean and standard deviation of historical atmospheric pollutant concentration monitoring site data, the mean and standard deviation of historical emission source data, and the mean and standard deviation of historical meteorological field data are used to anti-standardize the prediction results to obtain the actual applicable prediction results.

[0086] In other embodiments, in order to evaluate the accuracy of the prediction results, after step S205, the method further comprises: calculating an uncertainty interval of the prediction results, and evaluating the accuracy of the prediction results based on the uncertainty interval to obtain an evaluation result.

[0087] It can be understood that when predicting the pollutant concentration, due to various factors such as data incompleteness, model error, and environmental condition changes, the prediction results will have certain uncertainty. By calculating an uncertainty interval, it is indicated that the real pollutant concentration has a certain probability of falling within the interval. For example, a common method is to calculate a confidence interval based on a confidence interval. Assuming that the prediction value in the prediction result output by the pre-trained deep learning network model is x, by analyzing the error of the pre-trained deep learning network model and the statistical characteristics of the data, a confidence interval [a, b] with a confidence level of 95% is obtained, indicating that the real pollutant concentration has a 95% probability of falling between a and b.

[0088] It should be noted that the size of the uncertainty interval can be an important indicator for evaluating the prediction reliability. Generally, the narrower the uncertainty interval, the more accurate the prediction result and the higher the reliability; on the contrary, the wider the uncertainty interval, the greater the uncertainty of the prediction result and the lower the reliability. In addition, in other embodiments, the reliability of the prediction result can also be evaluated in combination with other indicators, such as mean square error, mean absolute error, etc.

[0089] Based on the same inventive concept, the embodiments of the present application also provide a multi-modal air pollutant prediction device for implementing the multi-modal air pollutant prediction method described above. The implementation scheme of the device for solving the problem is similar to the implementation scheme described in the above method, so the specific limitations in one or more multi-modal air pollutant prediction device embodiments provided below can be referred to the limitations of the multi-modal air pollutant prediction method in the above, which will not be repeated here.

[0090] In an exemplary embodiment, as Figure 4As shown, a multi-modal air pollutant prediction device is provided, which comprises: an acquisition module 401, a preprocessing module 402, a first extraction module 403, a second extraction module 404, and a prediction module 405; wherein,

[0091] The acquisition module 401 is configured to acquire to-be-processed multi-modal data of a target region.

[0092] The preprocessing module 402 is configured to fill in missing values of the to-be-processed multi-modal data and / or remove missing values of the to-be-processed multi-modal data, and perform standardization processing on the to-be-processed multi-modal data after filling in missing values and / or removing abnormal values, to obtain preprocessed multi-modal data; the preprocessed multi-modal data comprises current-time atmospheric pollutant concentration monitoring station data, emission source data, and meteorological field data in the target region.

[0093] The first extraction module 403 is configured to splice the last-time atmospheric pollutant concentration monitoring station data and the current-time atmospheric pollutant concentration monitoring station data along the time dimension, and extract data features of the spliced data, to obtain a station position-time-pollutant concentration feature vector.

[0094] The second extraction module 404 is configured to extract a feature fusion vector of the emission source data and the meteorological field data, and fuse the extracted feature fusion vector with the station position-time-pollutant concentration feature vector, to obtain a multi-modal feature vector.

[0095] The prediction module 405 is configured to input the multi-modal feature vector into a pre-trained deep learning network model, and use the pre-trained deep learning network model to predict the pollutant concentration of the target region, to obtain a prediction result.

[0096] As an optional implementation, the first extraction module 403 is specifically configured to splice the last-time atmospheric pollutant concentration monitoring station data and the current-time atmospheric pollutant concentration monitoring station data along the time dimension, and embed the latitude and longitude relative positions of each monitoring station and the time coding into the spliced last-time atmospheric pollutant concentration monitoring station data and the current-time atmospheric pollutant concentration monitoring station data, to obtain spatio-temporal joint data of different monitoring stations in the target region.

[0097] The data features of the spatio-temporal joint data of different monitoring stations are extracted based on a self-attention mechanism and a feedforward network, to obtain the station position-time-pollutant concentration feature vector.

[0098] As an optional implementation, the second extraction module 404 is specifically configured to perform feature extraction on the emission source data and the meteorological field data based on a preset convolutional neural network to obtain a meteorological-emission coupling multi-scale spatiotemporal feature vector; and perform fusion of the meteorological-emission coupling multi-scale spatiotemporal feature vector and a feature vector of site location-time-pollutant concentration based on a cross-attention mechanism to obtain a multi-modal feature vector.

[0099] As an optional implementation, the multi-modal air pollutant prediction device 400 further includes a calculation module configured to perform inverse standardization processing on the prediction result to obtain a predicted value of an actual pollutant concentration, or calculate an uncertainty interval of the prediction result and evaluate accuracy of the prediction result based on the uncertainty interval to obtain an evaluation result.

[0100] As an optional implementation, the multi-modal air pollutant prediction device 400 further includes a training module configured to obtain historical atmospheric pollutant concentration monitoring site data, historical emission source data and historical meteorological field data in a target region; fill in missing values of the historical atmospheric pollutant concentration monitoring site data, the historical emission source data and the historical meteorological field data and / or remove abnormal values of the historical atmospheric pollutant concentration monitoring site data, the historical emission source data and the historical meteorological field data, and perform standardization processing on the historical atmospheric pollutant concentration monitoring site data, the historical emission source data and the historical meteorological field data with the filled-in missing values and / or the removed abnormal values to obtain historical standard atmospheric pollutant concentration monitoring site data, historical standard emission source data and historical standard meteorological field data; embed longitude and latitude phase position codes and time codes between each monitoring site in the target region into the historical standard atmospheric pollutant concentration monitoring site data, and extract data features of the embedded historical standard atmospheric pollutant concentration monitoring site data to obtain a feature vector of site location-time-historical pollutant concentration; extract a feature fusion vector of the historical emission source data and the historical meteorological field data, and perform fusion of the extracted feature fusion vector and the feature vector of site location-time-historical pollutant concentration to obtain a historical multi-modal feature vector; input the historical multi-modal feature vector into a deep learning network model, perform single-step training on the deep learning network model with a first preset step to obtain a first-stage deep learning network model; input the historical multi-modal feature vector into the first-stage deep learning network model, and perform joint optimization of autoregressive prediction with a second preset step to obtain a pre-trained deep learning network model.

[0101] As an optional implementation, the preprocessing module 402 is configured to obtain the mean and standard deviation of historical atmospheric pollutant concentration monitoring site data, the mean and standard deviation of historical emission source data, and the mean and standard deviation of historical meteorological field data of the target area; fill in missing values of corresponding data of the to-be-processed multi-modal data based on the mean of the historical atmospheric pollutant concentration monitoring site data, the mean of the historical emission source data, and the mean of the historical meteorological field data; remove outliers of corresponding data of the to-be-processed multi-modal data based on the standard deviation of the historical atmospheric pollutant concentration monitoring site data, the standard deviation of the historical emission source data, and the standard deviation of the historical meteorological field data; and perform standardization processing on the corresponding data of the to-be-processed multi-modal data after filling in the missing values or removing the outliers based on the mean and standard deviation of the historical atmospheric pollutant concentration monitoring site data, the mean and standard deviation of the historical emission source data, and the mean and standard deviation of the historical meteorological field data, to obtain the current-time atmospheric pollutant concentration monitoring site data, the emission source data, and the meteorological field data.

[0102] In this implementation, the multi-modal data is constructed based on the current-time atmospheric pollutant concentration monitoring site data, the emission source data, and the meteorological field data in the target area, and the meteorological data, the pollutant monitoring data, and the emission source data can be integrated through the construction of the multi-modal data, so as to achieve the purpose of multi-source data cooperation. Specifically, based on the consideration of the interaction between the meteorological field data and the pollutant concentration field, the pollution information at the local scale and the regional scale is fused, the influence of regional transportation on the pollutant concentration in the target area can be effectively represented, and thus the problem of insufficient data fusion and low prediction accuracy in the prior art can be solved, and the prediction accuracy of the pollutant concentration can be improved to a certain extent.

[0103] In addition, there are a large amount of missing values and abnormal data in the actual pollutant monitoring data, and directly using strict quality control and missing value filling strategies may introduce additional errors, thereby affecting the actual prediction accuracy of the model. Therefore, in order to reduce the dependence of the model on the input data, the input of the prediction model only includes two time instances of Xt-6 and Xt, and the prediction is performed based on an autoregressive strategy to generate a 6-hour resolution future 72-hour pollutant concentration forecast (a total of 12 time steps).

[0104] In an exemplary embodiment, a computer device, which can be a server or a terminal, is provided, and an internal structure diagram of the computer device can be as shown in FIG. 1. Figure 5As shown in the figure. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the prediction data of the multi-modal air pollutants. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with the terminal outside through network connection. The computer program is executed by the processor to realize a multi-modal air pollutant prediction method.

[0105] Those skilled in the art can understand that, Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0106] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to realize the steps in each of the above method embodiments.

[0107] In one exemplary embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to realize the steps in each of the above method embodiments.

[0108] In one exemplary embodiment, a computer program product is provided, including a computer program, which is executed by a processor to realize the steps in each of the above method embodiments.

[0109] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data agreed by the user or agreed by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0110] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. The non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. The volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc.

[0111] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, etc., without being limited thereto.

[0112] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0113] The principles and implementation modes of the present application are described by applying specific examples herein, and the above-mentioned embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In conclusion, the content of the present application should not be understood as a limitation.

Claims

1. A method of prediction of multi-modal air pollutants, characterized in that, The method for predicting the multi-modal air pollutants comprises: acquiring multi-modal data to be processed of a target area; filling in missing values of the multi-modal data to be processed and / or removing missing values of the multi-modal data to be processed, and performing standardization processing on the multi-modal data to be processed after filling in missing values and / or removing abnormal values to obtain preprocessed multi-modal data; the preprocessed multi-modal data comprises current-time atmospheric pollutant concentration monitoring station data, emission source data and meteorological field data in the target area; splicing the atmospheric pollutant concentration monitoring station data of a previous time and the current-time atmospheric pollutant concentration monitoring station data along a time dimension, and extracting data features of the spliced data to obtain a feature vector of station position-time-pollutant concentration; extracting a feature fusion vector of the emission source data and the meteorological field data, and fusing the extracted feature fusion vector with the feature vector of station position-time-pollutant concentration to obtain a multi-modal feature vector; inputting the multi-modal feature vector into a pre-trained deep learning network model, and using the pre-trained deep learning network model to predict the pollutant concentration of the target area to obtain a prediction result; wherein the splicing the atmospheric pollutant concentration monitoring station data of a previous time and the current-time atmospheric pollutant concentration monitoring station data along a time dimension, and extracting data features of the spliced data to obtain a feature vector of station position-time-pollutant concentration comprises: splicing the atmospheric pollutant concentration monitoring station data of a previous time and the current-time atmospheric pollutant concentration monitoring station data along a time dimension, and embedding the relative position coding and time coding of the longitude and latitude between each monitoring station into the spliced atmospheric pollutant concentration monitoring station data of a previous time and the current-time atmospheric pollutant concentration monitoring station data to obtain spatio-temporal joint data of different monitoring stations in the target area; extracting data features of the spatio-temporal joint data of different monitoring stations based on a self-attention mechanism and a feedforward network to obtain the feature vector of station position-time-pollutant concentration; the extracting a feature fusion vector of the emission source data and the meteorological field data, and fusing the extracted feature fusion vector with the feature vector of station position-time-pollutant concentration to obtain a multi-modal feature vector comprises: extracting features of the emission source data and the meteorological field data based on a preset convolutional neural network to obtain a meteorological-emission coupled multi-scale spatio-temporal feature vector; fusing the meteorological-emission coupled multi-scale spatio-temporal feature vector with the feature vector of station position-time-pollutant concentration based on a cross-attention mechanism to obtain a multi-modal feature vector.

2. The method of prediction of multi-modal air pollutants according to claim 1, wherein, The method further comprises: performing inverse standardization processing on the prediction result to obtain a predicted value of actual pollutant concentration; or calculating an uncertainty interval of the prediction result, and evaluating the accuracy of the prediction result based on the uncertainty interval to obtain an evaluation result.

3. The method of prediction of multi-modal air pollutants according to claim 1, wherein, The training method of the pre-trained deep learning network model comprises: obtaining historical atmospheric pollutant concentration monitoring station data, historical emission source data and historical meteorological field data in a target area; filling in missing values of the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data and / or removing outliers of the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data, and performing standardization processing on the historical atmospheric pollutant concentration monitoring station data, the historical emission source data and the historical meteorological field data after filling in the missing values and / or removing the outliers to obtain historical standard atmospheric pollutant concentration monitoring station data, historical standard emission source data and historical standard meteorological field data; embedding longitude and latitude phase position codes and time codes between each monitoring station in the target area into the historical standard atmospheric pollutant concentration monitoring station data, and extracting data features of the embedded historical standard atmospheric pollutant concentration monitoring station data to obtain a feature vector of station position-time-historical pollutant concentration; extracting a feature fusion vector of the historical emission source data and the historical meteorological field data, and fusing the extracted feature fusion vector with the feature vector of station position-time-historical pollutant concentration to obtain a historical multi-modal feature vector; inputting the historical multi-modal feature vector into a deep learning network model, performing single-step training on the deep learning network model with a first preset step size to obtain a first-stage deep learning network model; inputting the historical multi-modal feature vector into the first-stage deep learning network model, and performing joint optimization of autoregressive prediction with a second preset step size to obtain the pre-trained deep learning network model.

4. The method of prediction of multi-modal air pollutants according to claim 1, wherein, The filling in missing values of the to-be-processed multi-modal data and / or removing outliers of the to-be-processed multi-modal data, and the standardization processing on the to-be-processed multi-modal data after filling in the missing values and / or removing the outliers to obtain preprocessed multi-modal data, include: obtaining mean values and standard deviations of historical atmospheric pollutant concentration monitoring station data, mean values and standard deviations of historical emission source data, and mean values and standard deviations of historical meteorological field data of a target area; filling in missing values of corresponding data of the to-be-processed multi-modal data based on the mean values of the historical atmospheric pollutant concentration monitoring station data, the mean values of the historical emission source data and the mean values of the historical meteorological field data; and / or removing outliers of corresponding data of the to-be-processed multi-modal data based on the standard deviations of the historical atmospheric pollutant concentration monitoring station data, the standard deviations of the historical emission source data and the standard deviations of the historical meteorological field data; based on the mean values and standard deviations of the historical atmospheric pollutant concentration monitoring station data, the mean values and standard deviations of the historical emission source data and the mean values and standard deviations of the historical meteorological field data, performing standardization processing on corresponding data of the to-be-processed multi-modal data after filling in the missing values and / or removing the outliers to obtain the current time atmospheric pollutant concentration monitoring station data, the emission source data and the meteorological field data.

5. A device for predicting a multi-modal air pollutant, characterized by, The multi-modal air pollutant prediction device comprises: An acquisition module configured to acquire multi-modal data to be processed of a target region; A preprocessing module configured to fill in or remove missing values and outliers of the multi-modal data to be processed, and to perform standardization processing on the multi-modal data to be processed after filling in the missing values or removing the outliers, to obtain preprocessed multi-modal data; the preprocessed multi-modal data comprises current-time atmospheric pollutant concentration monitoring station data, emission source data, and meteorological field data; A first extraction module configured to splice the current-time atmospheric pollutant concentration monitoring station data and previous-time atmospheric pollutant concentration monitoring station data along a time dimension, and to extract data features of the spliced data, to obtain a station position-time-pollutant concentration feature vector; specifically, the first extraction module splices the current-time atmospheric pollutant concentration monitoring station data and the previous-time atmospheric pollutant concentration monitoring station data along the time dimension, and embeds latitude and longitude relative positions between various monitoring stations and time codes into the spliced current-time atmospheric pollutant concentration monitoring station data and the previous-time atmospheric pollutant concentration monitoring station data, to obtain spatio-temporal joint data of different monitoring stations in the target region; the first extraction module extracts data features of the spatio-temporal joint data of the different monitoring stations based on a self-attention mechanism and a feedforward network, to obtain the station position-time-pollutant concentration feature vector; A second extraction module configured to extract a feature fusion vector of the emission source data and the meteorological field data, and to further fuse the extracted feature fusion vector with the station position-time-pollutant concentration feature vector, to obtain a multi-modal feature vector; specifically, the second extraction module extracts features of the emission source data and the meteorological field data based on a preset convolutional neural network, to obtain a meteorological-emission coupled multi-scale spatio-temporal feature vector; the second extraction module fuses the meteorological-emission coupled multi-scale spatio-temporal feature vector with the station position-time-pollutant concentration feature vector based on a cross-attention mechanism, to obtain the multi-modal feature vector; A prediction module configured to input the multi-modal feature vector into a pre-trained deep learning network model, and to predict a pollutant concentration of the target region by using the pre-trained deep learning network model, to obtain a prediction result.

6. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the multi-modal air pollutant prediction method according to any one of claims 1-4.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the multi-modal air pollutant prediction method according to any one of claims 1-4.

8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the multi-modal air pollutant prediction method according to any one of claims 1-4.