Meteorological Element Forecasting Method Based on Graph Neural Network and Multimodal Meteorological Data Fusion
By combining geostationary meteorological satellite data and ground meteorological station observation data, a graph neural network framework for multimodal data fusion was constructed, which solved the problems of single-level and single-structure graphs in ground meteorological station observation data, and realized all-weather, high-precision meteorological element forecasting.
Patent Information
- Application Number
- CN202310751074.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2043-06-25
AI Technical Summary
In existing technologies, the observation data from ground meteorological stations are of a single level, and the graph neural network in spatiotemporal forecasting has a single graph structure and defects in the multi-graph fusion method, which leads to limited accuracy and range of meteorological forecasts and makes it difficult to achieve high-precision all-weather meteorological element forecasts.
A multimodal meteorological data fusion method based on graph neural networks is adopted. By utilizing geostationary meteorological satellite data and ground meteorological station observation data, a satellite-observation station multimodal data fusion framework is constructed. The attention mechanism of Transformer is used for modal interaction, and a graph convolutional neural network framework for multi-graph fusion is constructed to explore the intrinsic relationship between station geographical location and meteorological elements, so as to realize multimodal feature fusion and adaptive learning.
It improves the accuracy and robustness of weather forecasts, enabling all-weather, high-precision forecasting of meteorological elements from ground observation stations, thus enhancing the accuracy of forecast results and the adaptability of models.
Smart Images

Figure CN116720156B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of numerical weather prediction technology, and in particular to a method for forecasting meteorological elements at ground observation stations based on graph neural network multimodal meteorological data fusion. Background Technology
[0002] Weather forecasting refers to the estimation and prediction of future trends in meteorological characteristics within a specified spatial range. It includes forecasting multiple meteorological elements and can be viewed as a complex multivariate spatiotemporal series prediction problem. In the past few decades, numerical weather prediction (NWP) has been a widely used method in the international meteorological forecasting field. Based on actual atmospheric conditions and certain initial and boundary conditions, NWP quantifies atmospheric states and performs numerical calculations to solve the fluid dynamics and thermodynamic equations governing weather evolution, predicting atmospheric motion and weather phenomena over a future period. The calculations are typically performed using supercomputers or distributed computing clusters.
[0003] NWP is a quantitative and objective forecast, but many problems remain, such as parameterization of the gridding process, uncertainty of the initial field, finiteness of numerical models in describing atmospheric motion physical processes, inaccuracy of numerical solutions to nonlinear equations, and high requirements for computer resources and capabilities.
[0004] In recent years, meteorological researchers have introduced data-driven approaches into weather forecasting tasks, achieving considerable breakthroughs and successes. Data-driven approaches utilize historical meteorological observation data accumulated over many years as input to learn the mapping relationship between input and output.
[0005] CN115271062A discloses a method and system for sea fog visibility forecasting based on a deep convolutional neural network (DCNN) model. The method divides a day into preset time periods to obtain forecast periods. A preset number of DCNN models are established for each forecast period, and these models are trained using a sample dataset to obtain sea fog visibility forecasting models corresponding to different forecast lead times for each forecast period. By constructing sea fog visibility forecasting models for each forecast period and corresponding forecast lead times within each day's forecast periods, the method allows for the selection of the appropriate sea fog visibility forecasting model based on the current forecast period and the required forecast lead time during actual forecasting. This effectively improves the accuracy of model predictions and solves the problem that current conventional forecast products have low accuracy and cannot adequately meet the meteorological support service needs of port vessel scheduling and berth operations. However, this method, starting from sea fog visibility forecasting, uses a relatively basic deep convolutional neural network model. While effective for single-point predictions, it cannot perform overall spatiotemporal forecasting for multiple observation stations or forecasting other meteorological elements besides visibility.
[0006] In recent years, with the rapid development of deep learning technology, Graph Neural Networks (GNNs) use neural networks to learn graph-structured data, extract and discover features and patterns in graph-structured data, and capture the internal dependencies of data by modeling data in non-Euclidean space. Its excellent ability to process unstructured data has been fully demonstrated in traffic flow prediction applications, while research on its application in the field of weather forecasting is still relatively limited.
[0007] CN112232543A discloses a multi-site prediction method based on graph convolutional networks (GCNNs). This method improves upon GCNNs by leveraging their advantage in processing non-Euclidean data, proposing feature extraction in both spatial and temporal dimensions. An attention mechanism is then introduced to enhance model performance, resulting in some improvement over other models in multi-site atmospheric visibility prediction. However, this method relies solely on atmospheric visibility data, resulting in a single data level. Using meteorological data from a single source makes accurate weather forecasts difficult. For example, relying solely on ground-based meteorological station observations presents challenges such as uneven geographical distribution, significant differences in sensor performance, missing data, and anomalies. The observation environment is also susceptible to interference from surrounding obstacles, buildings, and human activities, limiting the accuracy and range of weather forecasts. Furthermore, while this method uses a basic GCNN architecture for multi-site prediction, it doesn't delve deeply into the graph structure, leading to a limitation in spatiotemporal forecasting due to its simplistic graph structure.
[0008] In meteorological forecasting tasks, which are highly susceptible to interference from various factors, single modal data often cannot contain all the effective information required to produce accurate weather forecast results. CN113919231A discloses a method and system for predicting the spatiotemporal variation of PM2.5 concentration based on spatiotemporal graph neural networks. Using observational data from approximately 1500 atmospheric monitoring stations nationwide as a training set, and combining multiple data sources such as meteorology and elevation, a unified prediction framework is constructed using spatiotemporal graph neural networks. This framework can simultaneously predict PM2.5 concentration changes over a large area and improve prediction accuracy. Although this method uses meteorological data from multiple sources, it only uses a direct concatenation method to construct the feature matrix when building samples, fusing features from different modalities. This leads to data redundancy, dependency, and high dimensionality, ignoring the unique statistical properties of modalities and the interaction relationships between modalities. Its generalization ability and adaptability to specific tasks need improvement.
[0009] Therefore, in order to overcome the problem of the single level of observation data from ground meteorological stations, and the shortcomings of existing graph neural network technology in spatiotemporal forecasting such as the single graph structure and the defects in multi-graph fusion methods, it is necessary to propose a multimodal meteorological data fusion method for forecasting meteorological elements from ground observation stations. Summary of the Invention
[0010] The purpose of this invention is to propose a meteorological element forecasting method based on graph neural network multimodal meteorological data fusion. This method leverages the complementary advantages of multimodal meteorological data (geostationary meteorological satellite data and ground meteorological station observation data) (top-down remote sensing observation + bottom-up ground observation) to complete meteorological forecasting tasks, compensating for the shortcomings of single-source data. It constructs a satellite-observation station multimodal data fusion framework to fuse multimodal features and proposes a graph convolutional neural network framework based on multi-graph fusion. This framework mines the relationships between station geographical locations and the intrinsic connections between different meteorological elements from multiple perspectives, constructing various static and dynamic graphs. Through adaptive learning, it fuses time-series multi-graph features to achieve all-weather, high-precision ground observation station meteorological element forecasting.
[0011] To achieve the above objectives, the present invention provides the following technical solution:
[0012] This invention proposes a meteorological element forecasting method based on graph neural network multimodal meteorological data fusion, comprising the steps of preprocessing geostationary meteorological satellite data, preprocessing ground meteorological station observation data, satellite image feature extraction, satellite-observation station multimodal data fusion, graph structure construction, and construction of a graph convolutional neural network for multi-graph fusion; wherein:
[0013] The steps involved in satellite-observation station multimodal data fusion include:
[0014] S11. Input the sequences of ground observation data and satellite image features into a one-dimensional temporal convolutional layer and add position encoding to obtain low-level position-aware features of ground observation data and satellite image data respectively.
[0015] S12. For data from two modalities, use the cross-modal Transformer multi-head attention mechanism to enable modal interaction, allowing one modality to receive information from the other modality.
[0016] S13. Each modality continuously updates its sequence through external information from its own cross-modal Transformer multi-head attention mechanism. Then, through the Transformer self-attention layer, the outputs of the self-attention layers of the two modal branches are spliced together in the channel dimension to obtain the satellite-observation station multimodal fusion features.
[0017] The steps involved in constructing a graph convolutional neural network for multi-graph fusion include:
[0018] S21. Use learnable weights to describe the importance of the graph structure to each ground observation station to obtain the fused graph structure;
[0019] S22. The spatiotemporal sequence of satellite-observation station multimodal fusion features and the fused graph structure are used as inputs to a multi-graph convolutional neural network to predict the spatiotemporal sequence of a specified meteorological element in the future.
[0020] The multi-graph convolutional neural network described above employs stacked spatial-temporal convolutional blocks. These blocks use residual connections, employing graph convolution in the spatial dimension and multi-branch convolutional layers in the temporal dimension. Different branches of the multi-branch convolutional layers have different receptive fields to extract information at different scales. Then, information at different scales is fused through merging operations and convolutional layers.
[0021] Furthermore, in step S11, the low-level position-aware features of the ground observation data are represented as follows:
[0022]
[0023] Low-level position-aware features of satellite image data are represented as follows:
[0024]
[0025] Where X is the input sequence, Conv1D represents a one-dimensional temporal convolutional layer, and PE represents positional encoding.
[0026] Furthermore, in step S12, when using the cross-modal Transformer multi-head attention mechanism for modal interaction, the process of a single-head cross-modal attention mechanism from one modality β to another modality α is as follows: The modality α is associated with its weight matrix. Q is obtained through matrix multiplication. α Modality β is associated with its weight matrix. and K is obtained through matrix multiplication. β and V β The output of the cross-modal attention mechanism is:
[0027]
[0028] Where, d k The number of feature channels, For K β The transpose of .
[0029] Furthermore, in step S13, each cross-modal Transformer multi-head attention mechanism includes D layers of cross-modal multi-head attention modules, and the feedforward calculation of each layer of cross-modal multi-head attention modules is as follows:
[0030]
[0031]
[0032]
[0033] in, These represent the low-level position-aware features of modal α and β, respectively. LN represents the cross-modal multi-head attention module at layer i, and f represents layer normalization. θ It is a feedforward layer parameterized by θ. This represents the computation result of cross-modal multi-head attention for the i-th layer module. This represents the final output of the i-th layer module.
[0034] Furthermore, the steps for constructing a graph structure include:
[0035] The graph structure G is represented as G = (V, E, A), where V, E, and A represent the set of locations of ground observation stations, the set of edges connecting the stations, and the adjacency matrix, respectively; the constructed graph structure includes the distance graph G. D Pattern similarity graph G P and dynamic graph G K ;
[0036] Distance-based graph structure G D = (V, E, A) D The spherical distances between ground observation station locations were calculated using Gaussian kernels and then filtered by setting a threshold. D The element is defined as follows:
[0037]
[0038] Where, d ij Indicates ground observation station v i and v j The spherical distance between them, ε and Used to control A D sparsity and distribution;
[0039] A pattern similarity graph G is constructed based on the Pearson correlation coefficient. P = (V, E, A) P ), A P The element is defined as follows:
[0040]
[0041] Where f represents specific meteorological elements from ground observation station data or extracted satellite image features. It is a ground observation station v i Time series of length P It is v i The value of f at time step p;
[0042] Constructing dynamic graph G K = (V, E, A) K The nonlinear spatiotemporal correlation of the input data is modeled, and the ground observation station v i The input time series X at time t t,i ={x t-w+1,i , ..., x t-1,i , ..., x t,i}∈R W×D Where W is the sequence length, D is the total dimension of the features, and X... t,i Flattened into vector Z i ∈R WD A K The element is defined as follows:
[0043] D i =tanh(λZ) i W1)
[0044] D j =tanh(λZ) j W2)
[0045] A K,ij =ReLU(tanh(λ(D)) i ·D j -D j ·D i )))
[0046] Where W1 and W2 are linear layers, λ is a hyperparameter used to control the saturation of the activation function, and D... i and D j These represent the calculated hidden features.
[0047] Furthermore, the specific process of step S21 is as follows:
[0048] Use learnable weights W s ∈R N×N Let s be used to describe the importance of the graph structure s to each ground observation station, where N represents the number of ground observation stations, and s is S = {G} D G P G K In a graph structure G, D For distance graphs, G P For pattern similarity graphs, G K For dynamic graphs, the strategy for multi-graph fusion is as follows:
[0049]
[0050] Where ⊙ represents the element-wise multiplication operation of the matrix, A s This is the adjacency matrix corresponding to the graph structure;
[0051] The final merged graph structure is as follows:
[0052] G fused = (V, E, A) fused )
[0053] Among them, V, E, A fused These represent the set of locations of ground observation stations, the set of edges connecting the stations, and the adjacency matrix of the multi-graph fusion, respectively.
[0054] Furthermore, the specific process of step S22 is as follows:
[0055] At time t, the multimodal fusion feature spatiotemporal sequence {X′} with time length W′. (t-W′+1) , ..., X′ (t)} and the merged graph structure G fused These are used as inputs to a multi-graph convolutional neural network, with the goal of learning a function P to predict the spatiotemporal sequence of a given meteorological element F over future time W″. Right now:
[0056]
[0057] The information of meteorological element f at time t on the graphical structure is represented as follows: N is the number of ground observation stations, determined by the kernel function g. θ Filter the information on the graph structure G, that is, respectively for gθ Perform a spectral domain Fourier transform on x, multiply the transformed results, and then perform an inverse Fourier transform to obtain the output of the graph convolution, represented as:
[0058] g g *Gx=g g (L)x=g g (UAU T )x=Ug θ (Λ)U T x
[0059] Fourier basis U∈R of the graph N×N It is the normalized Laplace matrix The eigenvector matrix, I N It is the identity matrix, D is the degree matrix, Λ is the diagonal matrix of eigenvalues of L, and g θ (Λ) is a diagonal matrix.
[0060] Furthermore, the eigenvalue decomposition process is handled using Chebyshev polynomials, expressed as:
[0061]
[0062] in, It is a scaled Laplace matrix The k-th order Chebyshev polynomial obtained at λ max is the largest eigenvalue of the normalized Laplacian matrix L, and K is the kernel size of the graph convolution, which determines the maximum radius of the convolution starting from the center node.
[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0064] 1. This invention adopts a data-driven approach and proposes a meteorological element forecasting method based on graph neural network multimodal meteorological data fusion. It uses geostationary meteorological satellite data and ground meteorological station observation data as inputs, representing two major levels: top-down remote sensing observation and bottom-up ground observation, respectively. After extracting image semantic features and meteorological station sensor observation features, a satellite-observation station multimodal data fusion framework based on the Transformer attention mechanism is used. This fully utilizes multimodal information interaction to extract fusion features and perform multimodal feature fusion. This is an end-to-end model that emphasizes the mutual influence between modes, resulting in better multimodal data representation and more efficient use of meteorological data. It solves the problem of single-level ground meteorological station observation data and significantly improves the accuracy of prediction results and enhances the robustness of the prediction model.
[0065] 2. This invention addresses the shortcomings of existing graph neural network technologies in spatiotemporal forecasting, such as the single graph structure and the defects in multi-graph fusion methods. It constructs a graph convolutional neural network framework based on multi-graph fusion, which mines the relationships between geographical locations of stations and the intrinsic connections between different meteorological elements from multiple perspectives. It constructs graph structure relationships of various static and dynamic graphs respectively, and further fuses temporal multi-graph features through adaptive learning of the relationships between various graphs. The graph convolutional neural network framework is used to model spatial information and topological relationships to achieve all-weather, high-precision forecasting of meteorological elements from ground observation stations. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0067] Figure 1 The flowchart of the ground observation station forecasting method based on graph neural network for multimodal meteorological data fusion provided by the present invention is shown below.
[0068] Figure 2 This is a framework diagram for satellite-observation station multimodal data fusion provided by the present invention.
[0069] Figure 3 The diagram shows the multi-graph fusion graph convolutional neural network framework provided by this invention. Detailed Implementation
[0070] To better understand this technical solution, the method of the present invention will be described in detail below with reference to the accompanying drawings.
[0071] This invention proposes a meteorological element forecasting method based on graph neural network multimodal meteorological data fusion, the process of which is as follows: Figure 1 As shown, the process includes preprocessing of geostationary meteorological satellite data, preprocessing of ground meteorological station observation data, satellite image feature extraction, satellite-observation station multimodal data fusion, graph structure construction, and construction of a graph convolutional neural network for multi-graph fusion. Each step is described in detail below.
[0072] 1. Introduction to the data sources used in this invention
[0073] 1) Geostationary meteorological satellite data
[0074] Geostationary meteorological satellites orbit approximately 35,800 kilometers above the Earth's equator, operating synchronously with the Earth's rotation. Relative to the Earth, they are stationary and can continuously observe meteorological conditions over nearly one-third of the Earth's surface (approximately 100 million square kilometers). These satellites carry various meteorological remote sensing instruments capable of receiving and measuring visible light, infrared, and microwave radiation from the Earth and its atmosphere. This data is converted into electrical signals and transmitted to the ground in real time. These signals can be reconstructed into images of clouds, the Earth's surface, and ocean surfaces. Further processing and calculations yield various meteorological data. The satellite data used in this invention comes from Japan's next-generation geostationary meteorological satellites—Himawari 8 / 9—equipped with the advanced Himawari Imager (AHI).
[0075] 2) Observational data from ground meteorological stations
[0076] Ground-based meteorological stations are a major component of the weather, climate, and climate change observation network, enabling timely and reliable collection and analysis of various meteorological data. The CMA (China Meteorological Administration) provides hourly observation data from over 2,000 national-level ground-based meteorological stations, recording observations of meteorological elements such as temperature, humidity, air pressure, and precipitation.
[0077] 2. Meteorological data preprocessing
[0078] 1) Preprocessing of geostationary meteorological satellite data
[0079] To ensure the continuity of daytime and nighttime observation data, the infrared channels (channels 7-16) of the Himawari-8 / 9 meteorological satellites were selected for processing. By consulting the data calibration table provided by the Japan Meteorological Agency, the raw satellite observation data was transformed into meaningful physical variable data. Using common methods in digital image processing, the physical variable data were normalized, mapped to the 0-255 range, and specific areas were cropped. Grayscale images of each channel were saved with a temporal resolution of 1 hour and a spatial resolution of 0.02 degrees.
[0080] 2) Preprocessing of observation data from ground meteorological stations
[0081] Hourly observation data from national-level surface meteorological stations provided by the China Meteorological Administration were used to perform station statistics, meteorological element screening, and processing of missing and outlier values. The final results yielded 20 meteorological elements: air pressure, water vapor pressure, air temperature, 1-hour maximum temperature, 1-hour minimum temperature, dew point temperature, surface temperature, relative humidity, wind speed, wind direction, 1-hour maximum wind speed, 1-hour maximum wind direction, vertical visibility, 1-minute horizontal visibility, 10-minute horizontal visibility, and 1 / 3 / 6 / 12 / 24-hour cumulative precipitation, as well as 3 geographic location information: longitude, latitude, and altitude.
[0082] 3. Satellite image feature extraction
[0083] Image feature extraction refers to the process of extracting a set of feature vectors or feature descriptors representing the content of an image, typically for applications such as image classification, retrieval, and matching. Commonly used image features include color, shape, texture, and edges. This invention uses a ResNet101 network pre-trained on ImageNet to extract high-level semantic information from satellite images. Only image features of grids aligned with the geographical location of ground observation stations are saved. Furthermore, considering the multispectral characteristics of satellite images, this invention extracts image features from each channel of satellite data separately, stitches them together along the channel dimension, and uses them in subsequent processes.
[0084] 4. Satellite-observation station multimodal data fusion
[0085] The satellite-observation station multimodal data fusion framework used in this invention is as follows: Figure 2 As shown.
[0086] To ensure that each element of the input sequence X has sufficient awareness of its neighbors, this invention passes the input sequence through a 1D (one-dimensional) temporal convolutional layer. To enable the sequence to carry temporal information, positional encoding (PE) is added, resulting in low-level position-aware features Z of different modalities. [0] Conv1D represents a one-dimensional temporal convolutional layer.
[0087]
[0088] Low-level position-aware features of satellite image data are represented as follows:
[0089]
[0090] Where X is the input sequence, Conv1D represents a one-dimensional temporal convolutional layer, and PE represents positional encoding.
[0091] For data from two modalities α and β, the cross-modal Transformer self-attention mechanism is adjusted to enable modal interaction, allowing one modality to receive information from the other. Taking the single-head cross-modal attention from β to α as an example, modality α and its weight matrix are... Q is obtained through matrix multiplication. α Modality β is associated with its weight matrix. and K is obtained through matrix multiplication. β and V β The output of the cross-modal attention mechanism is:
[0092]
[0093] Each cross-modal Transformer self-attention mechanism contains D layers of cross-modal attention modules, and the feedforward calculation for each layer is as follows:
[0094]
[0095]
[0096]
[0097] in, These represent the low-level position-aware features of modal α and β, respectively. LN represents the cross-modal multi-head attention module at layer i, and f represents layer normalization. θ It is a feedforward layer parameterized by θ. This represents the computation result of cross-modal multi-head attention for the i-th layer module. This represents the final output of the i-th layer module.
[0098] In this process, each modality continuously updates its sequence with external information from the cross-modal multi-head attention module. After passing through the cross-modal attention layer, it also passes through a self-attention layer based on the original Transformer architecture. Finally, the outputs of the self-attention layers of the two branches are concatenated along the channel dimension to obtain the satellite-observation station multimodal fusion feature X′.
[0099] 5. Graph Structure Construction
[0100] The graph structure G is represented as G = (V, E, A), where V, E, and A represent the set of locations of ground observation stations, the set of edges connecting stations, and the adjacency matrix, respectively. The constructed graph structure includes the distance graph G. D Pattern similarity graph G P and dynamic graph G K .
[0101] 1) Distance graph G D
[0102] Distance-based graph structure G D = (V, E, A) D The distance map G in this invention can describe the topological structure of the geographical location of ground observation stations. D It is constructed using Gaussian kernels to calculate the spherical distances between ground observation station locations, and then filtered by setting a threshold. D The element is defined as follows:
[0103]
[0104] Where, d ij Indicates ground observation station vi and v j The spherical distance between them, ε and Used to control A D The sparsity and distribution.
[0105] 2) Pattern similarity graph G P
[0106] Meteorological elements from ground observation stations that are geographically distant may also exhibit highly consistent characteristics. Therefore, this invention constructs a pattern similarity graph G based on the Pearson correlation coefficient. P = (V, E, A) P To uncover this similarity relationship. P The element is defined as follows:
[0107]
[0108] Where f represents specific meteorological elements from ground observation station data or extracted satellite image features. It is a ground observation station v i Time series of length P It is v i The value of f at time step p.
[0109] 3) Dynamic graph G K
[0110] The present invention constructs a dynamic diagram G K = (V, E, A) K The nonlinear spatiotemporal correlation of the input data is modeled. Ground observation station v i The input time series X at time t t,i ={x t-W+1,i , ..., x t-1,i , ..., x t,i}∈R W×D Where W is the sequence length, D is the total dimension of the features, and X... t,i Flattened into vector Z i ∈R WD Similar to the construction of learnable graphs, A K The element is defined as follows:
[0111] D j =tanh(λZ) i W1)
[0112] D j =tanh(λZ) j W2)
[0113] A K,ij =ReLU(tanh(λ(D)) i·D j -D j ·D i )))
[0114] Where W1 and W2 are linear layers, λ is a hyperparameter used to control the saturation of the activation function, and D... i and D j These represent the calculated hidden features. G K Both the training and inference phases are dynamically generated.
[0115] 6. Graph Convolutional Neural Networks for Multi-Graph Fusion
[0116] Use learnable weights W s ∈R N×N Let s be used to describe the importance of the graph structure s to each ground observation station, where N represents the number of ground observation stations, and s is S = {G} D G P G K A graph structure within a given area. The strategy for multi-graph fusion is as follows:
[0117]
[0118] Where ⊙ represents the element-wise multiplication operation of the matrix, A s This is the adjacency matrix corresponding to the graph structure;
[0119] The final merged graph structure is as follows:
[0120] G fused = (V, E, A) fused )
[0121] Among them, V, E, A fused These represent the set of locations of ground observation stations, the set of edges connecting the stations, and the adjacency matrix of the multi-graph fusion, respectively.
[0122] The multi-graph fusion graph convolutional neural network framework used in this invention is as follows: Figure 3 As shown.
[0123] At time t, the multimodal fusion feature spatiotemporal sequence {X′} with time length W′. (t-W′+1) , ..., X′ (t)} and the merged graph structure G fused These are used as inputs to a multi-graph convolutional neural network, with the goal of learning a function P to predict the spatiotemporal sequence of a given meteorological element θ over future time W″. Right now:
[0124]
[0125] Considering the importance of spatiotemporal consistency in weather forecasting, the multi-graph convolutional neural network framework is designed as a stacked spatial-temporal convolutional block (ST-Block). The ST-Block uses residual connections and graph convolution in the spatial dimension. The information of meteorological element f at time t is represented in the graph structure as follows: N is the number of ground observation stations, determined by the kernel function g. θ Filter the information on the graph structure G, that is, respectively for g θ Perform a spectral domain Fourier transform on x, multiply the transformed results, and then perform an inverse Fourier transform to obtain the output of the graph convolution, represented as:
[0126] g θ *Gx=g θ (L)x=g θ (UΛU T )x=Ug θ (Λ)U T x
[0127] Fourier basis U∈R of the graph N×N It is the normalized Laplace matrix The eigenvector matrix, I N It is the identity matrix, D is the degree matrix, Λ is the diagonal matrix of eigenvalues of L, and g θ (Λ) is a diagonal matrix. The above formula can be understood as relating g to g respectively. θ Perform a spectral domain Fourier transform on x, multiply the transformed results, and then perform an inverse Fourier transform to obtain the output of the graph convolution.
[0128] Due to the large scale of the graph structure, Chebyshev polynomials are used to address the efficiency issue in eigenvalue decomposition.
[0129]
[0130] It is a scaled Laplace matrix The k-th order Chebyshev polynomial obtained at λ max is the largest eigenvalue of the normalized Laplacian matrix L, and K is the kernel size of the graph convolution, which determines the maximum radius of the convolution starting from the center node.
[0131] ST-Block employs multi-branch convolutional layers in the temporal dimension. Different branches of these convolutional layers have different receptive fields (e.g., 1×1, 1×3, 1×5) to extract information at different scales. This information is then fused through a merging operation and a 1×1 convolutional layer. Considering the importance of spatiotemporal consistency in weather forecasting, parameters in ST-Block are adjusted to balance the spatial and temporal spans. The kernel size K of the graph convolution and the receptive field of the temporal convolution increase linearly with the stacking of modules.
[0132] The output of the last ST-Block is passed through a fully connected layer to generate a forecast result for the specified meteorological element F. During the training phase, the true value Y was observed at the ground weather station of F. F Calculate the MAE loss, and output and save the forecast results directly during the inference phase.
[0133] Compared with the prior art, the advantages of the present invention are:
[0134] 1. This invention has the complementary advantages of multimodal data: it uses top-down geostationary meteorological satellite data and bottom-up ground meteorological station observation data to complete the weather forecasting task, and the raw data is easy to obtain.
[0135] 2. The present invention has high accuracy in forecasting meteorological elements: Based on the Transformer-based multimodal data fusion method and the graph convolutional neural network framework for multi-graph fusion, it effectively improves the evaluation indicators such as MAE and RMSE of time and space meteorological forecasts.
[0136] 3. The universality of the framework of this invention in the field of weather forecasting: It is applicable to the prediction of various meteorological elements without requiring too much modification to the overall algorithm.
[0137] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. However, these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A weather element prediction method based on graph neural network multi-modal meteorological data fusion, characterized in that, The steps include preprocessing of static meteorological satellite data, preprocessing of ground meteorological station observation data, satellite image feature extraction, satellite-observation station multi-modal data fusion, graph structure construction, and multi-graph fusion graph convolutional neural network construction. The satellite image feature extraction step includes extracting high-level semantic information of the satellite image using a ResNet101 network pre-trained on ImageNet, and saving only the image features of the grid aligned with the geographical location of the ground observation station; and extracting image features of each channel data of the satellite and splicing in the channel dimension. The satellite-observation station multi-modal data fusion step includes: S11, inputting the sequence of ground observation data and satellite image features into a one-dimensional time convolution layer respectively, and adding position encoding to obtain low-level position perception features of the ground observation data and satellite image data respectively; S12, using a cross-modal Transformer multi-head attention mechanism to interact between the two modalities, so that one modality can receive information from the other modality; S13, each modality updates its own sequence through the external information of the respective cross-modal Transformer multi-head attention mechanism, and then splices the outputs of the self-attention layers of the two modality branches in the channel dimension to obtain satellite-observation station multi-modal fusion features; The steps of the multi-graph fusion graph convolutional neural network construction include: S21, using a learnable weight to describe the importance of the graph structure to each ground observation station to obtain a fused graph structure; S22, inputting the satellite-observation station multi-modal fusion feature spatiotemporal sequence and the fused graph structure into a multi-graph convolutional neural network to predict the spatiotemporal sequence of a specified meteorological element at a future time; The multi-graph convolutional neural network adopts a stacked spatial convolution-temporal convolution block, which uses a residual connection method, uses graph convolution in the spatial dimension, and uses a multi-branch convolution layer in the time dimension. The convolution layers in different branches of the multi-branch convolution layer have different receptive fields for extracting information of different scales, and then fuse the information of different scales through a merging operation and a convolution layer.
2. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, In step S11, the low-level position perception features of the ground observation data are represented as: The low-level position perception features of the satellite image data are represented as: Where X is the input sequence, Conv1D represents a one-dimensional time convolution layer, and PE represents position encoding.
3. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, When the step S12 uses the cross-modal Transformer multi-head attention mechanism for modal interaction, the process of a single head cross-modal attention mechanism of one mode β to another mode α is as follows: the mode α is multiplied by its weight matrix Q is obtained by matrix multiplication α The mode β is multiplied by its weight matrix and K is obtained by matrix multiplication β and V β The output of the cross-modal attention mechanism is: where d k is the number of feature channels, is K β transpose.
4. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, In step S13, each cross-modal Transformer multi-head attention mechanism includes D layers of cross-modal multi-head attention modules. The feedforward calculation of each layer of cross-modal multi-head attention module is as follows: wherein, respectively denote modal a, b low-level position-aware features, represents the cross-modal multi-head attention module at the i-th layer, LN represents layer normalization, f θ is a feed-forward layer parameterized by θ, denotes the computation result of the cross-modal multi-head attention of the i-th layer module, denotes the final output of the i-th layer module.
5. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, The steps of the graph structure construction include: The graph structure G is represented as G = (V, E, A), where V, E, A represent the set of locations of ground observation stations, the set of edges connecting the stations, and the adjacency matrix, respectively; the constructed graph structure includes a distance graph G D , a pattern similarity graph G P , and a dynamic graph G K ; Graph structure G constructed based on distance D = (V, E, A D ) constructed based on spherical distance between ground observation station positions calculated using Gaussian kernel and filtered by setting threshold, A D Elements of A are defined as follows: where d ij represents the spherical distance between the ground observation station v i and v j ; ε and are used to control the sparsity and distribution of A D . constructing a pattern similarity graph G based on the Pearson correlation coefficient P = (V, E, A P ), A P is defined as follows: where f is a particular meteorological element in the ground observation station data or extracted satellite image feature, is the length of the time series i of ground observation station v is the length of the time series i f value at time step p; Constructing dynamic graph G K =(V,E,A) K The nonlinear spatiotemporal correlation of the input data is modeled, and the ground observation station v i The input time series X at time t t,i ={x t-W+1,i , ..., x t-1,i , ..., x t,i }∈R W×D Where W is the sequence length, D is the total dimension of the features, and X... t,i Flattened into vector Z i ∈R WD A K The element is defined as follows: D i = tanh(λZ i W1) D j = tanh(λZ j W2) A K,ij = ReLU(tanh(λ(D i ·D j -D j ·D i ))) where W1and W2are linear layers, l is a hyperparameter used to control the saturation of the activation function, D i and D j denote the computed hidden features, respectively.
6. The weather element prediction method based on graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, The specific process of step S21 is as follows: Using learnable weights W s ∈ R N×N to describe the importance of graph structure s to each ground observation station, where N represents the number of ground observation stations, s is a certain graph structure in S = {G D , G P , G K}, G D is a distance graph, G P is a pattern similarity graph, and G K is a dynamic graph. The strategy of multi-graph fusion is as follows: where is the matrix multiplication operation of the corresponding elements of matrices A s is the adjacency matrix of the corresponding graph structure; The final fused graph structure is as follows: G fused = (V, E, A fused ) wherein, wherein V, E, A fused respectively represent a set of locations of ground observation stations, a set of edges connecting the stations, and an adjacency matrix of multi-graph fusion.
7. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 1, characterized in that, The specific process of step S22 is as follows: At time t, the multimodal fusion feature spatiotemporal sequence {X' t (t-W′+1) , X' t (t)} with a time length of W' and the fused graph G fused are respectively taken as inputs of a multi-graph convolutional neural network, and the goal is to learn a function P to predict the spatiotemporal sequence of a specified meteorological element F in the future W'' time , that is: The information of the meteorological element f at time t on the graph structure is represented as N is the number of ground observation stations, and g θ Filtering the information on the graph structure G, i.e., respectively performing spectral domain Fourier transform on g θ And x, multiplying the transformed results, and performing Fourier inverse transform to obtain the output of the graph convolution, represented as: g θ *Gx = g θ (L)x = g θ (UΛU T )x = Ug θ (Λ)U T x The Fourier basis U e R of a graph N×N is the normalized Laplacian matrix is the eigenvector matrix of the normalized Laplacian matrix L, I N is the identity matrix, D is the degree matrix, A is the diagonal matrix of the eigenvalues of L, g θ (Λ) is the diagonal matrix.
8. The weather element prediction method based on the graph neural network multi-modal meteorological data fusion according to claim 7, characterized in that, The eigenvalue decomposition process is handled using Chebyshev polynomials, represented as: where, are k-th order Chebyshev polynomials obtained at the scaled Laplacian matrix are k-th order Chebyshev polynomials obtained at the scaled Laplacian matrix max is the largest eigenvalue of the normalized Laplacian matrix L, and K is the kernel size of the graph convolution, which determines the maximum radius of the convolution starting from the center node.
Citation Information
Patent Citations
Multi-site prediction method based on graph convolution network
CN112232543A
PM2.5 concentration space-time change prediction method and system based on space-time diagram neural network
CN113919231A