A ground-based meteorological element forecasting method based on hybrid expert models and graph attention.
By combining expert models with graph attention mechanisms, the problem of spatiotemporal heterogeneity adaptation in meteorological element forecasting at ground stations was solved, achieving high-accuracy forecasting in complex weather scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-02-12
- Publication Date
- 2026-04-21
AI Technical Summary
Existing models struggle to dynamically adapt to the spatiotemporal heterogeneity of meteorological elements at ground stations, resulting in low forecast accuracy and poor generalization ability in complex weather scenarios.
By employing a hybrid expert model and graph attention mechanism, and through the collaborative optimization of multi-expert architecture, scene-aware routing, and joint loss function, the spatiotemporal correlation between meteorological elements at ground stations is captured, thereby improving forecast accuracy.
It effectively improves the accuracy of site element forecasts in complex scenarios, dynamically adapts to spatiotemporal dependencies, and enhances forecast precision.
Smart Images

Figure CN121724065B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of surface weather forecasting technology, and in particular relates to a method for forecasting meteorological elements at surface stations based on a hybrid expert model and graph attention mechanism. Background Technology
[0002] Weather forecasting is a long-term and practically significant scientific challenge, closely related to human society and economic activities. With the increasing frequency of extreme weather events, the demand for accurate forecasts of local meteorological elements has grown significantly. Traditional numerical weather prediction (NWP) models are prone to oversmoothing of local small-scale weather processes and have high computational costs, limiting their effectiveness in simulating local meteorological characteristics. In recent years, data-driven deep learning methods have provided a new paradigm for weather forecasting. These methods learn potential patterns by mining the spatiotemporal evolution mechanisms in massive amounts of historical weather data to predict future weather conditions. However, these methods typically rely on gridded data, such as reanalysis products, for training. Reanalysis data is usually generated by assimilating multi-source observational data, which may introduce uncertainty errors due to issues such as data scale mismatch or approximate algorithms. Especially in areas with sparse observation stations, insufficient observational data exacerbates systematic errors in the assimilation process, making models trained on such data unable to accurately reflect local weather conditions.
[0003] Due to these limitations, research has focused more on using measured data from automatic weather stations to conduct local meteorological spatiotemporal modeling and forecasting studies. Compared to reanalysis data, station-measured data has advantages such as high fidelity, low latency, and low collection costs. However, station-measured meteorological data exhibits sparsity and spatiotemporal heterogeneity, leading to diverse and complex spatiotemporal dependencies, making it difficult for existing models to dynamically adapt to its characteristics. Specifically, spatiotemporal heterogeneity refers to significant differences in the distribution characteristics and evolution patterns of meteorological elements across the spatiotemporal dimensions. Some stations are located in isolated geographical units (such as remote mountainous areas), spatially distant from other stations, and their meteorological models exhibit strong independence. Conversely, some stations form observational networks with upstream and downstream connections, showing significant correlations and co-evolutionary characteristics among their meteorological models. Furthermore, the dynamic and non-stationary nature of weather systems further exacerbates the complexity of spatiotemporal dependencies. For example, under stable weather scenarios, the spatiotemporal correlations between stations are relatively stable, while under volatile scenarios such as extreme weather, these correlations may change rapidly, even exhibiting nonlinear abrupt changes. However, existing models have significant shortcomings in dealing with the aforementioned spatiotemporal heterogeneity. They are unable to dynamically adapt to complex spatiotemporal dependencies, resulting in insufficient modeling of meteorological spatiotemporal models, limited generalization ability in complex weather scenarios, and low forecast accuracy. Summary of the Invention
[0004] To address the problems existing in the traditional methods, this invention proposes a ground station meteorological element forecasting method based on a hybrid expert model and graph attention mechanism. This method can solve the problems of low forecast accuracy and poor generalization ability caused by the inability of existing station element forecasting methods to dynamically adapt to specific weather scenarios.
[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0006] On the one hand, a method for forecasting meteorological elements at ground stations based on a hybrid expert model and graph attention mechanism is provided, including the following steps:
[0007] Step 1: Acquire observation data from ground-based automatic weather stations and preprocess the observation data.
[0008] Step 2: Use a linear mapping layer to map the preprocessing results into high-dimensional features.
[0009] Step 3: Process the high-dimensional features using a temporal information embedding layer to obtain temporally enhanced input features.
[0010] Step 4: Input the time-enhanced input features into the expert network module. The three expert networks with different architectures model the evolution of the time dimension, capture static spatiotemporal dependencies, and capture variable spatial dependencies, respectively.
[0011] Step 5: Input the preprocessing results and the output features of the three expert networks into the scene-aware router, establish the correlation between the preprocessing results and the expert network outputs based on the meta-node library, evaluate the adaptation of the three expert networks, and obtain the matching score of each expert for the current time step scene; normalize the matching scores through Softmax to obtain the route probability of selecting each expert.
[0012] Step 6: Use the Top-K hardware routing strategy to select the expert output with the highest routing probability as the final output result.
[0013] One of the above technical solutions has the following advantages and beneficial effects:
[0014] The aforementioned ground-based meteorological element forecasting method based on a hybrid expert model and graph attention mechanism integrates these two mechanisms. Through the collaborative optimization of a multi-expert architecture, scene-aware routing, and a joint loss function, it captures the spatiotemporal relationships between meteorological elements at ground stations, improving the accuracy of station element forecasts in complex scenarios. This method employs an integrated hybrid expert model and graph attention mechanism, composed of a differentiated expert network and a scene-aware router, effectively capturing the spatiotemporal dependencies between meteorological elements at ground stations. Through the collaborative optimization of dynamic routing and accurate prediction, it effectively improves the accuracy of station element forecasts. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a ground station meteorological element forecasting method based on a hybrid expert model and graph attention mechanism in one embodiment.
[0017] Figure 2 This is a structural diagram of a ground station meteorological element forecasting model based on a hybrid expert model and graph attention mechanism in one embodiment;
[0018] Figure 3 Here is a diagram of the timing expert network structure in one embodiment;
[0019] Figure 4 This is a static spatiotemporal graph expert network structure diagram in one embodiment;
[0020] Figure 5 This is a diagram of the dynamic spatiotemporal graph expert network structure in one embodiment;
[0021] Figure 6 This is a diagram of the scene-aware router structure in one embodiment. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0024] It should be noted that, in this document, the reference to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The presentation of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments. The term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, including such combinations.
[0025] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] In one embodiment, such as Figure 1 As shown, a method for forecasting meteorological elements at ground stations based on a hybrid expert model and graph attention mechanism is provided, which may include the following processing steps 1 to 6:
[0027] Step 1: Acquire observation data from ground-based automatic weather stations and preprocess the observation data.
[0028] Specifically, historical time series observation data are collected from multiple ground automatic weather stations distributed in different geographical locations. The observation data includes at least: (1) basic meteorological elements, such as temperature, air pressure, humidity, wind speed, wind direction and precipitation; (2) derived statistical features calculated by sliding window based on the original observation data, such as the average temperature of the past 3 hours and the cumulative precipitation of the past 6 hours; (3) the data record of each station includes the station's unique identifier, latitude and longitude coordinates, altitude and the observation timestamp of each data point.
[0029] The collected raw observation data undergoes preprocessing such as quality control, missing value data is imputed, and then feature standardization is performed.
[0030] Step 2: Use a linear mapping layer to map the preprocessing results into high-dimensional features.
[0031] Step 3: Process the high-dimensional features using a temporal information embedding layer to obtain temporally enhanced input features.
[0032] Specifically, a time-series information embedding layer is constructed, and the Time2Vec or TS2Vec embedding method is used to learn the information in the timestamps to capture the periodicity and linear patterns in meteorological data.
[0033] Step 4: Input the temporal enhanced input features into the expert network module. The three expert networks with different architectures model the evolution of the time dimension, capture static spatiotemporal dependencies, and capture variable spatial dependencies, respectively.
[0034] Specifically, the expert network module consists of three expert networks with differentiated architectures: the Temporal Expert Network (T-ExNet), which focuses on modeling the evolutionary patterns over time; the Static Spatiotemporal Graph Expert Network (S-ExNet), which focuses on capturing static spatiotemporal dependencies; and the Dynamic Spatiotemporal Graph Expert Network (D-ExNet), which adapts to dynamic spatiotemporal dependencies.
[0035] The temporal expert network adopts a hierarchical structure design based on the Transformer encoder, consisting of a temporal encoder (T-Encoder), a feedforward network (FFN), residual connections, and layer normalization. The temporal encoder (T-Encoder) employs a multi-head attention mechanism, utilizing dynamic attention to fully consider the historical information of the entire meteorological data sequence, and redistributing adaptive weights to capture the importance differences at different time steps. Its main components include feature extraction, score calculation, weight normalization, and feature aggregation stages.
[0036] The static spatiotemporal graph expert network S-ExNet has an overall architecture similar to T-ExNet, the difference being the introduction of a static graph encoder layer SG-Encoder. The static adjacency matrix input to the SG-Encoder is constructed by a memory network, and then spatial dependencies are captured using a graph convolutional network (GCN).
[0037] The dynamic spatiotemporal graph expert network D-ExNet shares the same overall architecture as S-ExNet, but differs in its introduction of a static graph encoder layer, SG-Encoder. SG-Encoder computes the dynamic adjacency weights of each node at each time step in real time, capturing variable spatial dependencies. The dynamic adjacency matrix input to the dynamic graph encoder layer, DG-Encoder, is also constructed using a memory network, followed by spatial dependency capture using a graph convolutional network (GCN).
[0038] Step 5: Input the preprocessing results and the output features of the three expert networks into the scene-aware router, establish the correlation between the preprocessing results and the expert network outputs based on the meta-node library, evaluate the adaptation of the three expert networks, and obtain the matching score of each expert for the current time step scene; normalize the matching scores through Softmax to obtain the route probability of selecting each expert.
[0039] Specifically, the scene-aware router is used to establish the correlation between input features and expert outputs based on the meta-node library, evaluate the fit of the three expert networks, and output the matching score of each expert for the current time step scene. Its structural design mainly consists of a memory query mechanism, routing probability calculation, and joint loss function optimization.
[0040] Step 6: Use the Top-K hardware routing strategy to select the expert output with the highest routing probability as the final output result.
[0041] As a preferred approach, a Top-1 hard routing strategy is adopted, selecting the expert output with the highest probability as the final output result.
[0042] A ground-based meteorological element forecasting model based on a hybrid expert model and graph attention mechanism consists of a linear mapping layer, a temporal information embedding layer, an expert network module, a scene-aware router, and a Top-K hardware routing strategy. The structure of the ground-based meteorological element forecasting model based on a hybrid expert model and graph attention mechanism is as follows: Figure 2 As shown.
[0043] The aforementioned ground-based meteorological element forecasting method based on a hybrid expert model and graph attention mechanism integrates these two mechanisms. Through the collaborative optimization of a multi-expert architecture, scene-aware routing, and a joint loss function, it captures the spatiotemporal relationships between meteorological elements at ground stations, improving the accuracy of station element forecasts in complex scenarios. This method employs an integrated hybrid expert model and graph attention mechanism, composed of a differentiated expert network and a scene-aware router, effectively capturing the spatiotemporal dependencies between meteorological elements at ground stations. Through the collaborative optimization of dynamic routing and accurate prediction, it effectively improves the accuracy of station element forecasts.
[0044] In one embodiment, step 1 specifically includes:
[0045] Step 1-1: Collect raw observation data from multiple automatic weather stations distributed in different geographical locations. The raw observation data includes historical time-series observation data, derived statistical features calculated from the raw observation data using a sliding window, and data records for each station. The observation data should contain at least basic meteorological elements. The derived statistical features should include at least the average temperature over the past 3 hours and the cumulative precipitation over the past 6 hours. The data records for each station should include a unique station identifier, latitude and longitude coordinates, altitude, and the observation timestamp for each data point.
[0046] Steps 1-2: Perform quality checks, data cleaning, outlier handling, imputation of missing values, and feature standardization on the raw observation data to obtain the preprocessed results. Specifically, this includes:
[0047] Step 1-2-1: Perform physical rationality checks and eliminations on the original dataset. Based on prior meteorological knowledge, set an absolute reasonable range for each meteorological element. Physically impossible values that exceed this range are directly identified as physical outliers and eliminated.
[0048] Step 1-2-2 involves using a dynamic extreme value test method based on historical statistical information to perform climate extreme value testing on the original dataset. For each station and each meteorological element, the daily climate mean and standard deviation of the long-term historical data are calculated. For any given date, the reasonable dynamic threshold is defined as the climate mean ± K times the standard deviation for that date. Data points exceeding the dynamic threshold are marked as climate outliers and set as missing values, to be corrected by subsequent time-series interpolation algorithms.
[0049] Steps 1-2-3 involve performing time consistency checks and smoothing operations on the original dataset, checking for inconsistencies between related features, calculating the difference between data at adjacent time steps (e.g., 1 hour before and after), setting a mutation threshold based on feature characteristics, marking suspicious data points exceeding the mutation threshold as time-series outliers, setting them as missing values, and correcting them using subsequent time-series interpolation algorithms.
[0050] Steps 1-2-4 involve using a temporal interpolation algorithm to process missing values and using the ARIMA temporal model to fill in the missing values, ensuring that the data is complete and free of missing data, thus meeting the quality requirements of the experiment.
[0051] Steps 1-2-5 involve z-score feature standardization of the processed dataset. Each meteorological element feature channel is standardized separately and converted into a standardized distribution with a mean of 0 and a variance of 1.
[0052] In one embodiment, step 2: using a fully connected layer to map the preprocessing results into high-dimensional features.
[0053] In one embodiment, step 3 includes: using the Time2Vec embedding method to learn information in the timestamps of the high-dimensional features, capturing the periodicity and linear patterns in the meteorological data, and obtaining a time-series embedding vector; concatenating the time-series embedding vector with the high-dimensional features to obtain time-enhanced input features.
[0054] Specifically, the Time2Vec embedding method is used to learn information from timestamps, capturing periodicity and linear patterns in meteorological data. Given the time-series information of a timestamp... ,use h dimensional embedding vector To compute temporal embedding vectors Specifically, through each dimension Learnable parameters and To parameterize it, the formula is as follows:
[0055]
[0056] in This represents the dimension of the temporal embedding vector. This represents a periodic activation function. The sine function is chosen to capture periodic, linear terms. To capture aperiodic features.
[0057] The timing embedding vector is then used. Will be related to input meteorological elements The high-dimensional features mapped by the fully connected layer are concatenated into a temporal-enhanced input feature vector. The data is then input into the subsequent expert network; a sliding window approach is used for a given time step. Select the past Historical observation dataset at each time step Each meteorological element is used as an input meteorological element. .
[0058] In one embodiment, the expert network module includes: a temporal expert network focusing on modeling the evolutionary patterns of the time dimension, a static spatiotemporal graph expert network focusing on capturing static spatiotemporal dependencies, and a dynamic spatiotemporal graph expert network adapted to dynamic spatiotemporal dependencies; step 4 includes: inputting the temporal enhanced input features into the temporal expert network, processing them using a hierarchical network structure based on a Transformer encoder to obtain the output features of the temporal expert network; and inputting the temporal enhanced input features into the static spatiotemporal graph expert network, processing them using a hierarchical network structure based on a Transformer encoder that incorporates a static graph encoder layer. The output features of the static spatiotemporal graph expert network are obtained. The static graph encoder layer is used to capture spatial dependencies using a graph convolutional network (GCN), and the temporal attention output features are weighted by a static adjacency matrix to obtain the neighbor features of each node. The temporal enhancement input features are then input into the dynamic spatiotemporal graph expert network, and processed using a layered network structure based on a Transformer encoder that incorporates a dynamic graph encoder layer to obtain the output features of the dynamic spatiotemporal graph expert network. The dynamic graph encoder layer is used to capture spatial dependencies using a graph convolutional network (GCN), and the temporal attention output features are weighted by a dynamic adjacency matrix to obtain the neighbor features of each node.
[0059] In one embodiment, such as Figure 3 As shown, the temporal expert network includes: a temporal encoder, a feedforward network, residual connections, and layer normalization. The temporal encoder employs a multi-head attention mechanism, utilizing dynamic attention to focus on historical information throughout the entire meteorological data sequence, and then redistributes adaptive weights to capture the importance differences at different time steps. The temporal-enhanced input features are input into the temporal expert network and processed using a hierarchical network structure based on a Transformer encoder to obtain the output features of the temporal expert network, including:
[0060] The temporally enhanced input features are fed into the temporal encoder to obtain the attention score:
[0061]
[0062] in, To score attention, For time-enhanced input features, , , Representing three space-feature interaction matrices, Represents the time transformation matrix. Represents the Sigmoid activation function. This is the bias vector.
[0063] Attention scores are aggregated with temporally enhanced input features to obtain temporal attention features. These temporal attention features are then input into a feedforward network. The resulting output is added to and fused with the temporal attention features, and then subjected to layer normalization to obtain the output features of the temporal expert network.
[0064] Specifically, the temporal expert network consists of a temporal encoder (T-Encoder), a feedforward network (FFN), residual connections, and layer normalization. The T-Encoder employs a multi-head attention mechanism, utilizing dynamic attention to fully consider the historical information of the entire meteorological data sequence, and then redistributing adaptive weights to capture the importance differences at different time steps. Its main stages include feature extraction, score calculation, weight normalization, and feature aggregation. Specifically, the temporal enhanced input vector is first processed... Dimensionality compression is performed on both the spatial and feature dimensions, decomposing the original features into independently processable subspaces to reduce the computational burden on high-dimensional spatiotemporal data. A learnable spatial-feature interaction matrix W is then introduced to reveal the relationship between modeling node features and temporal features. Next, higher-order transformations in the temporal dimension are learned through matrix V, and matrix b provides bias correction, enabling the attention weights to capture temporal dependencies. Finally, the attention score is calculated using the following formula:
[0065]
[0066] in , , Representing three space-feature interaction matrices, Represents the time transformation matrix. This represents the Sigmoid activation function. A normalization mechanism combining Sigmoid pre-activation and Softmax is used to normalize all data. Finally, the attention scores are aggregated with the input to obtain the temporal attention features. for:
[0067]
[0068] Subsequently attention characteristics The input is fed into the feedforward network, and through residual connections and layer normalization, the output of T-ExNet is finally obtained. .
[0069] In one embodiment, such as Figure 4 As shown, the static spatiotemporal graph expert network is an extension of the temporal expert network, which introduces a static graph encoder, residual connections, and layer normalization between the residual connections after the temporal encoder and the feedforward network. In the static graph encoder:
[0070] A memory network based on a meta-node library is constructed to generate a static adjacency matrix. The memory network is used to memorize and store different weather patterns by introducing a meta-node library, and then a static adjacency matrix is generated through a hypernetwork.
[0071] A graph convolutional network is used to capture spatial dependencies in the static adjacency matrix, and the temporal attention features are weighted by the neighbor features of each node in the static adjacency matrix to obtain the output features of the static graph encoder layer.
[0072] Specifically, the overall architecture of the Static Spatiotemporal Graph Expert Network (S-ExNet) is similar to that of T-ExNet, the difference being the introduction of a static graph encoder layer (SG-Encoder). The S-ExNet consists of a temporal encoder (T-Encoder), a feedforward network (FFN), the static graph encoder layer (SG-Encoder), residual connections, and layer normalization. The static adjacency matrix input to the SG-Encoder is constructed using a memory network, and then spatial dependencies are captured using a graph convolutional network (GCN).
[0073] Furthermore, this specifically includes image processing steps:
[0074] (1) Construct a memory network based on the meta-node bank to generate a static graph topology. The memory network uses the meta-node bank to memorize and store different weather patterns, and then generates a static adjacency matrix through a hyper-network. Specifically, the initialization of the meta-node bank vector is as follows: ,in and These represent the number of memory items and the dimension, respectively. Then, a hypernetwork is used to connect the meta-node library. Through linear transformation matrix Mapping yields meta-node embeddings Then, the similarity between nodes is calculated using the inner product and normalized to generate a learnable static adjacency matrix. The calculation formula is as follows:
[0075]
[0076]
[0077] in, This is the embedding vector for the meta node.
[0078] (2) Graph Convolutional Network (GCN) is used to capture spatial dependencies and output features with temporal attention. Use adjacency matrix The neighbor features of each node in the weighted matrix are calculated using the following formula:
[0079]
[0080] in The output features of the static graph encoder layer SG-Encoder are then fed into the feedforward network, and through residual connections and layer normalization, the output of S-ExNet is finally obtained. .
[0081] In one embodiment, such as Figure 5 As shown, the dynamic spatiotemporal graph expert network is an extension of the temporal expert network, which introduces a dynamic graph encoder, residual connections, and layer normalization between the residual connections after the temporal encoder and the feedforward network. In the dynamic graph encoder:
[0082] A memory network based on a meta-node library is constructed. Memory items in the memory network at specific time steps are read through a memory matching mechanism. Then, a hypernetwork is used to generate a dynamic adjacency matrix. A graph convolutional network is used to capture spatial dependencies. The temporal attention features are weighted by the dynamic adjacency matrix, and the neighbor features of each node are weighted to obtain the output features of the dynamic graph encoder layer.
[0083] Specifically, the dynamic spatiotemporal graph expert network D-ExNet has the same overall network architecture as S-ExNet, the difference being the introduction of a static graph encoder layer SG-Encoder. SG-Encoder can compute the dynamic adjacency weights of each node at each time step in real time, capturing variable spatial dependencies. The D-ExNet layered structure consists of a temporal encoder T-Encoder, a feedforward network FFN, a dynamic graph encoder layer DG-Encoder, residual connections, and layer normalization. The dynamic adjacency matrix input to the dynamic graph encoder layer DG-Encoder is also constructed using a memory network, and then a graph convolutional network GCN is used to capture spatial dependencies.
[0084] Specifically, the following steps are included:
[0085] (1) A dynamic graph topology is generated using a memory network built on a meta-node bank. Memory entries in the memory network at specific time steps are read using a memory match mechanism (3M), and then a dynamic adjacency matrix is generated using a hypernetwork. Specifically, first... t Temporal attention characteristics at any given moment Projection as query vector Subsequently, a memory matching mechanism is used to perform a memory retrieval operation from the meta-node library to calculate the query vector. With each memory item Similarity score The memory items are then weighted and summed based on their scores to obtain the metanode vector. The calculation formula is as follows:
[0086]
[0087]
[0088]
[0089] in express The Middle i Node vectors This is the projection transformation matrix; Represents a local query vector; As a scalar, reflecting and The similarity between them; Indicates the first i The meta-node vector of each node is constantly updated.
[0090] Then, the updated metanode vector Input into the hypernetwork and generate dynamic meta-node embeddings. Then, the similarity between nodes is calculated using the inner product and normalized to generate a learnable dynamic adjacency matrix. The calculation formula is as follows:
[0091]
[0092]
[0093] (2) Graph Convolutional Network (GCN) is used to capture spatial dependencies and output features with temporal attention. Use adjacency matrix The neighbor features of each node in the weighted matrix are calculated using the following formula:
[0094]
[0095] in The output features of the DG-Encoder are then fed into a feedforward network, and through residual connections and layer normalization, the output of D-ExNet is obtained. .
[0096] In one embodiment, step 5 includes: projecting the preprocessing result into a query vector through a linear layer; calculating the similarity between the query vector and each memory item in the meta-node library using a memory matching mechanism; obtaining attention weights through softmax normalization, and then using the attention weights to perform a weighted summation of the memory items to generate a memory query result; calculating the cosine similarity between the output features of each expert network and the memory query result to obtain an adaptation score; and converting the adaptation score into the routing probability of selecting each expert through softmax normalization.
[0097] Specifically, such as Figure 6 As shown, the Context-Aware Router establishes the correlation between input features and expert outputs based on the meta-node library, evaluates the fit of three expert networks, and inputs the prediction result of the optimal expert. Its structural design mainly consists of a memory query mechanism, route probability calculation, and joint loss function optimization.
[0098] Specifically, the steps include the following:
[0099] (1) The memory query mechanism is designed to establish a connection between input features and expert outputs through a meta-node library, providing a basis for routing decisions. Specifically, the input meteorological data at the current time step is used to... Projecting the query vector through a linear layer:
[0100]
[0101] Then, the similarity between the query vector and each memory item in the memory bank is calculated using a memory matching mechanism, and the attention weights are obtained through softmax normalization. Then, a weighted sum is used to generate the memory query result. The calculation formula is as follows:
[0102]
[0103]
[0104] in This reflects the degree of matching between the current input data and typical weather patterns, providing a basis for subsequent routing.
[0105] (2) Calculate the routing probability and calculate the output features of each expert network. With memory query results The similarity, which quantifies the degree to which experts match the current time step scenario, is calculated using the following formula:
[0106]
[0107] in This represents the similarity function; cosine similarity is chosen. The matching score will then be calculated. By normalizing using softmax, it is transformed into the route probability of selecting each expert:
[0108]
[0109] in E Given the number of expert networks, a "Top-1" hard routing strategy is adopted, selecting the expert input with the highest probability as the final input result:
[0110]
[0111] in Representative experts e The prediction results.
[0112] In one embodiment, a ground station meteorological element forecasting model based on a hybrid expert model and graph attention mechanism is constructed, consisting of a linear mapping layer, a temporal information embedding layer, an expert network module, a scene-aware router, and a Top-K hardware routing strategy. The joint loss function expression used in the training process of the ground station meteorological element forecasting model is as follows:
[0113]
[0114] in, For the joint loss function, To predict the MAE loss for the final model, To minimize losses, The optimal choice loss, worst-case avoidance loss, and optimal choice loss are all of the same type as cross-entropy loss. Weights for the worst-case avoidance loss and the best-case selection loss.
[0115] Specifically, the "routing problem of regression task" is transformed into a "label matching problem of classification task." "Correct routes" and "incorrect routes" are defined using "pseudo-labels," and cross-entropy loss is used to penalize bias, guiding the router to learn the optimal mapping relationship between the "scene" and the expert. Joint loss function. This includes the individual loss (MAE) ultimately predicted by the model. And the worst-case avoidance loss. And Best Choice Loss The weighted sum of the terms, the joint loss function is shown in the joint loss function expression above.
[0116] Specifically, the following steps are included:
[0117] To design the worst-case loss avoidance function, first quantize the error and set the error quantiles. Labels are generated, and then cross-entropy loss is calculated. Specifically, MAE is first used to quantify the prediction error of each expert. The calculation formula is:
[0118]
[0119] in Indicates site i exist t Real-time meteorological elements; Experts e The predicted value; then, an error quantile is set. By judging the error Has it exceeded the limit? Did the decision regarding the router involve experts? e Generate pseudo tags The rules are as follows:
[0120]
[0121] Then, cross-entropy is used to measure the routing probability. With pseudo-tags The difference is used to penalize incorrect decisions, as shown in the following formula:
[0122]
[0123] in, This represents the routing probability.
[0124] It should be understood that, although the above Figure 1 The steps are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated in this document, there is no strict order in which these steps are executed; they can be performed in other orders. Furthermore, the above... Figure 1 At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0125] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0126] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and all such modifications and improvements fall within the scope of protection of this application.
Claims
1. A method for forecasting meteorological elements at ground stations based on a hybrid expert model and graph attention mechanism, characterized in that, Including the following steps: Step 1: Acquire observation data from ground-based automatic weather stations and preprocess the observation data; Step 2: Use a linear mapping layer to map the preprocessing results into high-dimensional features; Step 3: Process the high-dimensional features using a temporal information embedding layer to obtain temporally enhanced input features; Step 4: The temporal-enhanced input features are input into the expert network module. Three expert networks with differentiated architectures model the temporal evolution patterns, capture static spatiotemporal dependencies, and capture variable spatial dependencies, respectively. The expert network module includes: a temporal expert network focusing on modeling the temporal evolution patterns, a static spatiotemporal graph expert network focusing on capturing static spatiotemporal dependencies, and a dynamic spatiotemporal graph expert network adapted to dynamic spatiotemporal dependencies. Step 4 specifically includes: The temporal enhanced input features are input into the temporal expert network and processed using a hierarchical network structure based on a Transformer encoder to obtain the output features of the temporal expert network. The temporal enhanced input features are input into the static spatiotemporal graph expert network and processed using a layered network structure based on a Transformer encoder that incorporates a static graph encoder layer to obtain the output features of the static spatiotemporal graph expert network. The static graph encoder layer is used to capture spatial dependencies using a graph convolutional network (GCN) and to weight the temporal attention output features with the neighbor features of each node using a static adjacency matrix. The temporal enhanced input features are input into the dynamic spatiotemporal graph expert network, and processed using a layered network structure based on a Transformer encoder that introduces a dynamic graph encoder layer to obtain the output features of the dynamic spatiotemporal graph expert network. The dynamic graph encoder layer is used to capture spatial dependencies using a graph convolutional network (GCN), and the temporal attention output features are weighted by a dynamic adjacency matrix with the neighbor features of each node. Step 5: Input the preprocessing results and the output features of the three expert networks into the scene-aware router, establish the correlation between the preprocessing results and the expert network outputs based on the meta-node library, evaluate the adaptation of the three expert networks, and obtain the matching score of each expert for the current time step scene; normalize the matching scores through Softmax to obtain the route probability of selecting each expert. Step 6: Use the Top-K hardware routing strategy to select the expert output with the highest routing probability as the final output result.
2. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 1, characterized in that, Step 1 includes: Raw observation data from multiple automatic weather stations distributed across different geographical locations were collected. The raw observation data included historical time-series observation data, derived statistical features calculated from the raw observation data using a sliding window, and data records for each station. The observation data contained at least basic meteorological elements. The derived statistical features included at least the average temperature over the past 3 hours and the cumulative precipitation over the past 6 hours. The data records for each station included a unique station identifier, latitude and longitude coordinates, altitude, and an observation timestamp for each data point. The original observation data is subjected to quality verification, data cleaning, outlier handling, imputation of missing values, and feature standardization to obtain the preprocessed results.
3. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 1, characterized in that, Step 2: Use a fully connected layer to map the preprocessing results into high-dimensional features.
4. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 1, characterized in that, Step 3 includes: The Time2Vec embedding method is used to learn the information in the timestamps of the high-dimensional features to capture the periodicity and linear patterns in the meteorological data and obtain the time-series embedding vector. The temporal embedding vector is concatenated with the high-dimensional feature to obtain the temporal-enhanced input feature.
5. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 4, characterized in that, The time series expert network includes: a time series encoder, a feedforward network, residual connections, and layer normalization; the time series encoder is used to employ a multi-head attention mechanism, utilize dynamic attention to focus on the historical information of the entire meteorological data sequence, and redistribute adaptive weights to capture the importance differences at different time steps. The temporally enhanced input features are input into the temporal expert network and processed using a hierarchical network structure based on a Transformer encoder to obtain the output features of the temporal expert network, including: The temporally enhanced input features are fed into the temporal encoder to obtain the attention score: in, To score attention, For time-enhanced input features, , , Represents the space-feature interaction matrix. Represents the time transformation matrix. Represents the Sigmoid activation function. It is the bias vector; The attention score is aggregated with the temporal enhanced input features to obtain temporal attention features; The temporal attention features are input into the feedforward network, and the resulting output is added to and fused with the temporal attention features. Then, layer normalization is performed to obtain the output features of the temporal expert network.
6. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 5, characterized in that, The static spatiotemporal graph expert network is formed by introducing a static graph encoder, residual connections, and layer normalization between the residual connections after the time sequence encoder and the feedforward network in the time sequence expert network. In the static graph encoder: A memory network based on a meta-node library is constructed to generate a static adjacency matrix. The memory network is used to memorize and store different weather patterns by introducing a meta-node library, and then a static adjacency matrix is generated through a hypernetwork. The static adjacency matrix is used to capture spatial dependencies using a graph convolutional network. The temporal attention features are weighted by the static adjacency matrix to obtain the neighbor features of each node, thus obtaining the output features of the static graph encoder layer.
7. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 5, characterized in that, The dynamic spatiotemporal graph expert network is formed by introducing a dynamic graph encoder, residual connection, and layer normalization into the residual connection after the time sequence encoder and the feedforward network in the time sequence expert network. In the dynamic graph encoder: A memory network based on a meta-node library is constructed. Memory items in the memory network at specific time steps are read through a memory matching mechanism, and then a dynamic adjacency matrix is generated using a hypernetwork. A graph convolutional network is used to capture spatial dependencies. The temporal attention features are weighted by a dynamic adjacency matrix, which is used to calculate the neighbor features of each node, thus obtaining the output features of the dynamic graph encoder layer.
8. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 1, characterized in that, Step 5 includes: The preprocessed results are projected into a query vector through a linear layer; The similarity between the query vector and each memory item in the meta-node library is calculated using a memory matching mechanism. Attention weights are obtained by normalizing with softmax, and then the attention weights are used to perform a weighted summation of the memory items to generate memory query results. Calculate the cosine similarity between the output features of each expert network and the memory query results to obtain the fit score; The adaptation score is normalized using softmax to convert it into the route probability of selecting each expert.
9. The ground station meteorological element forecasting method based on hybrid expert model and graph attention mechanism according to claim 1, characterized in that, A ground-based meteorological element forecasting model based on a hybrid expert model and graph attention mechanism is constructed, consisting of a linear mapping layer, a temporal information embedding layer, an expert network module, a scene-aware router, and a Top-K hardware routing strategy. The joint loss function used in the training process of the ground station meteorological element forecast model is: in, For the joint loss function, To predict the MAE loss for the final model, To minimize losses, The optimal choice loss, worst-case avoidance loss, and optimal choice loss are all of the same type as cross-entropy loss. Weights for the worst-case avoidance loss and the best-case selection loss.
Citation Information
Patent Citations
Time sequence prediction method based on multi-source heterogeneous data
CN117743884A
PM10 concentration prediction method based on space-time diagram neural network and expert hybrid model
CN120492896A