PM10 concentration prediction method based on space-time diagram neural network and expert hybrid model

Through the spatial and temporal graph neural network and expert hybrid model, combined with multi-scale time feature extraction, residual attention and dynamic multimodal weighted graph convolution, the dynamic periodicity and non-stationarity problems of spatial and temporal features of PM10 concentration prediction are solved, and high-precision PM10 concentration prediction is achieved.

CN120492896APending Publication Date: 2025-08-15INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510587734.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the long-term dependence and instantaneous correlation of the space-time dimension of PM10 concentration, and traditional methods are difficult to capture the dynamic periodicity and non-stationarity of its space-time characteristics at the same time, resulting in inaccurate prediction of PM10 concentration.

Method used

Using a method based on a spatiotemporal graph neural network and an expert hybrid model, a complex PM10 propagation mode is dynamically modeled through multi-scale time feature extraction, residual attention mechanism, dynamic multimodal weighted graph and graph convolutional neural network, combined with an expert hybrid network, and PM10 concentration prediction is performed to dynamically model complex PM10 propagation mode.

Benefits of technology

It realizes high-precision PM10 concentration prediction in the next 24 hours, improves the robustness and generalization ability of the prediction, and can adapt to changes in different pollution modes and meteorological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492896A_ABST
    Figure CN120492896A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of PM10 concentration prediction, and discloses a PM10 concentration prediction method based on a space-time diagram neural network and an expert hybrid model, and the method comprises the following specific steps: S1, time feature extraction (RTAF): the PM10 concentration is influenced by a plurality of time factors, including short-term fluctuation, medium-term trend and long-term trend; a dynamic multi-modal weighted graph is constructed, meteorological factors, geographic positions and historical pollution similarities are coded into features of edges and nodes, a PM10 spatial propagation mechanism is modeled based on an adaptive graph neural network, a residual attention fusion module is introduced into the model in the time dimension, multi-scale time dependence features are effectively extracted, and the time-dependent features are extracted. According to the method, a long-term trend and a short-time fluctuation process are captured, finally, dynamic modeling and expert selection are performed on a complex PM10 propagation mode by using an expert hybrid network, the prediction robustness and generalization ability are improved, the model fully fuses a PM transmission mechanism and a depth space-time modeling ability, and high-precision prediction of the PM10 concentration in the next 24 hours is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of PM10 concentration prediction, and specifically provides a PM10 concentration prediction method based on a spatiotemporal graph neural network and an expert mixture model. Background Art

[0002] PM10 concentration refers to the concentration of particles with an aerodynamic equivalent diameter of less than or equal to 10 microns in the air. PM10 persists in the ambient air for a long time and can affect people's health and air visibility. Long-term exposure to high PM10 concentrations will not only affect urban traffic, but also harm people's cardiovascular and respiratory systems. In recent years, air governance has achieved great success, but frequent sandstorms still have an impact on people's lives. Studies have shown that the spatiotemporal evolution characteristics of sandstorms are directly related to the level of PM10 concentration.

[0003] The concentration of inhalable particulate matter in the air increases significantly during sandstorms, which will reduce surface solar radiation and cause dust accumulation on photovoltaic panels, reducing the photoelectric conversion efficiency. For wind turbine operation, key components are easily affected by particles in extreme environments during sandstorms, resulting in damage to the aerodynamic performance and structural safety of the wind turbine, which has a serious impact on the centralized development of new energy power generation bases and the safe and stable long-distance transmission. Therefore, accurately predicting PM10 concentration values in advance is of great significance to the sustainable development of society. First, the formation and evolution of PM10 involves a complex process: from pollution source emissions to transmission and diffusion affected by meteorological conditions and geographic information, its spatiotemporal distribution shows highly nonlinear and multi-scale coupling characteristics. In addition, the spatiotemporal characteristics of PM10 often have dynamic periodicity and non-stationarity. Traditional methods find it difficult to simultaneously capture the long-term dependence and instantaneous correlation of its spatiotemporal dimensions, so they need to be improved. Summary of the Invention

[0004] The purpose of the present invention is to provide a PM10 concentration prediction method based on a spatiotemporal graph neural network and an expert mixture model to solve the problems raised in the above background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a PM10 concentration prediction method based on a spatiotemporal graph neural network and an expert mixture model, the specific steps of which are as follows:

[0006] S1. Temporal feature extraction (RTAF)

[0007] PM10 concentration is affected by multiple temporal factors, including short-term fluctuations, medium-term trends, and long-term trends. To effectively model these temporal characteristics, we use multi-scale temporal feature extraction to obtain information at different time scales, and utilize a residual attention mechanism for feature enhancement. Finally, adaptive feature fusion is used to select the most important time steps for dynamic weighting.

[0008] S1.1 Basic temporal feature modeling

[0009] Given the PM10 concentrations and related meteorological data of multiple stations, the input sequence is defined as:

[0010]

[0011] in, Represents the feature vector of the i-th station at time t, including information such as PM10 concentration, temperature, humidity, and wind speed;

[0012] Since the variation pattern of PM10 concentration is complex, involving a combination of short-term fluctuations and long-term trends, a single-scale feature extraction method is difficult to meet the prediction requirements. Therefore, a multi-layer temporal feature extraction network is first constructed. This network can capture global temporal dependencies and provide robust temporal feature representation. The time series of multi-scale input is represented as:

[0013] H t =f(X t )

[0014] Where, X t is the input time series data, H t is the extracted temporal feature, f(·) is the mapping function of the feature extraction network;

[0015] In order to enhance the interactivity between time steps and ensure that information at different time scales can be effectively transmitted, a residual attention mechanism is introduced to replace the traditional sequence modeling method.

[0016] S1.2 Residual Attention Mechanism

[0017] By introducing the residual attention mechanism, information between different time steps can interact more effectively;

[0018] The residual attention mechanism calculation formula is as follows:

[0019]

[0020] Among them, Q, K, V are query, key and value matrices respectively, d k is the dimension of the key, PrevLayerOutput represents the output of the previous layer, and the softmax normalization operation is used to assign attention weights at different time steps;

[0021] The residual attention mechanism not only enhances the interactivity between time steps, but also avoids the gradient vanishing problem in deep networks through residual connections;

[0022] The final time features are calculated as follows:

[0023] H residual =Residual Attention(H t )

[0024] S1.3 Adaptive feature fusion

[0025] By introducing the adaptive feature fusion mechanism, the model can automatically learn which time steps are most important for the final prediction. The calculation process of adaptive feature fusion is as follows:

[0026] α t =softmax(W x x t +W h h t-1 )

[0027]

[0028] Among them, α t is the attention weight at time step t, W x ,W h is a trainable parameter, It is the temporal feature after residual attention enhancement, and the softmax operation ensures that the sum of the weights of all time steps is 1;

[0029] Through the adaptive feature fusion mechanism, the most critical time steps for PM10 prediction can be dynamically screened. The time step selection strategy is continuously optimized during the model training process, allowing the model to adapt to different time feature patterns. The final output time feature calculation is:

[0030]

[0031] Time characteristic H exttime It will be used as input to the spatial feature extraction module for subsequent site relationship modeling;

[0032] S2. Spatial feature extraction (STGNN)

[0033] S2.1. Spatial Dependence Modeling

[0034] PM10 concentrations are affected not only by temporal variations but also by the geographic location of different regions, exhibiting significant spatial dependence. The diffusion of PM values between adjacent monitoring stations is influenced by factors such as wind speed, wind direction, topography, and emission sources, which in turn dynamically change over time. These factors jointly determine the pollution similarity and the intensity of mutual influence between monitoring stations. By introducing a dynamic multimodal weighted graph (DMWG), we can more comprehensively model the spatial dependence of PM10 concentrations.

[0035] To do this, consider the following three key spatial dependencies:

[0036] Geographical proximity: The physical distance between monitoring stations determines the spatial diffusion capacity of pollutants. Adjacent monitoring stations usually have strong pollution correlations.

[0037] Meteorological factors such as wind speed and direction affect the transmission path of pollutants, making it possible for strong pollution relationships to exist between some distant measuring stations.

[0038] Historical similarity: PM10 concentration variation patterns at different stations may be similar, even if they are geographically distant. However, strong spatial dependencies may still exist due to similar pollution sources or seasonal variations. DMWG will use a weighted fusion strategy to combine these three spatial dependencies, allowing the model to not only capture static geographic proximity information but also adaptively adjust for the contributions of meteorological influences and historical pollution patterns.

[0039] S2.2 Graph Construction Method

[0040] Dynamic adjacency matrix generation: A dynamic multimodal weighted graph (DMWG) is constructed, where nodes represent monitoring stations and edge weights are calculated based on a combination of factors. The core of the graph construction lies in defining the adjacency matrix A, which determines how information is propagated between stations. Incorporating multimodal factors, the adjacency matrix is generated using the following strategy:

[0041] A=αA1+βA2+γA3

[0042] Among them, α, β, and γ are trainable parameters, ensuring that the sum of their weights is 1, so that the model can adapt to the pollution propagation mode in different situations; to avoid numerical explosion, the adjacency matrix needs to be normalized. Where D is the degree matrix of the adjacency matrix A, which is used to ensure that the information of different stations will not be overly concentrated or dispersed due to the imbalance of the adjacency matrix;

[0043] Adaptive adjacency matrix update mechanism: By introducing an adaptive adjacency matrix update mechanism, the adjacency matrix can be adjusted according to the current pollution propagation pattern:

[0044] A (t) =ηA(t-1) +(1-η)A new

[0045] Where A (t) is the adjacency matrix at time step t, A (t-1) is the adjacency matrix of the previous time step, A new is the currently calculated adjacency matrix, and η is used to control the smoothness of the adjacency matrix update, which is used to ensure that the model can adapt to the dynamic changes of PM10 propagation over time and improve the accuracy of spatiotemporal prediction;

[0046] Graph structure construction: Based on the calculated dynamic adjacency matrix, the PM10 prediction problem is represented as a spatiotemporal graph G = (V, E, A). The node set V represents air quality monitoring stations, each serving as a node in the graph. The edge set E represents the spatial connectivity between monitoring stations, determined by the adjacency matrix A. The adjacency matrix A is used to describe the PM10 propagation pattern between monitoring stations.

[0047] In the spatiotemporal diagram, the PM10 concentration data of the measuring station changes dynamically over time, forming a spatiotemporal data sequence, namely:

[0048] X={X1,X2,…,X T}

[0049] Where, X t represents the PM10 data of all monitoring stations at time step t; finally, STGNN is applied on the constructed graph structure to learn spatiotemporal features to obtain the pollution interaction pattern between monitoring stations;

[0050] S2.3. Spatial feature extraction and calculation

[0051] After constructing the spatiotemporal graph structure and defining the dynamic adjacency matrix, the model needs to extract effective spatial features based on this to learn the pollution propagation pattern between monitoring stations. By using the spatiotemporal graph neural network (STGNN) for calculation, PM10 information between monitoring stations can be efficiently transmitted and used for the final PM10 prediction.

[0052] S2.3.1. Graph Convolutional Neural Network (GCN) Computation

[0053] By utilizing the graph structure information, each station node can not only aggregate information based on its own PM10 concentration, but also aggregate information from neighboring stations. The GCN calculation formula is as follows:

[0054]

[0055] Where, is the normalized adjacency matrix, which is used to ensure the stability of feature propagation; H (l) is the node feature matrix of the lth layer; W(l) is a trainable parameter matrix; σ is a nonlinear activation function;

[0056] The graph convolution operation ensures that the information of PM10 concentration can be diffused along the propagation path and can effectively learn the mutual influence between different monitoring sites;

[0057] S2.3.2 Optimization strategy for spatial feature extraction

[0058] GCN may encounter the problem of information over-smoothing when learning spatial features. That is, as the number of network layers increases, the features of different monitoring sites will gradually converge, resulting in a decrease in prediction ability. To solve this problem, the following optimization strategy is introduced:

[0059] Skip Connection: Add residual connections between GCN layers to preserve information in the deep network while preventing gradient vanishing. The calculation method is as follows, which is used to ensure that the original feature information is not completely submerged in multi-layer propagation.

[0060]

[0061] Graph Attention Network (GAT): GCN uses a fixed adjacency matrix, while GAT allows the model to dynamically adjust weights based on the feature similarity between stations, improving the model's expressiveness. Its formula is as follows:

[0062]

[0063] The final updated version is as follows:

[0064]

[0065] In this way, monitoring stations can dynamically pay attention to the neighboring stations that have the greatest impact on them, improving the ability to model the spread of PM values;

[0066] Multi-Scale Graph Convolution (Multi-Scale GCN): PM10 diffusion may involve different scales in space. Therefore, by using multi-scale graph convolution, the model can learn both local and global propagation features. The calculation method is as follows:

[0067]

[0068] Where H (s) Represents graph convolution features of different scales; W s Trainable parameters of the same scale;

[0069] S2.3.3 Final output of spatial features

[0070] After applying the above GCN calculation and optimization strategy, the final spatial feature representation is obtained:

[0071] H s =STGNN(H time ,A)

[0072] Among them: H time is the output of the temporal feature extraction module; A is the dynamic adjacency matrix constructed by DMWG; the final spatial feature H (s) It will be input into MoE together with the time features for the final expert hybrid decision;

[0073] S3, Mixture of Experts (MOE)

[0074] By adopting a Mixture of Experts (MoE) network, multiple expert sub-models are used to process different propagation modes. The gating network dynamically selects the most appropriate expert and generates the final prediction result through weighted fusion of experts, thereby improving the accuracy of prediction.

[0075] S3.1. Overview of MoE Structure

[0076] The MoE consists of three parts: an expert network (experts), which consists of multiple independent sub-models responsible for learning the characteristics of different pollution patterns. By integrating multiple experts, the model's prediction accuracy and generalization ability are improved, making it suitable for PM10 concentration prediction under different pollution patterns; a gating network, which dynamically adjusts the weights of each expert based on input features; and a weighted expert fusion, which performs a weighted summation of the outputs of multiple experts to form the final prediction. The calculation process of the entire MoE structure is as follows:

[0077]

[0078] Among them, K is the number of experts, G k (H s ) is the weight of the kth expert calculated by the gating network; E k (H s ) is the prediction result of expert k, and S is the set of K selected experts. This means that among all the experts, the gating network selects the most important K experts for the calculation of the current sample, rather than having all experts participate in the calculation at the same time;

[0079] S3.2 Expert Networks

[0080] Each expert network E kResponsible for modeling different PM10 change patterns, the calculation method is as follows:

[0081] E k (H s )=σ(W k H s +b k )

[0082] Where H s is the spatiotemporal network extracted by STGNN, W k , b k is the trainable parameter of expert k; σ is the nonlinear activation function;

[0083] S3.3 Gating Network

[0084] The gating network is used to dynamically select the most suitable expert. It calculates the weight of each expert through Softmax:

[0085]

[0086] Among them, W g , b g is a trainable parameter of the gating network; τ is a temperature parameter used to control the uniformity of expert selection;

[0087] S3.4. Final output of MOE prediction

[0088] MoE calculates PM10 predictions through weighted fusion of experts. Uncertainty modeling is introduced to improve MoE's effectiveness in PM10 prediction: Bayesian MoE models the uncertainty output by the expert network, ensuring that the prediction results include not only concentration estimates but also confidence intervals; Uncertainty Quantification, combined with Monte Carlo Dropout or Deep Ensemble, calculates the uncertainty of the prediction, thereby improving the credibility of the model.

[0089] Preferably, the adjacency matrix based on geographic distance described in S2.1 assumes that the pollution diffusion capacity between monitoring stations is inversely proportional to their geographic distance, and defines the spatial relationship between monitoring station i and monitoring station j as:

[0090]

[0091] where d ijis the Euclidean distance between sites i and j; σ is a hyperparameter that controls the distance decay rate; it ensures that the connection strength between adjacent monitoring stations is high, while the influence of distant monitoring stations is small.

[0092] Preferably, the adjacency matrix based on meteorological factors described in S2.1 needs to be further constructed based on meteorological variables because the diffusion of PM10 is greatly affected by meteorological factors, such as wind speed and wind direction:

[0093]

[0094] Among them, w i and w j represent the meteorological factors at monitoring site i and monitoring site j respectively, and λ is a smoothing factor used to control the influence of meteorological variables.

[0095] Preferably, the adjacency matrix formula based on historical similarity described in S2.1 is:

[0096]

[0097] in, and denote the PM10 concentrations at monitoring sites i and j at time step t, respectively, and δ controls the effect of PM concentration similarity.

[0098] Preferably, the diversity of the expert network is improved by introducing the following mechanisms as described in S3.2: Expert Regularization, which constrains the weights of experts through KL divergence or L2 regularization so that they learn different feature patterns and prevent all experts from converging; Expert Specialization, which samples the expert input data so that each expert learns a specific PM10 change pattern, such as short-term bursts, long-term trend changes, etc.

[0099] Preferably, in the MOE structure described in S3.2, different experts can adopt different structures, and each architecture is suitable for different change pattern modeling requirements, for example, MLP experts are used to learn basic pollution patterns; LSTM experts are used to model long-term dependencies of pollution time series; Transformer experts are used to model complex spatiotemporal features.

[0100] Preferably, the stability and diversity of the gating network are improved by introducing the following strategies as described in S3.3: Load Balancing Regularization, which prevents some experts from being overloaded while other experts are hardly activated; Temperature Annealing, which gradually reduces and enhances competition between experts, so that the model explores different experts in the early training stage and converges to the optimal expert combination in the later stage; Gating Gradient Clipping, which prevents drastic fluctuations in gating weights during gradient updates and improves the stability of the model; Stochastic Gating: adding a small amount of noise to the Softmax calculation of the gating network to make expert selection more robust.

[0101] Preferably, during the construction of the spatial relationship diagram described in S2.2, the system will dynamically adjust the correlation strength between different monitoring sites based on real-time meteorological data and geographical location relationships. If the weather in a certain area changes suddenly, the information transmission weight between related sites will be automatically enhanced to improve the accuracy of pollution diffusion modeling.

[0102] Preferably, when analyzing data, the time feature extraction module in S1 will automatically adjust the feature extraction strategy according to different time periods, and reversely optimize the time analysis range through actual prediction errors during training, so that the model is more in line with the time change law of the actual scene.

[0103] The beneficial effects of the present invention are as follows:

[0104] By constructing a dynamic multimodal weighted graph, meteorological factors, geographical location and historical pollution similarity are encoded as edge and node features, and the spatial propagation mechanism of PM10 is modeled based on an adaptive graph neural network. In the time dimension, the model introduces a residual attention fusion module to effectively extract multi-scale time-dependent features and capture long-term trends and short-term fluctuation processes. Finally, an expert mixture network is used to dynamically model complex PM10 propagation patterns and select experts to improve the robustness and generalization ability of the prediction. The model fully integrates the PM transmission mechanism and deep spatiotemporal modeling capabilities to achieve high-precision prediction of PM10 concentrations in the next 24 hours. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 Schematic diagram of the model of the present invention. DETAILED DESCRIPTION

[0106] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0107] like Figure 1 As shown, the embodiment of the present invention provides a PM10 concentration prediction method based on a spatiotemporal graph neural network and an expert mixture model, and the specific steps are as follows:

[0108] S1. Temporal feature extraction (RTAF)

[0109] PM10 concentration is affected by multiple temporal factors, including short-term fluctuations, medium-term trends, and long-term trends. To effectively model these temporal characteristics, we use multi-scale temporal feature extraction to obtain information at different time scales, and utilize a residual attention mechanism for feature enhancement. Finally, adaptive feature fusion is used to select the most important time steps for dynamic weighting.

[0110] S1.1 Basic temporal feature modeling

[0111] Given the PM10 concentrations and related meteorological data of multiple stations, the input sequence is defined as:

[0112]

[0113] in, Represents the feature vector of the i-th station at time t, including information such as PM10 concentration, temperature, humidity, and wind speed;

[0114] Since the variation pattern of PM10 concentration is complex, involving a combination of short-term fluctuations and long-term trends, a single-scale feature extraction method is difficult to meet the prediction requirements. Therefore, a multi-layer temporal feature extraction network is first constructed. This network can capture global temporal dependencies and provide robust temporal feature representation. The time series of multi-scale input is represented as:

[0115] H t =f(X t )

[0116] Where, X t is the input time series data, H t is the extracted temporal feature, f(·) is the mapping function of the feature extraction network;

[0117] In order to enhance the interactivity between time steps and ensure that information at different time scales can be effectively transmitted, a residual attention mechanism is introduced to replace the traditional sequence modeling method.

[0118] S1.2 Residual Attention Mechanism

[0119] By introducing the residual attention mechanism, information between different time steps can interact more effectively;

[0120] The residual attention mechanism calculation formula is as follows:

[0121]

[0122] Among them, Q, K, V are query, key and value matrices respectively, d k is the dimension of the key, PrevLayerOutput represents the output of the previous layer, and the softmax normalization operation is used to assign attention weights at different time steps;

[0123] The residual attention mechanism not only enhances the interactivity between time steps, but also avoids the gradient vanishing problem in deep networks through residual connections;

[0124] The final time features are calculated as follows:

[0125] H residual =ResidualAttention(H t )

[0126] S1.3 Adaptive feature fusion

[0127] By introducing the adaptive feature fusion mechanism, the model can automatically learn which time steps are most important for the final prediction. The calculation process of adaptive feature fusion is as follows:

[0128] α t =softmax(W x x t +W h h t-1 )

[0129]

[0130] Among them, α t is the attention weight at time step t, W x ,W h is a trainable parameter, It is the temporal feature after residual attention enhancement, and the softmax operation ensures that the sum of the weights of all time steps is 1;

[0131] Through the adaptive feature fusion mechanism, the most critical time steps for PM10 prediction can be dynamically screened. The time step selection strategy is continuously optimized during the model training process, allowing the model to adapt to different time feature patterns. The final output time feature calculation is:

[0132]

[0133] Time characteristic H exttime It will be used as input to the spatial feature extraction module for subsequent site relationship modeling;

[0134] S2. Spatial feature extraction (STGNN)

[0135] S2.1. Spatial Dependence Modeling

[0136] PM10 concentrations are affected not only by temporal variations but also by the geographic location of different regions, exhibiting significant spatial dependence. The diffusion of PM values between adjacent monitoring stations is influenced by factors such as wind speed, wind direction, topography, and emission sources, which in turn dynamically change over time. These factors jointly determine the pollution similarity and the intensity of mutual influence between monitoring stations. By introducing a dynamic multimodal weighted graph (DMWG), we can more comprehensively model the spatial dependence of PM10 concentrations.

[0137] To do this, consider the following three key spatial dependencies:

[0138] Geographical proximity: The physical distance between monitoring stations determines the spatial diffusion capacity of pollutants. Adjacent monitoring stations usually have strong pollution correlations.

[0139] Meteorological factors such as wind speed and direction affect the transmission path of pollutants, making it possible for strong pollution relationships to exist between some distant measuring stations.

[0140] Historical similarity: PM10 concentration variation patterns at different stations may be similar, even if they are geographically distant. However, strong spatial dependencies may still exist due to similar pollution sources or seasonal variations. DMWG will use a weighted fusion strategy to combine these three spatial dependencies, allowing the model to not only capture static geographic proximity information but also adaptively adjust for the contributions of meteorological influences and historical pollution patterns.

[0141] S2.2 Graph Construction Method

[0142] Dynamic adjacency matrix generation: A dynamic multimodal weighted graph (DMWG) is constructed, where nodes represent monitoring stations and edge weights are calculated based on a combination of factors. The core of the graph construction lies in defining the adjacency matrix A, which determines how information is propagated between stations. Incorporating multimodal factors, the adjacency matrix is generated using the following strategy:

[0143] A=αA1+βA2+γA3

[0144] Among them, α, β, and γ are trainable parameters, ensuring that the sum of their weights is 1, so that the model can adapt to the pollution propagation mode in different situations; to avoid numerical explosion, the adjacency matrix needs to be normalized. Where D is the degree matrix of the adjacency matrix A, which is used to ensure that the information of different stations will not be overly concentrated or dispersed due to the imbalance of the adjacency matrix;

[0145] Adaptive adjacency matrix update mechanism: By introducing an adaptive adjacency matrix update mechanism, the adjacency matrix can be adjusted according to the current pollution propagation pattern:

[0146] A (t) =ηA (t-1) +(1-η)A new

[0147] Where A (t) is the adjacency matrix at time step t, A (t-1) is the adjacency matrix of the previous time step, A new is the currently calculated adjacency matrix, and η is used to control the smoothness of the adjacency matrix update, which is used to ensure that the model can adapt to the dynamic changes of PM10 propagation over time and improve the accuracy of spatiotemporal prediction;

[0148] Graph structure construction: Based on the calculated dynamic adjacency matrix, the PM10 prediction problem is represented as a spatiotemporal graph G = (V, E, A). The node set V represents air quality monitoring stations, each serving as a node in the graph. The edge set E represents the spatial connectivity between monitoring stations, determined by the adjacency matrix A. The adjacency matrix A is used to describe the PM10 propagation pattern between monitoring stations.

[0149] In the spatiotemporal diagram, the PM10 concentration data of the measuring station changes dynamically over time, forming a spatiotemporal data sequence, namely:

[0150] X={X1,X2,…,X T}

[0151] Where, X t represents the PM10 data of all monitoring stations at time step t; finally, STGNN is applied on the constructed graph structure to learn spatiotemporal features to obtain the pollution interaction pattern between monitoring stations;

[0152] S2.3. Spatial feature extraction and calculation

[0153] After constructing the spatiotemporal graph structure and defining the dynamic adjacency matrix, the model needs to extract effective spatial features based on this to learn the pollution propagation pattern between monitoring stations. By using the spatiotemporal graph neural network (STGNN) for calculation, PM10 information between monitoring stations can be efficiently transmitted and used for the final PM10 prediction.

[0154] S2.3.1. Graph Convolutional Neural Network (GCN) Computation

[0155] By utilizing the graph structure information, each station node can not only aggregate information based on its own PM10 concentration, but also aggregate information from neighboring stations. The GCN calculation formula is as follows:

[0156]

[0157] Where, is the normalized adjacency matrix, which is used to ensure the stability of feature propagation; H (l) is the node feature matrix of the lth layer; W (l) is a trainable parameter matrix; σ is a nonlinear activation function;

[0158] The graph convolution operation ensures that the information of PM10 concentration can be diffused along the propagation path and can effectively learn the mutual influence between different monitoring sites;

[0159] S2.3.2 Optimization strategy for spatial feature extraction

[0160] GCN may encounter the problem of information over-smoothing when learning spatial features. That is, as the number of network layers increases, the features of different monitoring sites will gradually converge, resulting in a decrease in prediction ability. To solve this problem, the following optimization strategy is introduced:

[0161] Skip Connection: Add residual connections between GCN layers to preserve information in the deep network while preventing gradient vanishing. The calculation method is as follows, which is used to ensure that the original feature information is not completely submerged in multi-layer propagation.

[0162]

[0163] Graph Attention Network (GAT): GCN uses a fixed adjacency matrix, while GAT allows the model to dynamically adjust weights based on the feature similarity between stations, improving the model's expressiveness. Its formula is as follows:

[0164]

[0165] The final updated version is as follows:

[0166]

[0167] In this way, monitoring stations can dynamically pay attention to the neighboring stations that have the greatest impact on them, improving the ability to model the spread of PM values;

[0168] Multi-Scale Graph Convolution (Multi-Scale GCN): PM10 diffusion may involve different scales in space. Therefore, by using multi-scale graph convolution, the model can learn both local and global propagation features. The calculation method is as follows:

[0169]

[0170] Where H (s) Represents graph convolution features of different scales; W s Trainable parameters of the same scale;

[0171] S2.3.3 Final output of spatial features

[0172] After applying the above GCN calculation and optimization strategy, the final spatial feature representation is obtained:

[0173] H s =STGNN(H time ,A)

[0174] Among them: H time is the output of the temporal feature extraction module; A is the dynamic adjacency matrix constructed by DMWG; the final spatial feature H (s) It will be input into MoE together with the time features for the final expert hybrid decision;

[0175] S3, Mixture of Experts (MOE)

[0176] By adopting a Mixture of Experts (MoE) network, multiple expert sub-models are used to process different propagation modes. The gating network dynamically selects the most appropriate expert and generates the final prediction result through weighted fusion of experts, thereby improving the accuracy of prediction.

[0177] S3.1. Overview of MoE Structure

[0178] The MoE consists of three parts: an expert network (experts), which consists of multiple independent sub-models responsible for learning the characteristics of different pollution patterns. By integrating multiple experts, the model's prediction accuracy and generalization ability are improved, making it suitable for PM10 concentration prediction under different pollution patterns; a gating network, which dynamically adjusts the weights of each expert based on input features; and a weighted expert fusion, which performs a weighted summation of the outputs of multiple experts to form the final prediction. The calculation process of the entire MoE structure is as follows:

[0179]

[0180] Among them, K is the number of experts, G k (Hs ) is the weight of the kth expert calculated by the gating network; E k (H s ) is the prediction result of expert k, and S is the set of K selected experts. This means that among all the experts, the gating network selects the most important K experts for the calculation of the current sample, rather than having all experts participate in the calculation at the same time;

[0181] S3.2 Expert Networks

[0182] Each expert network E k Responsible for modeling different PM10 change patterns, the calculation method is as follows:

[0183] E k (H s )=σ(W k H s +b k )

[0184] Where H s is the spatiotemporal network extracted by STGNN, W k , b k is the trainable parameter of expert k; σ is the nonlinear activation function;

[0185] S3.3 Gating Network

[0186] The gating network is used to dynamically select the most suitable expert. It calculates the weight of each expert through Softmax:

[0187]

[0188] Among them, W g , b g is a trainable parameter of the gating network; τ is a temperature parameter used to control the uniformity of expert selection;

[0189] S3.4. Final output of MOE prediction

[0190] MoE calculates PM10 predictions through weighted fusion of experts. Uncertainty modeling is introduced to improve MoE's effectiveness in PM10 prediction: Bayesian MoE models the uncertainty output by the expert network, ensuring that the prediction results include not only concentration estimates but also confidence intervals; Uncertainty Quantification, combined with Monte Carlo Dropout or Deep Ensemble, calculates the uncertainty of the prediction, thereby improving the credibility of the model.

[0191] The adjacency matrix based on geographic distance quantifies the impact of physical distance between monitoring sites on pollutant diffusion, prioritizes modeling of pollution correlation between adjacent sites, accurately reflects the natural law of pollutant attenuation with distance, effectively captures the dominant mode of pollution transmission in local areas, significantly improves the physical rationality of spatial dependence modeling, and enhances the model's adaptability to static geographic structures, thereby providing a more realistic spatial dynamic representation basis for PM10 concentration prediction in complex environments.

[0192] The adjacency matrix based on geographic distance in S2.1 assumes that the pollution diffusion capacity between monitoring stations is inversely proportional to their geographic distance, and defines the spatial relationship between monitoring stations i and j as:

[0193]

[0194] where d ij is the Euclidean distance between sites i and j; σ is a hyperparameter that controls the distance decay rate; it ensures that the connection strength between adjacent monitoring stations is high, while the influence of distant monitoring stations is small.

[0195] Constructing an adjacency matrix based on geographic distance can accurately reflect the natural law that pollutant diffusion decays with physical distance, so that the pollution correlation between adjacent monitoring stations is modeled first, improving the physical rationality of spatial dependency modeling and thus enhancing the model's adaptability to static geographic structures.

[0196] Among them, the adjacency matrix based on meteorological factors in S2.1 needs to be further constructed based on meteorological variables because the diffusion of PM10 is greatly affected by meteorological factors such as wind speed and wind direction:

[0197]

[0198] Among them, w i and w j represent the meteorological factors at monitoring site i and monitoring site j respectively, and λ is a smoothing factor used to control the influence of meteorological variables.

[0199] Dynamic meteorological variables such as wind speed and wind direction are introduced to construct the adjacency matrix, which can ensure that the adjacency weights of monitoring stations are high when they are under similar meteorological conditions, thereby simulating the spatial propagation pattern of pollutants more accurately.

[0200] Among them, the adjacency matrix formula based on historical similarity in S2.1 is:

[0201]

[0202] in, and denote the PM10 concentrations at monitoring sites i and j at time step t, respectively, and δ controls the effect of PM concentration similarity.

[0203] By constructing an adjacency matrix based on the similarity of historical PM10 concentrations, it is possible to identify sites with similar pollution patterns, thus breaking through the limitations of traditional spatial adjacency and enhancing the model's ability to model complex pollution source distribution and nonlinear propagation mechanisms. It is particularly suitable for predicting sudden pollution events or cross-regional transmission scenarios.

[0204] Among them, S3.2 improves the diversity of the expert network by introducing the following mechanisms: Expert Regularization, which constrains the weights of experts through KL divergence or L2 regularization, so that they learn different feature patterns and prevent all experts from converging; Expert Specialization, which samples the expert input data so that each expert learns a specific PM10 change pattern, such as short-term bursts and long-term trend changes.

[0205] This mechanism effectively avoids the homogenization of expert models, improves the model's generalization ability for multi-scale, highly heterogeneous spatiotemporal data, and reduces the risk of overfitting.

[0206] In the MOE structure in S3.2, different experts can adopt different structures, and each architecture is suitable for different change pattern modeling requirements. For example, the MLP expert is used to learn basic pollution patterns; the LSTM expert is used to model the long-term dependence of pollution time series; and the Transformer expert is used to model complex spatiotemporal features.

[0207] By adopting heterogeneous expert architectures such as MLP, LSTM, and Transformer, we can model basic pollution patterns, long-term temporal dependencies, and complex spatiotemporal interactions respectively, so as to flexibly adapt to the characteristic requirements of different pollution scenarios and improve the comprehensiveness and accuracy of the prediction as a whole.

[0208] Among them, S3.3 improves the stability and diversity of the gating network by introducing the following strategies: Load Balancing Regularization, which prevents some experts from being overloaded while other experts are hardly activated; Temperature Annealing, which gradually reduces and enhances competition between experts, allowing the model to explore different experts in the early training stage and converge to the optimal expert combination in the later stage; Gating Gradient Clipping, which prevents drastic fluctuations in gating weights during gradient updates and improves model stability; Stochastic Gating: Adding a small amount of noise to the Softmax calculation of the gating network makes expert selection more robust.

[0209] By introducing load balancing regularization to prevent expert overload, temperature attenuation to enhance competitive selection, gradient clipping to suppress weight fluctuations, and noise perturbation to improve robustness, the stability and dynamic selection capabilities of the gating network are enhanced, ensuring the rational allocation of expert resources. Especially under complex and changeable pollution patterns, the model can quickly converge to the optimal expert combination, thereby improving the reliability and efficiency of the prediction.

[0210] Among them, during the construction of the spatial relationship diagram in S2.2, the system will dynamically adjust the correlation strength between different monitoring sites based on real-time meteorological data and geographical location relationships. If the weather in a certain area changes suddenly, the information transmission weight between related sites will be automatically enhanced to improve the accuracy of pollution diffusion modeling.

[0211] The weights between sites are dynamically adjusted according to real-time meteorological data, so that the model can quickly respond to the interference of sudden meteorological events on the pollutant transmission path.

[0212] Among them, the time feature extraction module in S1 will automatically adjust the feature extraction strategy according to different time periods when analyzing data, and reversely optimize the time analysis range through actual prediction errors during training, so that the model is more in line with the time change laws of actual scenarios.

[0213] By dynamically adjusting the temporal analysis range through backpropagation of errors and combining it with a multi-scale feature fusion strategy, the model can adapt to the pollution variation patterns in different time periods, making the temporal feature extraction more in line with the dynamic characteristics of actual data and improving the stability of long-term predictions.

[0214] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0215] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A PM10 concentration prediction method based on spatiotemporal graph neural network and expert mixture model, characterized by: The specific steps are as follows: S1. Temporal feature extraction (RTAF) PM10 concentration is affected by multiple temporal factors, including short-term fluctuations, medium-term trends, and long-term trends. To effectively model these temporal characteristics, we use multi-scale temporal feature extraction to obtain information at different time scales, and utilize a residual attention mechanism for feature enhancement. Finally, adaptive feature fusion is used to select the most important time steps for dynamic weighting. S1.1 Basic temporal feature modeling Given the PM10 concentrations and related meteorological data of multiple stations, the input sequence is defined as: in, Represents the feature vector of the i-th station at time t, including information such as PM10 concentration, temperature, humidity, and wind speed; Since the variation pattern of PM10 concentration is complex, involving a combination of short-term fluctuations and long-term trends, a single-scale feature extraction method is difficult to meet the prediction requirements. Therefore, a multi-layer temporal feature extraction network is first constructed. This network can capture global temporal dependencies and provide robust temporal feature representation. The time series of multi-scale input is represented as: H t =f(X t ) Where, X t is the input time series data, H t is the extracted temporal feature, f(·) is the mapping function of the feature extraction network; In order to enhance the interactivity between time steps and ensure that information at different time scales can be effectively transmitted, a residual attention mechanism is introduced to replace the traditional sequence modeling method. S1.2 Residual Attention Mechanism By introducing the residual attention mechanism, information between different time steps can interact more effectively; The residual attention mechanism calculation formula is as follows: Among them, Q, K, V are query, key and value matrices respectively, d k is the dimension of the key, PrevLayerOutput represents the output of the previous layer, and the softmax normalization operation is used to assign attention weights at different time steps; The residual attention mechanism not only enhances the interactivity between time steps, but also avoids the gradient vanishing problem in deep networks through residual connections; The final time features are calculated as follows: H residual =Residual Attention(H t ) S1.3 Adaptive feature fusion By introducing the adaptive feature fusion mechanism, the model can automatically learn which time steps are most important for the final prediction. The calculation process of adaptive feature fusion is as follows: α t =softmax(W x x t +W h h t-1 ) Among them, α t is the attention weight at time step t, W x ,W h is a trainable parameter, It is the temporal feature after residual attention enhancement, and the softmax operation ensures that the sum of the weights of all time steps is 1; Through the adaptive feature fusion mechanism, the most critical time steps for PM10 prediction can be dynamically screened. The time step selection strategy is continuously optimized during the model training process, allowing the model to adapt to different time feature patterns. The final output time feature calculation is: Time characteristic H exttime It will be used as input to the spatial feature extraction module for subsequent site relationship modeling; S2. Spatial feature extraction (STGNN) S2.

1. Spatial Dependence Modeling PM10 concentrations are affected not only by temporal variations but also by the geographic location of different regions, exhibiting significant spatial dependence. The diffusion of PM values between adjacent monitoring stations is influenced by factors such as wind speed, wind direction, topography, and emission sources, which in turn dynamically change over time. These factors jointly determine the pollution similarity and the intensity of mutual influence between monitoring stations. By introducing a dynamic multimodal weighted graph (DMWG), we can more comprehensively model the spatial dependence of PM10 concentrations. To do this, consider the following three key spatial dependencies: Geographical proximity: The physical distance between monitoring stations determines the spatial diffusion capacity of pollutants. Adjacent monitoring stations usually have strong pollution correlations. Meteorological factors such as wind speed and direction affect the transmission path of pollutants, making it possible for strong pollution relationships to exist between some distant measuring stations. Historical similarity: PM10 concentration variation patterns at different stations may be similar, even if they are geographically distant. However, strong spatial dependencies may still exist due to similar pollution sources or seasonal variations. DMWG will use a weighted fusion strategy to combine these three spatial dependencies, allowing the model to not only capture static geographic proximity information but also adaptively adjust for the contributions of meteorological influences and historical pollution patterns. S2.2 Graph Construction Method Dynamic adjacency matrix generation: A dynamic multimodal weighted graph (DMWG) is constructed, where nodes represent monitoring stations and edge weights are calculated based on a combination of factors. The core of the graph construction lies in defining the adjacency matrix A, which determines how information is propagated between stations. Incorporating multimodal factors, the adjacency matrix is generated using the following strategy: A=αA1+βA2+γA3 Among them, α, β, and γ are trainable parameters, ensuring that the sum of their weights is 1, so that the model can adapt to the pollution propagation mode in different situations; to avoid numerical explosion, the adjacency matrix needs to be normalized. Where D is the degree matrix of the adjacency matrix A, which is used to ensure that the information of different stations will not be overly concentrated or dispersed due to the imbalance of the adjacency matrix; Adaptive adjacency matrix update mechanism: By introducing an adaptive adjacency matrix update mechanism, the adjacency matrix can be adjusted according to the current pollution propagation pattern: A (t) =ηA (t-1) +(1-n)A new Where A (t) is the adjacency matrix at time step t, A (t-1) is the adjacency matrix of the previous time step, A new is the currently calculated adjacency matrix, and η is used to control the smoothness of the adjacency matrix update, which is used to ensure that the model can adapt to the dynamic changes of PM10 propagation over time and improve the accuracy of spatiotemporal prediction; Graph structure construction: Based on the calculated dynamic adjacency matrix, the PM10 prediction problem is represented as a spatiotemporal graph G = (V, E, A). The node set V represents air quality monitoring stations, each serving as a node in the graph. The edge set E represents the spatial connectivity between monitoring stations, determined by the adjacency matrix A. The adjacency matrix A is used to describe the PM10 propagation pattern between monitoring stations. In the spatiotemporal diagram, the PM10 concentration data of the measuring station changes dynamically over time, forming a spatiotemporal data sequence, namely: X={X1,X2,…,X T } Where, X t represents the PM10 data of all monitoring stations at time step t; finally, STGNN is applied on the constructed graph structure to learn spatiotemporal features to obtain the pollution interaction pattern between monitoring stations; S2.

3. Spatial feature extraction and calculation After constructing the spatiotemporal graph structure and defining the dynamic adjacency matrix, the model needs to extract effective spatial features based on this to learn the pollution propagation pattern between monitoring stations. By using the spatiotemporal graph neural network (STGNN) for calculation, PM10 information between monitoring stations can be efficiently transmitted and used for the final PM10 prediction. S2.3.

1. Graph Convolutional Neural Network (GCN) Computation By utilizing the graph structure information, each station node can not only aggregate information based on its own PM10 concentration, but also aggregate information from neighboring stations. The GCN calculation formula is as follows: Where, is the normalized adjacency matrix, which is used to ensure the stability of feature propagation; H (l) is the node feature matrix of the lth layer; W (l) is a trainable parameter matrix; σ is a nonlinear activation function; The graph convolution operation ensures that the information of PM10 concentration can be diffused along the propagation path and can effectively learn the mutual influence between different monitoring sites; S2.3.2 Optimization strategy for spatial feature extraction GCN may encounter the problem of information over-smoothing when learning spatial features. That is, as the number of network layers increases, the features of different monitoring sites will gradually converge, resulting in a decrease in prediction ability. To solve this problem, the following optimization strategy is introduced: Skip Connection: Add residual connections between GCN layers to preserve information in the deep network while preventing gradient vanishing. The calculation method is as follows, which is used to ensure that the original feature information is not completely submerged in multi-layer propagation. Graph Attention Network (GAT): GCN uses a fixed adjacency matrix, while GAT allows the model to dynamically adjust weights based on the feature similarity between stations, improving the model's expressiveness. Its formula is as follows: The final updated version is as follows: In this way, monitoring stations can dynamically pay attention to the neighboring stations that have the greatest impact on them, improving the ability to model the spread of PM values; Multi-Scale Graph Convolution (Multi-Scale GCN): PM10 diffusion may involve different scales in space. Therefore, by using multi-scale graph convolution, the model can learn both local and global propagation features. The calculation method is as follows: Where H (s) Represents graph convolution features of different scales; W s Trainable parameters of the same scale; S2.3.3 Final output of spatial features After applying the above GCN calculation and optimization strategy, the final spatial feature representation is obtained: H s =STGNN(H time ,A) Among them: H time is the output of the temporal feature extraction module; A is the dynamic adjacency matrix constructed by DMWG; the final spatial feature H (s) It will be input into MoE together with the time features for the final expert hybrid decision; S3, Mixture of Experts (MOE) By adopting a Mixture of Experts (MoE) network, multiple expert sub-models are used to process different propagation modes. The gating network dynamically selects the most appropriate expert and generates the final prediction result through weighted fusion of experts, thereby improving the accuracy of prediction. S3.

1. Overview of MoE Structure The MoE consists of three parts: an expert network (experts), which consists of multiple independent sub-models responsible for learning the characteristics of different pollution patterns. By integrating multiple experts, the model's prediction accuracy and generalization ability are improved, making it suitable for PM10 concentration prediction under different pollution patterns; a gating network, which dynamically adjusts the weights of each expert based on input features; and a weighted expert fusion, which performs a weighted summation of the outputs of multiple experts to form the final prediction. The calculation process of the entire MoE structure is as follows: Among them, K is the number of experts, G k (H s ) is the weight of the kth expert calculated by the gating network; E k (H s ) is the prediction result of expert k, and S is the set of K selected experts. This means that among all the experts, the gating network selects the most important K experts for the calculation of the current sample, rather than having all experts participate in the calculation at the same time; S3.2 Expert Networks Each expert network E k Responsible for modeling different PM10 change patterns, the calculation method is as follows: E k (H s )=σ(W k H s +b k ) Where H s is the spatiotemporal network extracted by STGNN, W k , b k is the trainable parameter of expert k; σ is the nonlinear activation function; S3.3 Gating Network The gating network is used to dynamically select the most suitable expert. It calculates the weight of each expert through Softmax: Among them, W g , b g is a trainable parameter of the gating network; τ is a temperature parameter used to control the uniformity of expert selection; S3.

4. Final output of MOE prediction MoE calculates PM10 predictions through weighted fusion of experts. Uncertainty modeling is introduced to improve MoE's effectiveness in PM10 prediction: Bayesian MoE models the uncertainty output by the expert network, ensuring that the prediction results include not only concentration estimates but also confidence intervals; Uncertainty Quantification, combined with Monte Carlo Dropout or Deep Ensemble, calculates the uncertainty of the prediction, thereby improving the credibility of the model.

2. The PM10 concentration prediction method based on a spatiotemporal graph neural network and a mixture of experts model according to claim 1 is characterized by: The adjacency matrix based on geographic distance described in S2.1 assumes that the pollution diffusion capacity between monitoring stations is inversely proportional to their geographic distance. The spatial relationship between monitoring stations i and j is defined as: where d ij is the Euclidean distance between sites i and j; σ is a hyperparameter that controls the distance decay rate; it ensures that the connection strength between adjacent monitoring stations is high, while the influence of distant monitoring stations is small.

3. The PM10 concentration prediction method based on a spatiotemporal graph neural network and a mixture of experts model according to claim 1 is characterized by: The adjacency matrix based on meteorological factors described in S2.1 needs to be further constructed based on meteorological variables because the diffusion of PM10 is greatly affected by meteorological factors such as wind speed and wind direction: Among them, w i and w j represent the meteorological factors at monitoring site i and monitoring site j respectively, and λ is a smoothing factor used to control the influence of meteorological variables.

4. The PM10 concentration prediction method based on a spatiotemporal graph neural network and a mixture of experts model according to claim 1 is characterized by: The adjacency matrix formula based on historical similarity described in S2.1 is: in, and denote the PM10 concentrations at monitoring sites i and j at time step t, respectively, and δ controls the effect of PM concentration similarity.

5. The PM10 concentration prediction method based on spatiotemporal graph neural network and expert mixture model according to claim 1 is characterized by: As described in S3.2, the diversity of the expert network is improved by introducing the following mechanisms: Expert Regularization, which constrains the weights of experts through KL divergence or L2 regularization, so that they learn different feature patterns and prevent all experts from converging; Expert Specialization, which samples the expert input data so that each expert learns a specific PM10 change pattern, such as short-term bursts and long-term trend changes.

6. The PM10 concentration prediction method based on spatiotemporal graph neural network and expert mixture model according to claim 1 is characterized by: As described in S3.2, in the MOE structure, different experts can adopt different structures. Each architecture is suitable for different change pattern modeling requirements. For example, the MLP expert is used to learn basic pollution patterns; the LSTM expert is used to model the long-term dependence of pollution time series; and the Transformer expert is used to model complex spatiotemporal features.

7. The PM10 concentration prediction method based on a spatiotemporal graph neural network and a mixture of experts model according to claim 1 is characterized by: As described in S3.3, the following strategies are introduced to improve the stability and diversity of the gating network: Load Balancing Regularization, which prevents some experts from being overloaded while others are barely activated; Temperature Annealing, which gradually reduces the temperature and enhances competition between experts, allowing the model to explore different experts in the early stages of training and converge on the optimal expert combination in the later stages; Gating Gradient Clipping, which prevents drastic fluctuations in gating weights during gradient updates and improves model stability. Stochastic Gating: Adds a small amount of noise to the Softmax calculation of the gating network to make the expert selection more robust.

8. The PM10 concentration prediction method based on a spatiotemporal graph neural network and a mixture of experts model according to claim 1 is characterized by: During the construction of the spatial relationship diagram described in S2.2, the system will dynamically adjust the correlation strength between different monitoring sites based on real-time meteorological data and geographical location relationships. If the weather in a certain area changes suddenly, the information transmission weight between related sites will be automatically enhanced to improve the accuracy of pollution diffusion modeling.

9. The PM10 concentration prediction method based on spatiotemporal graph neural network and expert mixture model according to claim 1 is characterized by: When analyzing data, the time feature extraction module described in S1 will automatically adjust the feature extraction strategy according to different time periods, and reversely optimize the time analysis range through actual prediction errors during training, so that the model is more in line with the time change laws of actual scenarios.

Citation Information

Cited By

  • Air quality prediction method, device and equipment based on multi-modal space-time fusion and medium

    CN121278686A

  • Refining device intelligent prediction method based on physical information neural network

    CN121365205A

  • Power load prediction method based on dynamic expert pool and load balancing mechanism MoE

    CN121525994A

  • Ground station meteorological element forecasting method based on hybrid expert model and graph attention

    CN121724065A

  • A ground-based meteorological element forecasting method based on hybrid expert models and graph attention.

    CN121724065B