Space-time sand storm event AI forecasting method fusing satellite remote sensing and meteorological data

Through the AI ​​prediction method of space-time sandstorm events that integrates satellite remote sensing and meteorological data in sandstorm prediction, the DustMamba model with a multi-task architecture is adopted to solve the problem of multi-source space-time sandstorm modeling in the existing technology, realize high-precision sandstorm prediction, and provide scientific basis for meteorological disaster warning and environmental protection.

CN120162584AActive Publication Date: 2025-06-17LANZHOU UNIV

Patent Information

Application Number
CN202510225408.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently model multi-source spatio-temporal sand and dust data, resulting in insufficient accuracy and reliability of sandstorm prediction, especially when real-time and high-frequency predictions are carried out in large areas, which is expensive.

Method used

A spatiotemporal and dust storm event AI prediction method integrating satellite remote sensing and meteorological data is proposed. DustMamba, a sandstorm forecast model with a multi-task architecture, includes a spatiotemporal encoder, feature aggregation layer and a task-specific layer. Through deep learning technology, the space-time dependence information in multi-source data can be effectively extracted to achieve efficient and multi-scale prediction of sandstorm events.

Benefits of technology

It has achieved high-precision prediction of sandstorm events, can accurately predict the occurrence time, intensity and impact range of sandstorms, and provides scientific basis for meteorological disaster warning, environmental protection and public health prevention, so as to reduce ecological, social and economic losses caused by sandstorms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162584A_ABST
    Figure CN120162584A_ABST
Patent Text Reader

Abstract

The invention discloses a space-time sandstorm event AI forecasting method fusing satellite remote sensing and meteorological data, and relates to the technical field of atmospheric pollution prediction and artificial intelligence. In order to solve the problem that multi-source space-time sand and dust data is difficult to model efficiently in the existing method, the invention designs a sand and dust storm forecasting model DustMama based on a multi-task architecture, and aims to predict the future PM10 concentration and the occurrence condition of the sand and dust storm through the combination of satellite remote sensing and meteorological data. The DustMama is composed of a space-time encoder, a feature aggregation layer and a task specific layer. The space-time encoder is combined with the visual Mamba and the three-dimensional convolutional network, so that the space-time dependent information in the multi-source data can be efficiently extracted. The feature aggregation layer adopts a global attention module to enhance the capability of the model in cross-dimension feature interaction. And the task specific layer is based on an independent two-dimensional convolution predictor so as to realize accurate prediction of different sand and dust tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of air pollution prediction and artificial intelligence, and more specifically, to an AI forecasting method for spatio-temporal sandstorm events that integrates satellite remote sensing and meteorological data. Background Art

[0002] A sandstorm is a natural phenomenon in which dust particles suspended and spread in the air are caused by strong winds, and often occurs in arid and semi-arid regions. The occurrence mechanism of sandstorms involves multiple factors, including soil dryness, increasing wind speed, scarce surface vegetation, and adverse meteorological conditions, etc. These factors act together to cause the lifting, spreading, and settling processes of dust particles. Sandstorms not only cause soil erosion and vegetation damage, but also significantly reduce air quality, increase the concentration of suspended particulate matter (such as PM10), and thus pose a serious threat to human health. In addition, the occurrence of sandstorms may also lead to economic impacts such as agricultural production losses, increased traffic accidents, and infrastructure damage, and may even affect the normal operation of society in severe cases. Therefore, accurately predicting sandstorm events can not only take protective measures in advance and reduce health risks, but also provide a scientific basis for decision-making, thereby reducing the ecological, social, and economic losses brought by sandstorms.

[0003] Currently, numerical forecasting methods are traditional means for sandstorm prediction. This method mainly relies on meteorological models to analyze and simulate various factors that may affect the occurrence of sandstorms, such as the selection of dust source areas, wind speed, soil moisture, vegetation coverage, and atmospheric circulation, etc., and then quantitatively describes the generation, propagation path, and spatial distribution of sandstorms. Such numerical models can provide relatively intuitive sandstorm prediction results, but they have several limitations. First, numerical forecasting methods are extremely sensitive to the accuracy of initial conditions, and the high complexity and dynamic instability of the atmospheric system often lead to error accumulation, thereby affecting the accuracy and reliability of the prediction. Second, in order to achieve high spatio-temporal resolution prediction, numerical forecasting models usually require a large amount of computing resources and high-performance computing infrastructure, which makes it difficult and costly to conduct real-time and high-frequency sandstorm prediction in a large-scale area.

[0004] With the rapid development of artificial intelligence technologies, especially machine learning and deep learning technologies, more and more research has begun to apply these technologies to solve meteorological prediction tasks. In recent years, machine learning methods have made remarkable progress in the fields of meteorological element prediction, air quality prediction, etc., especially in the prediction of local sandstorms and fine particulate matter (such as PM2.5, PM10, etc.), showing strong prediction capabilities. However, these machine learning-based studies usually focus on local area prediction and mostly rely on ground meteorological observation data or local meteorological data with short time series. In terms of modeling the global dynamics and complex spatio-temporal characteristics of sandstorms, the effectiveness and adaptability of these methods still have certain deficiencies.

[0005] The formation and propagation of sandstorms involve multiple complex factors, including wind speed, air pressure, humidity, soil conditions, and the hierarchical structure of the atmosphere, etc. These factors usually exhibit high nonlinearity and mutual coupling in space and time. Traditional machine learning methods often struggle to comprehensively capture their spatio-temporal variation patterns, thereby limiting their prediction accuracy and applicability. In addition, existing machine learning models generally fail to fully integrate multi-source data (such as satellite remote sensing data and meteorological data) and lack effective spatio-temporal modeling mechanisms. When dealing with large-scale sandstorm predictions, it is difficult for them to fully exploit the relevance and complementarity between different data sources, thus affecting the prediction performance of sandstorms. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a spatio-temporal sandstorm event AI forecasting method that fuses satellite remote sensing and meteorological data. Aiming at the problem that existing methods are difficult to efficiently model multi-source spatio-temporal dust data, a sandstorm forecasting model DustMamba based on a multi-task architecture is proposed, aiming to jointly predict the future PM10 concentration and the occurrence of sandstorms through satellite remote sensing and meteorological data. DustMamba consists of three parts: a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. The spatio-temporal encoder combines Vision Mamba and a three-dimensional convolutional network to efficiently extract spatio-temporal dependence information in multi-source data. The feature aggregation layer uses a global attention module to enhance the model's ability in cross-dimensional feature interaction. The task-specific layer is based on an independent two-dimensional convolutional predictor to achieve accurate prediction of different dust tasks. The method of the present invention has significant technical advantages, can achieve efficient and multi-scale prediction of sandstorm events, and provides more accurate and timely data support for meteorological forecasting and emergency response. By accurately predicting the occurrence time, intensity, and impact range of sandstorms, the present invention provides a reliable scientific basis for the early warning of meteorological disasters, environmental protection, and public health prevention.

[0007] In a first aspect, the present invention provides a spatio-temporal sandstorm event AI forecasting method that fuses satellite remote sensing and meteorological data, and the method includes:

[0008] Obtain a data set and preprocess the data set to obtain a preprocessed data set; wherein, the data set includes satellite image data, meteorological reanalysis data, daily PM10 data, and air quality data;

[0009] Perform spatio-temporal alignment and interpolation on the preprocessed data to obtain unified input data;

[0010] Build a dust forecast model; among them, the dust forecast model includes a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. Based on unified input data, it processes the data into an input tensor with dimensions (B, C, T, H, W), where B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After the input passes through the dust forecast model, the outputs of four tasks are obtained, namely two regression prediction tasks and two classification prediction tasks. The output shape of each task is (B, H, W), representing the dust distribution at the next moment;

[0011] For multiple dust event forecast tasks, build a multi-task loss function, and train the dust forecast model based on the multi-task loss function;

[0012] During the model training process, use the comprehensive loss as the optimization target, continuously adjust the model parameters through the backpropagation algorithm, gradually minimize the multi-task loss, and obtain the trained dust forecast model.

[0013] Furthermore, obtain a dataset and preprocess the dataset to obtain a preprocessed dataset, including:

[0014] Clean the data in the dataset, use statistical methods to detect and remove data points with abnormal fluctuations, eliminate outliers and noise data, and ensure the accuracy and integrity of the data; for missing values caused by equipment failures or data transmission problems, use the cubic spline interpolation method to complete them to ensure the integrity of the data;

[0015] Adopt a normalization method to compress the cleaned data into a preset numerical range to eliminate the dimensional differences between different data dimensions.

[0016] Furthermore, perform spatio-temporal alignment and interpolation on the preprocessed data to obtain unified input data, including:

[0017] Project various types of data in the preprocessed data onto a unified spatial grid to ensure the spatial alignment of all data. For PM10 concentration data with a resolution lower than the set threshold, use bilinear interpolation or spline interpolation to adjust it to the same high-resolution grid as the satellite remote sensing data;

[0018] Unify all data to an hourly time interval, and fill in the missing data at different times through interpolation methods;

[0019] For meteorological data and satellite remote sensing data, use the linear interpolation method to smoothly fill the data with a short time interval to ensure the continuity and consistency of the time series;

[0020] For daily PM10 data with only daily intervals, the air quality station data at hourly granularity is used to correct it to obtain gridded PM10 data at hourly granularity.

[0021] Furthermore, the spatio-temporal encoder comprehensively extracts spatio-temporal features through a dual-channel architecture integrating 3D convolution and the Mamba module, and couples them through the Hadamard product to ensure effective capture of local details and global context; among them, the dual-channel architecture includes a convolution channel and a Mamba channel;

[0022] In the convolution channel, the encoder uses two layers of three-dimensional convolution to extract local spatio-temporal features. Given a first input tensor x ∈ R B×C×T×H×W , where R represents the set of real numbers, the convolution channel obtains a first output h c1 through the following convolution operation:

[0023] h c1 = Conv3D2(ReLU(Conv3D1(x)))

[0024] where Conv3D1 and Conv3D2 both represent three-dimensional convolution modules with a convolution kernel size of 3×3×3, and ReLU is the rectified linear activation function;

[0025] In the Mamba channel, the first input tensor x ∈ R B×C×T×H×W is divided into non-overlapping patches of size P×P through a patch embedding-based method:

[0026] x′ = FC(Rearrange(x, P)),

[0027] where FC represents the fully connected layer, P is the patch size, Rearrange represents the patch embedding operation, x′ is the reshaped tensor, and x′ ∈ R B×N×D , where is the number of embedded patches, and D = P 2 · C is the embedding dimension;

[0028] After dividing into non-overlapping patches, it is processed through multiple Vim blocks;

[0029] In each Vim block, the reshaped tensor x′ generates a second output x″ by combining bidirectional sequence modeling and structured SSM; the reshaped tensor x′ is first normalized and linearly projected into two feature representations, namely the first feature u and the second feature v. Subsequently, in the forward and backward directions respectively, a one-dimensional convolution operation is applied to the first feature u to generate an intermediate feature u′ o , and the intermediate feature u′ o is converted into a set of learnable parameters A o , B o , Co and Δ o , where Δ o is made positive by applying the softplus activation function, and Δ o has the following specific calculation formula:

[0030]

[0031] where, Δ o is obtained by linearly transforming Δ o and adding a learnable bias parameter b, and then using the softplus activation function to ensure its value is positive, thereby adjusting the non-linear transformation of the model;

[0032] The latent state h o is recursively updated through the SSM:

[0033]

[0034] where, represents the matrix multiplication operation discretized by time step;

[0035] The output y in each direction o is calculated by the following formula:

[0036]

[0037] The forward output y forward and the backward output y backward are gated using v and combined as:

[0038] y combined = y forward ⊙SiLU(v)+y backward ⊙SiLU(v)

[0039] where, SiLU represents the Sigmoid linear unit activation function; ⊙ represents the Hadamard product operation;

[0040] Finally, the combined result is linearly transformed and added to x′ through a residual connection to generate the second output x″:

[0041] x″ = Linear(y combined ) + x′

[0042] where, Linear represents the linear layer;

[0043] For the output y T ∈R B×N×D at the last moment of the Mamba channel, a reverse embedding operation is used to reconstruct the feature map to the initial data dimension, obtaining the final output h c2 ∈RB×C×T×H×W , match the data of the convolutional channels.

[0044] Furthermore, the feature aggregation layer uses a GAM module to enhance the feature representation provided by the spatio-temporal encoder; wherein, the GAM module integrates channel attention and spatial attention for weighted feature mapping, dynamically emphasizes channel dependencies and spatial correlations, and suppresses irrelevant noise information;

[0045] For a given second input tensor h c ∈R B×(C·T)×H×W , channel attention transforms it into a three-dimensional tensor h' c ∈R B×(H·W)×(C·T) , uses a two-layer fully connected network to identify the importance of features, and applies it to the original input through element-wise multiplication:

[0046] h CA = h c ·σ(FC2(ReLU(FC1(h c '))))

[0047] where FC1 and FC2 represent fully connected layers for adjusting the channel dimension, ReLU represents the rectified linear activation function, σ represents the sigmoid function; h CA represents the feature map;

[0048] Adopt a spatial attention module to capture the local spatial correlation in the feature map h CA , and emphasize the key areas through convolution:

[0049] h SA = h CA ·σ(Conv2(ReLU(BN(Conv1(h CA )))))

[0050] where Conv1 and Conv2 represent two layers of two-dimensional convolution for adjusting the channel dimension, BN represents batch normalization, and h SA represents the feature map finally generated by the GAM module.

[0051] Furthermore, the task-specific layer processes the aggregated features through a separate predictor or classifier to generate prediction results for multiple output tasks; for a regression task that directly outputs a numerical value, the predictor consists of two two-dimensional convolutional layers:

[0052]

[0053] where and represent task-specific convolutional layers; represents the model output result of a certain task;

[0054] The classifier adds a binary classification activation function on the basis of the predictor, expressed as:

[0055]

[0056] Furthermore, for multiple sand and dust event forecasting tasks, a multi-task loss function is constructed, and the sand and dust forecasting model is trained based on the multi-task loss function, including:

[0057] In each training cycle, the input data and the true label y i used to calculate the prediction of the i-th task The regression task is calculated using the mean squared error loss:

[0058]

[0059] where L i represents that of the i-th task; represents the prediction result of the j-th sample of the i-th task; N represents the total number of label data samples; j represents the sample index;

[0060] The classification task is trained using binary cross-entropy loss:

[0061]

[0062] The total loss L of a batch is calculated by the following formula batch :

[0063]

[0064] where w i is the weight of the i-th task, and Task is the total number of tasks;

[0065] To dynamically balance tasks, w is dynamically adjusted during training according to the gradient norm i ; the gradient norm g of the i-th task i is:

[0066]

[0067] where, represents the gradient with respect to the parameter θ;

[0068] The average gradient norm for all tasks is The task weights are updated by the following formula to minimize the difference between g i and :

[0069]

[0070] where η is the learning rate of the weights;

[0071] The weights are normalized to ensure that their sum is 1:

[0072]

[0073] where w j represents the j-th weight value;

[0074] In each training iteration, the total loss L is minimized batch , and the model parameters are updated through backpropagation to ensure balanced optimization for all tasks.

[0075] Furthermore, after obtaining the trained dust storm prediction model, the method further includes:

[0076] Evaluating the trained dust storm prediction model using evaluation metrics;

[0077] For regression tasks (predictions of PM10 and BADI), the mean squared error MSE, root mean squared error RMSE, mean absolute error MAE, and / or coefficient of determination R 2 are used to evaluate the prediction accuracy of the model;

[0078] For classification tasks, accuracy, recall, precision, and / or F1 score are used for evaluation.

[0079] In a second aspect, the present invention provides a spatio-temporal dust storm event AI prediction device that fuses satellite remote sensing and meteorological data. The device includes:

[0080] A data preprocessing unit configured to obtain a data set and preprocess the data set to obtain a preprocessed data set; wherein, the data set includes satellite image data, meteorological reanalysis data, daily PM10 data, and air quality data;

[0081] An alignment and interpolation unit configured to perform spatio-temporal alignment and interpolation on the preprocessed data to obtain unified input data;

[0082] A model construction unit configured to construct a dust storm prediction model; wherein, the dust storm prediction model includes a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. Based on the unified input data, it processes the input data into an input tensor of dimension (B, C, T, H, W), where B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After passing through the dust storm prediction model, the input data obtains the outputs of four tasks, namely two regression prediction tasks and two classification prediction tasks. The output shape of each task is (B, H, W), representing the dust distribution at the next moment;

[0083] A loss function determination unit, configured to construct a multi-task loss function for multiple dust event forecasting tasks, and train the dust forecasting model based on the multi-task loss function;

[0084] A model training unit, configured to use the comprehensive loss as the optimization objective during the model training process, continuously adjust the model parameters through the backpropagation algorithm, gradually minimize the multi-task loss, and obtain the trained dust forecasting model.

[0085] In a third aspect, the present invention provides a readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.

[0086] The present invention has at least the following beneficial effects:

[0087] The present invention utilizes the multi-task learning framework in deep learning to successfully construct the DustMamba model, achieving high-precision prediction of sandstorm events. By integrating multi-source data, including remote sensing data from FY-4A satellites, meteorological reanalysis data, PM10 concentration data, etc., and by fusing information at different time scales and spatial resolutions, the spatio-temporal characteristics of sandstorm occurrences are comprehensively captured.

[0088] First, in the data preprocessing stage, by cleaning, standardizing, and spatio-temporally aligning the multi-source data, the differences and inconsistencies between the data are eliminated, ensuring the high quality and structurality of the input data. This stage includes filling missing values, removing outliers, and spatio-temporal alignment between different data sources to ensure data consistency and availability, providing a reliable data basis for subsequent modeling.

[0089] Next, the DustMamba model realizes accurate prediction of multiple tasks (such as PM10 concentration prediction, sandstorm occurrence probability prediction, etc.) by designing a multi-task architecture of a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer, and optimizing the model parameters with a multi-task loss function. In the spatio-temporal encoder, the model can efficiently extract spatio-temporal dependencies in satellite remote sensing and meteorological data by introducing the Vision Mamba structure and a three-dimensional convolutional network, making full use of the spatio-temporal information of the data. The feature aggregation layer adopts a global attention mechanism to further enhance the model's ability in cross-dimensional feature interaction. The task-specific layer is independently designed for each prediction task to ensure the prediction accuracy of different tasks.

[0090] Through the joint optimization of multitask losses, the DustMamba model can not only accurately predict PM10 concentrations, but also effectively predict the occurrence probability of sandstorms and other related meteorological indicators, such as BADI, DRBTD, DST, etc., thus providing a scientific basis for sandstorm early warning and emergency response.

[0091] In summary, the present invention has significant advantages. First, by utilizing the framework of multi-source data fusion and multitask learning, the potential correlations between data are fully exploited, enhancing the generalization ability and accuracy of the prediction model. Second, while ensuring efficient computation, the DustMamba model effectively processes large-scale spatio-temporal data and achieves good performance in spatio-temporal prediction of sandstorms, with higher accuracy and practicality compared to traditional methods. Finally, the present invention can provide more scientific and accurate technical support for the prediction of extreme weather events such as sandstorms, reduce disaster losses, and ensure the safety of life and property, having broad application prospects and important social value. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 Shows the overall flowchart of a spatio-temporal sandstorm event AI prediction method that fuses satellite remote sensing and meteorological data according to an embodiment of the present invention.

[0093] Figure 2 Shows the schematic diagram of the model structure according to an embodiment of the present invention;

[0094] Figure 3 Shows the structure diagram of a spatio-temporal sandstorm event AI prediction device that fuses satellite remote sensing and meteorological data according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0095] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific examples, but shall not be construed as a limitation to the present invention. For the various steps described herein, if there is no necessity for a sequential relationship between them, the order in which they are described as examples herein should not be regarded as a limitation, and those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed and the entire process cannot be realized.

[0096] In the field of atmospheric science and meteorological forecasting, dust storms, as extreme weather events with severe environmental and social impacts, have received extensive attention. Dust storms not only lead to the deterioration of air quality, increase the concentration of suspended particulate matter (such as PM10), and seriously threaten human health, but also cause significant losses to agricultural production, transportation, and infrastructure. Therefore, accurately and timely predicting the occurrence and intensity of dust storms is of great significance for reducing disaster losses and ensuring the safety of people's lives and property. However, due to the suddenness, locality, and complexity of dust storm events, existing prediction methods still have certain limitations.

[0097] Although existing numerical prediction methods can simulate atmospheric processes, they have high requirements for computing resources, and the process of optimizing model parameters is complex. For weather events with strong locality and suddenness like dust storms, it is often difficult to achieve high-precision predictions. In addition, numerical methods usually rely on accurate knowledge of initial conditions, and the occurrence of dust storms is affected by various complex factors, which makes numerical models have great uncertainties in practical applications. On the other hand, although machine learning methods based on statistics have made certain progress in predicting pollutant concentrations, it is difficult to effectively provide a comprehensive spatio-temporal dust distribution, and it is difficult to handle complex dust storm events. This makes existing prediction methods still have problems of insufficient accuracy and low efficiency when dealing with large-scale dust storm events with high spatio-temporal resolution.

[0098] In view of the above problems, the embodiments of the present invention provide an AI forecasting method for spatio-temporal dust storm events that integrates satellite remote sensing and meteorological data. Through the effective integration of multi-source data, the spatio-temporal accuracy of dust storm prediction is improved, and the problems of data isolation and insufficient spatio-temporal information extraction in traditional dust storm prediction methods are overcome. Secondly, a deep learning multi-task modeling framework is proposed. Through the design of spatio-temporal encoders, feature aggregation layers, and task-specific layers, it is possible to accurately predict multiple dust storm-related tasks (such as the probability of dust storm occurrence, PM10 concentration, dust storm intensity, etc.) while capturing the spatio-temporal characteristics and non-linear relationships of dust storms. Extreme weather events such as dust storms have a serious impact on human production and life. Accurate prediction can provide a scientific basis for disaster prevention and mitigation, emergency response, and social and economic activities. Through the implementation of the present invention, it will make an important contribution to reducing the disaster losses caused by dust storms, ensuring the safety of people's lives and property, further improving the intelligent level of meteorological forecasting and environmental management, and promoting the scientific response to meteorological disasters and social security.

[0099] It should be noted that the Mamba model is a deep learning architecture based on the structured state space model (SSM), specifically designed for efficiently processing long sequence data. By introducing a selection mechanism to reparameterize the SSM, Mamba can filter out irrelevant information while retaining key features, and has excellent linear scalability. Different from traditional convolutional operations, Mamba adopts a hardware-aware algorithm, which improves computational efficiency through a scanning mechanism, especially significantly accelerating computations on large-scale GPUs. When processing long sequence data, its computational overhead grows linearly with the sequence length, having obvious advantages compared to models such as Transformer, and is particularly suitable for fields such as natural language processing, computer vision, and healthcare. These advantages make it an ideal choice for processing dust spatio-temporal modeling tasks.

[0100] Therefore, the core of the method in this embodiment lies in proposing an innovative multi-task dust prediction model DustMamba based on deep learning hybrid modeling. This model fully exploits the potential dust features in satellite remote sensing and meteorological data through an efficient Mamba architecture to achieve accurate sandstorm prediction. By designing a multi-task architecture of a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer, the proposed model can effectively capture the spatio-temporal characteristics of sandstorms while achieving accurate prediction of different dust detection indicators, thus providing support for the accurate forecasting and emergency response of sandstorm events.

[0101] Specifically, as Figure 1 shown, the AI forecasting method for spatio-temporal sandstorm events integrating satellite remote sensing and meteorological data includes the following steps S10 to S50.

[0102] S10. Obtain a dataset and preprocess the dataset to obtain a preprocessed dataset; wherein, the dataset includes satellite image data, meteorological reanalysis data, daily PM10 data, and air quality data.

[0103] In an exemplary embodiment, in the data collection and preprocessing stage, high-quality datasets from multiple sources are used to construct a sandstorm prediction model. Specifically, it includes FY-4A satellite image data and meteorological reanalysis data, 1 km high-resolution daily PM10 data in China, and national urban air quality data. By integrating these multi-source data, a high-quality input set with spatio-temporal consistency is constructed.

[0104] First, clean the data, use statistical methods to detect and remove data points with abnormal fluctuations, eliminate outliers and noise data to ensure the accuracy and integrity of the data. For missing values caused by equipment failures or data transmission problems, the cubic spline interpolation method is used for filling to ensure the integrity of the data.

[0105] Subsequently, due to the inconsistent data dimensions from different sources and the large difference in their numerical ranges (such as the order of magnitude difference between wind speed and PM10 concentration), the normalization (Min-Max Scaling) method is adopted to compress the data into a unified numerical range to eliminate the dimensional differences between different data dimensions. In this way, it can be ensured that in the subsequent model training process, the influence of various input features on the model is balanced.

[0106] The preprocessed dataset not only improves the data quality but also ensures the effective integration of multi-source data in sandstorm prediction, providing strong support for accurately capturing the spatio-temporal characteristics and meteorological impacts of sandstorms.

[0107] S20. Perform spatio-temporal alignment and interpolation on the preprocessed data to obtain unified input data.

[0108] In an exemplary embodiment, in the spatio-temporal alignment and interpolation step (i.e., step S20), the time frequency and spatial resolution of these data are further unified. First, due to the difference in spatial resolution of different data sources, various types of data need to be projected onto a unified spatial grid to ensure the spatial alignment of all data. For the PM10 concentration data with a lower resolution, interpolation methods (such as bilinear interpolation or spline interpolation) are used to adjust it to the same high-resolution grid as the satellite remote sensing data. Second, to handle the time mismatch problem between different data sources, all data are unified to an hourly time interval, and the missing data at different times are filled by interpolation methods. For meteorological data and satellite remote sensing data, the linear interpolation method is used to smoothly fill the data with a shorter time interval to ensure the continuity and consistency of the time series. For the 1 km high-resolution daily PM10 data in China with only daily intervals, the hourly-granularity national urban air quality station data are used to correct it to obtain hourly-granularity gridded PM10 data.

[0109] Through these spatio-temporal alignment and interpolation operations, the spatio-temporal differences between different data sources are successfully eliminated, providing high-quality and unified input data for subsequent feature fusion and model training, ensuring that the model can fully learn the spatio-temporal characteristics of sandstorm events and their complex relationships with meteorological elements.

[0110] S30. Construct a sandstorm prediction model; among them, the sandstorm prediction model includes a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. Based on the unified input data, it is processed into an input tensor with dimensions (B, C, T, H, W), where B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After passing through the sandstorm prediction model, the input obtains the outputs of four tasks, namely two regression prediction tasks and two classification prediction tasks. The output shape of each task is (B, H, W), representing the sandstorm distribution at the next moment.

[0111] It should be noted that in this paper, the "dust forecast model" is named "DustMamba". This model is a multi-task spatio-temporal sequence prediction framework with Mamba and convolution as its backbone, used to simultaneously process multiple prediction tasks of dust events. The model structure is as Figure 2 shown. DustMamba consists of three core components: a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. For a given input data X containing meteorological, satellite remote sensing, and dust index data, it is processed into an input tensor of dimension (B, C, T, H, W), where B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After passing through DustMamba, the input can obtain the outputs of four tasks: two regression prediction tasks (PM10 and BADI) and two classification prediction tasks (DRBTD and DST). The output shape of each task is (B, H, W), representing the dust distribution at the next moment.

[0112] In an exemplary embodiment, the spatio-temporal encoder comprehensively extracts spatio-temporal features through a dual-channel architecture integrating 3D convolution and the Mamba module, and couples them through the Hadamard product to ensure effective capture of local details and global context. In the convolutional channel, the encoder uses two layers of three-dimensional convolution to extract local spatio-temporal features. Given a first input tensor x∈R B×C×T×H×W , where R represents the set of real numbers, the first output h c1 obtained by this channel can be obtained through the following convolution operation:

[0113] h c1 = Conv3D2(ReLU(Conv3D1(x)))

[0114] where Conv3D1 and Conv3D2 represent three-dimensional convolution modules with a kernel size of (3×3×3), and ReLU is the rectified linear activation function. This channel is good at encoding local spatio-temporal relationships and is used to capture sudden changes in meteorological elements or satellite remote sensing images.

[0115] In the Mamba channel, first, the input tensor x∈R B×C×T×H×W is divided into non-overlapping patches of size P×P through a patch embedding-based method:

[0116] x′ = FC(Rearrange(x, P))

[0117] where FC represents the fully connected layer, P is the patch size, and Rearrange represents the patch embedding operation. These patches are flattened and linearly projected into a high-dimensional latent space. The reshaped tensor is x′∈R B×N×D , where is the number of embedded patches, D = P 2 · C is the embedding dimension. Subsequently, it is processed through multiple Vim blocks.

[0118] In each Vim block, the input tensor x′ generates the output x″ by combining bidirectional sequence modeling and structured SSM, thus achieving effective modeling of spatial and temporal features in visual tasks. The input tensor x′ is first normalized and linearly projected into two feature representations: the first feature u and the second feature v. Subsequently, in the forward and backward directions respectively, a one-dimensional convolution operation is applied to the first feature u to generate the intermediate feature u′ o , where o ∈ {forward, backward}. Then, the intermediate feature u′ o is converted into a set of learnable parameters A o , B o , C o and Δ o , where Δ o makes it positive by applying the softplus activation function. The specific calculation formula of Δ o is:

[0119]

[0120] where, Δ o is obtained by linearly transforming Δ o and adding a learnable bias parameter b, and then using the softplus activation function to ensure its value is positive, thereby adjusting the nonlinear transformation of the model.

[0121] The latent state h o is recursively updated through SSM:

[0122]

[0123] where, represents the matrix multiplication operation discretized by time step;

[0124] The output in each direction can be calculated by the following formula:

[0125]

[0126] The forward and backward outputs y forward and y backward are gated using v and combined as:

[0127] y combined = y forward ⊙ SiLU(v) + y backward ⊙ SiLU(v)

[0128] Among them, SiLU represents the Sigmoid linear unit activation function.

[0129] Finally, the merged result is linearly transformed and added to x′ through a residual connection to generate the second output x″:

[0130] x″ = Linear(y combined ) + x′

[0131] Among them, Linear represents the linear layer.

[0132] This Vim structure effectively integrates temporal dynamics and spatial modeling, ensuring that the model captures long-range dependencies and context relationships. Subsequently, for the output y T ∈R B×N×D at the last moment of Mamba, an inverse embedding operation is used to reconstruct the feature map to the initial data dimension to obtain the final output h c2 ∈R B×C×T×H×W , which is matched with the data in the convolutional channels.

[0133] By integrating the first output h c1 of the convolutional channel and the final output h c2 of the Mamba channel, the fused tensor h c is obtained:

[0134] h c = h c1 ⊙h c2

[0135] In the formula, ⊙ represents the Hadamard product operation.

[0136] This dual-channel fusion strategy comprehensively utilizes the advantages of two spatio-temporal modeling methods. Among them, the convolutional channel features provide high-resolution, local spatio-temporal information, while the Mamba channel features introduce large-scale context understanding, ensuring the effective encoding of local and global spatio-temporal patterns and providing rich feature representations for the accurate identification and prediction of sandstorms.

[0137] In an exemplary embodiment, the feature aggregation layer uses GAM to enhance the feature representation provided by the spatio-temporal encoder. GAM integrates channel attention and spatial attention for weighted feature mapping, dynamically emphasizing channel dependencies and spatial correlations, and suppressing irrelevant noise information. First, for the given second input tensor h c ∈R B×(C·T)×H×W , the channel attention transforms it into a three-dimensional tensor h′ c ∈R B×(H·W)×(C·T) , uses a two-layer fully connected network to identify the importance of features, and applies it to the original input through element-wise multiplication:

[0138] h CA= h c ·σ(FC2(ReLU(FC1(h c ′))))

[0139] Among them, the fully connected layers FC1 and FC2 are used to adjust the channel dimension, ReLU represents the rectified linear activation function, and σ represents the sigmoid function.

[0140] Furthermore, a spatial attention module is adopted to capture the local spatial correlation in the feature map h CA , and convolution is used to emphasize the key regions:

[0141] h SA = h CA ·σ(Conv2(ReLU(BN(Conv1(h CA )))))

[0142] Among them, the two-dimensional convolutional layers Conv1 and Conv2 are used to adjust the channel dimension, BN represents batch normalization, ReLU represents the rectified linear activation function, and σ represents the sigmoid function. The integration of channel and spatial attention ensures that the network can capture global and local dependencies. By combining these two complementary attention mechanisms, the GAM module enables the model to prioritize key information while maintaining computational efficiency, achieving effective aggregation of the spatio-temporal characteristics of dust.

[0143] In an exemplary embodiment, the task-specific layer processes the aggregated features through a separate predictor or classifier to generate prediction results for multiple output tasks. For a regression task that directly outputs a numerical value, the predictor consists of two two-dimensional convolutional layers with a kernel size of (1×1):

[0144]

[0145] Among them, Conv1 and Conv2 are task-specific convolutional layers, and the classifier adds a binary classification activation function on the basis of the predictor:

[0146]

[0147] In the formula, σ represents the sigmoid function; represents the model output result of a certain task. By independently processing the input feature h SA through multiple task-specific modules, the task-specific layer stacks the outputs of all tasks along a new dimension to generate the final prediction tensor

[0148] S40. For multiple dust event forecasting tasks, a multi-task loss function is constructed, and the dust forecasting model is trained based on the multi-task loss function.

[0149] In an exemplary embodiment, for multiple sand and dust event forecasting tasks, a multi-task loss function is constructed to ensure that all tasks are fully optimized without interference. In each training cycle, the input data x and the true label y i is used to calculate the prediction of the i-th task For the regression task, the mean square error (MSE) loss is used for calculation:

[0150]

[0151] where L i represents the i-th task; represents the prediction result of the j-th sample of the i-th task; N represents the total number of label data samples; j represents the sample index;

[0152] For the classification task, the binary cross entropy (BCE) loss is used for training:

[0153]

[0154] where N represents the total number of label data samples. The total loss for a batch is the weighted sum of the losses of all tasks:

[0155]

[0156] where w i is the weight of task i, and Task is the total number of tasks.

[0157] To dynamically balance tasks, the GradNorm method is introduced to dynamically adjust w during the training process according to the gradient norm i . The gradient norm of task i is:

[0158]

[0159] where, represents the gradient with respect to the parameter θ.

[0160] The average gradient norm for all tasks is Update the task weights to minimize the difference between g i and :

[0161]

[0162] where η is the learning rate of the weights. The weights are normalized to ensure that their sum is 1:

[0163]

[0164] where \(T\) represents the entire training duration and \(j\) is the training time index; \(w\) j represents the \(j\)-th weight value. In each training iteration, the total loss \(L\) is minimized batch , and the model parameters are updated through backpropagation, ensuring balanced optimization for all tasks.

[0165] S50. During the model training process, the comprehensive loss is used as the optimization objective, and the model parameters are continuously adjusted through the backpropagation algorithm to gradually minimize the multi-task loss, obtaining the trained dust storm prediction model.

[0166] In an exemplary embodiment, during the model training process, the comprehensive loss function \(L\) batch is used as the optimization objective, and the model parameters are continuously adjusted through the backpropagation algorithm to gradually minimize the multi-task loss, thereby improving the prediction accuracy of the model for each task. This comprehensive loss function integrates the losses of multiple tasks, ensuring that both regression tasks (such as the prediction of PM10 and BADI concentrations) and classification tasks (such as the classification of DRBTD and DST values) are considered simultaneously during the optimization process. In the regression task, by minimizing loss functions such as the mean squared error (MSE), the model can accurately fit the concentration values of PM10 and BADI; while in the classification task, by minimizing the cross-entropy loss, the classification performance of the model is optimized to ensure the accurate classification of DRBTD and DST values. Through such multi-task learning, the model can not only make accurate predictions of sandstorm-related indicators but also effectively share information between different tasks, improving the overall prediction accuracy. In each training iteration, the optimization of the loss function enables the model to adaptively adjust the parameters to minimize the prediction error and gradually approach the actual observed values. Through continuous iteration, the model can accurately predict the key indicators related to sandstorm events (such as PM10, BADI concentrations, and DRBTD, DST values), providing a more accurate scientific basis for sandstorm prediction.

[0167] In an exemplary embodiment, after the model training is completed, in the evaluation stage, a series of standardized indicators are used to comprehensively measure the prediction performance of the model. For the regression tasks (prediction of PM10 and BADI), indicators such as the mean squared error MSE, root mean squared error RMSE, mean absolute error MAE, and coefficient of determination \(R\) 2 are mainly used to evaluate the prediction accuracy of the model. MSE and RMSE measure the difference between the model prediction values and the true values, MAE evaluates the average absolute value of the prediction error, and \(R\) 2It is used to measure the goodness of fit and prediction accuracy of the model. For classification tasks (prediction of DRBTD and DST values), classification performance metrics such as accuracy, recall, precision, and F1 score are used for evaluation. Accuracy measures the proportion of correctly classified predictions, recall measures the ability of the model to identify positive class samples, precision evaluates the proportion of truly positive samples among those predicted as positive by the model, and the F1 score is the harmonic mean of precision and recall, comprehensively considering the performance balance of the classification model on different tasks. Through these evaluation metrics, the performance of the model in sandstorm prediction can be fully reflected, helping to further optimize and improve the model. Through this multi-dimensional evaluation method, the sandstorm prediction model of the present invention can fully verify its application effects on different tasks, providing scientific support for actual sandstorm event warnings.

[0168] In summary, the present invention innovatively designs the DustMamba model for sandstorm events. The following will elaborate on the key technical points of this technical solution in detail:

[0169] Multi-task learning model design: The innovation of the present invention lies in constructing a sandstorm prediction model DustMamba based on multi-task learning and introducing a multi-task loss optimization strategy to achieve the simultaneous prediction of multiple sandstorm-related indicators in the same model, such as PM10 concentration, sandstorm occurrence probability (BADI, DRBTD, DST), etc. By fusing satellite remote sensing data, meteorological reanalysis data, and high-resolution PM10 data, the DustMamba model can fully utilize the spatio-temporal information in these multi-source data to accurately capture the dynamic changes and spatial distribution characteristics of sandstorms, thereby achieving high-precision sandstorm prediction.

[0170] Spatio-temporal encoder and global attention mechanism: This technology effectively processes the complex spatio-temporal dependence of sandstorms through a spatio-temporal encoder and a global attention mechanism (GlobalAttention Mechanism, GAM). The spatio-temporal encoder combines 3D convolution and Vim blocks to extract spatio-temporal features at local and global scales from the input data, ensuring comprehensive feature representation. The global attention mechanism enhances the learning ability of the model during the feature aggregation stage, integrating channel attention and spatial attention to capture key spatio-temporal information, highlighting the key features affecting the occurrence and propagation of sandstorms, and improving the prediction accuracy of the model.

[0171] Task-specific prediction layer: The DustMamba model adopts a task-specific prediction layer, and each task (such as PM10 concentration prediction, sandstorm occurrence probability prediction, etc.) has an independent predictor. This design enables each task to be independently optimized based on the shared spatio-temporal features, avoiding interference between tasks and improving the robustness and prediction accuracy of the tasks.

[0172] Through these technological innovations, the DustMamba model not only improves the accuracy of sandstorm prediction but also significantly enhances the ability to process large-scale and high-complexity data, providing strong technical support for accurate forecasting and timely emergency response.

[0173] An embodiment of the present invention also provides a spatio-temporal sandstorm event AI forecasting device that integrates satellite remote sensing and meteorological data, as Figure 3 shown. The device includes:

[0174] A data preprocessing unit 301, configured to obtain a data set and preprocess the data set to obtain a preprocessed data set; wherein, the data set includes satellite image data, meteorological reanalysis data, daily PM10 data, and air quality data;

[0175] An alignment and interpolation unit 302, configured to perform spatio-temporal alignment and interpolation on the preprocessed data to obtain unified input data;

[0176] A model construction unit 303, configured to construct a sandstorm forecasting model; wherein, the sandstorm forecasting model includes a spatio-temporal encoder, a feature aggregation layer, and a task-specific layer. Based on the unified input data, it processes it into an input tensor of dimension (B, C, T, H, W), where B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After passing through the sandstorm forecasting model, the input obtains the outputs of four tasks, namely two regression prediction tasks and two classification prediction tasks, and the output shape of each task is (B, H, W), representing the sandstorm distribution at the next moment;

[0177] A loss function determination unit 304, configured to construct a multi-task loss function for multiple sandstorm event forecasting tasks and train the sandstorm forecasting model based on the multi-task loss function;

[0178] A model training unit 305, configured to use the comprehensive loss as the optimization objective during the model training process, continuously adjust the model parameters through the backpropagation algorithm, gradually minimize the multi-task loss, and obtain the trained sandstorm forecasting model.

[0179] In some embodiments, the data preprocessing unit is further configured to:

[0180] Clean the data in the data set, use statistical methods to detect and remove data points with abnormal fluctuations, eliminate outliers and noise data, and ensure the accuracy and integrity of the data; for missing values caused by equipment failures or data transmission problems, use the cubic spline interpolation method for filling to ensure the integrity of the data;

[0181] Using a normalization method, the cleaned data is compressed into a preset numerical range to eliminate the dimensional differences between different data dimensions.

[0182] In some embodiments, the alignment interpolation unit is further configured to:

[0183] Project various types of data in the preprocessed data onto a unified spatial grid to ensure spatial alignment of all data. For PM10 concentration data with a resolution lower than the set threshold, bilinear interpolation or spline interpolation is used to adjust it to the same high-resolution grid as the satellite remote sensing data;

[0184] Unify all data into hourly time intervals and fill in the missing data at different times through an interpolation method;

[0185] For meteorological data and satellite remote sensing data, linear interpolation is used to smoothly fill the data with a shorter time interval to ensure the continuity and consistency of the time series;

[0186] For daily PM10 data with only daily intervals, hourly-granularity air quality station data is used to correct it to obtain hourly-granularity gridded PM10 data.

[0187] In some embodiments, the spatio-temporal encoder comprehensively extracts spatio-temporal features through a dual-channel architecture integrating 3D convolution and the Mamba module, and is coupled through the Hadamard product to effectively capture local details and global context; wherein, the dual-channel architecture includes a convolutional channel and a Mamba channel;

[0188] In the convolutional channel, the encoder uses two layers of three-dimensional convolution to extract local spatio-temporal features. Given a first input tensor x ∈ R B×C×T×H×W , where R represents the set of real numbers, the convolutional channel obtains a first output h c1 through the following convolution operation:

[0189] h c1 = Conv3D2(ReLU(Conv3D1(x)))

[0190] where both Conv3D1 and Conv3D2 represent three-dimensional convolution modules with a kernel size of 3×3×3, and ReLU is the rectified linear activation function;

[0191] In the Mamba channel, the first input tensor x ∈ R B×C×T×H×W is divided into non-overlapping patches of size P×P through a patch embedding-based method:

[0192] x′ = FC(Rearrange(x, P)),

[0193] Among them, FC represents the fully connected layer, P is the patch size, Rearrange represents the patch embedding operation, x' is the reshaped tensor, and x' ∈ R B×N×D , where is the number of embedded patches, and D = P 2 ·C is the embedding dimension;

[0194] After dividing non-overlapping patches, it is processed through multiple Vim blocks;

[0195] In each Vim block, the reshaped tensor x' generates the second output x'' by combining bidirectional sequence modeling and structured SSM; the reshaped tensor x' is first normalized and linearly projected into two feature representations, namely the first feature u and the second feature v. Subsequently, in the forward and backward directions respectively, a one-dimensional convolution operation is applied to the first feature u to generate the intermediate feature u' o , and the intermediate feature u' o is converted into a set of learnable parameters A o , B o , C o and Δ o , where Δ o ensures its positive value by applying the softplus activation function, and Δ o The specific calculation formula of is:

[0196]

[0197] Among them, Δ o is obtained by linearly transforming Δ o and adding a learnable bias parameter b, and then using the softplus activation function to ensure its value is positive, thereby adjusting the nonlinear transformation of the model;

[0198] The latent state h o is recursively updated through SSM:

[0199]

[0200] Among them, represents the matrix multiplication operation discretized by time step;

[0201] The output y in each direction o is calculated by the following formula:

[0202]

[0203] The forward output y forward and the backward output y backward are gated using v and combined as:

[0204] y combined = yforward ⊙SiLU(v) + y backward ⊙SiLU(v)

[0205] Where SiLU represents the Sigmoid linear unit activation function; ⊙ represents the Hadamard product operation;

[0206] Finally, perform a linear transformation on the merged result and add it to x′ through a residual connection to generate the second output x″:

[0207] x″ = Linear(y combined ) + x′

[0208] Where Linear represents the linear layer;

[0209] For the output y T ∈ R B×N×D at the last moment of the Mamba channel, use the reverse embedding operation to reconstruct the feature map to the initial data dimension to obtain the final output h c2 ∈ R B×C×T×H×W , which is matched with the data of the convolutional channel.

[0210] In some embodiments, the feature aggregation layer uses a GAM module to enhance the feature representation provided by the spatio-temporal encoder; wherein, the GAM module integrates channel attention and spatial attention for weighted feature mapping, dynamically emphasizes channel dependencies and spatial correlations, and suppresses irrelevant noise information;

[0211] For a given second input tensor h c ∈ R B×(C·T)×H×W , channel attention transforms it into a three-dimensional tensor h′ c ∈ R B×(H·W)×(C·T) , uses a two-layer fully connected network to identify the importance of features, and applies it to the original input through element-wise multiplication:

[0212] h CA = h c ·σ(FC2(ReLU(FC1(h c ′))))

[0213] Where FC1 and FC2 represent fully connected layers for adjusting the channel dimension, ReLU represents the rectified linear activation function, σ represents the sigmoid function; h CA represents the feature map;

[0214] Use the spatial attention module to capture the local spatial correlations in the feature map h CA and emphasize the key regions through convolution:

[0215] h SA = hCA · σ(Conv2(ReLU(BN(Conv1(h CA )))))

[0216] where Conv1 and Conv2 represent two layers of two-dimensional convolutions for adjusting the channel dimension, BN represents batch normalization, and h SA represents the feature map finally generated by the GAM module.

[0217] In some embodiments, the task-specific layer processes the aggregated features through a separate predictor or classifier to generate prediction results for multiple output tasks; for a regression task that directly outputs a numerical value, the predictor consists of two two-dimensional convolutional layers:

[0218]

[0219] where and represent task-specific convolutional layers; represents the model output result for a certain task;

[0220] The classifier adds a binary classification activation function on the basis of the predictor, expressed as:

[0221]

[0222] In some embodiments, the model training unit is configured to:

[0223] In each training cycle, the input data and the true label y i are used to calculate the prediction for the i-th task The regression task is calculated using the mean squared error loss:

[0224]

[0225] where L i represents the loss for the i-th task; represents the prediction result for the j-th sample of the i-th task; N represents the total number of label data samples; j represents the sample index;

[0226] The classification task is trained using binary cross-entropy loss:

[0227]

[0228] The total loss L for a batch is calculated by the following formula batch :

[0229]

[0230] where w iis the weight of the i-th task, and Task is the total number of tasks;

[0231] To dynamically balance tasks, w is dynamically adjusted during training according to the gradient norm i ; The gradient norm gi of the i-th task i is:

[0232]

[0233] where represents the gradient with respect to the parameter θ;

[0234] The average gradient norm for all tasks is Update the task weights through the following formula to minimize g i and the difference between:

[0235]

[0236] where η is the learning rate of the weights;

[0237] The weights are normalized to ensure that their sum is 1:

[0238]

[0239] where w j represents the j-th weight value;

[0240] In each training iteration, minimize the total loss L batch , update the model parameters through backpropagation, and ensure balanced optimization for all tasks.

[0241] In some embodiments, the device further includes a model evaluation unit, and the model evaluation unit is configured to:

[0242] Evaluate the trained dust forecast model using evaluation metrics;

[0243] For regression tasks (prediction of PM10 and BADI), use the mean squared error MSE, root mean squared error RMSE, mean absolute error MAE, and / or coefficient of determination R 2 to evaluate the prediction accuracy of the model;

[0244] For classification tasks, use accuracy, recall, precision, and / or F1 score for evaluation.

[0245] It should be noted that the structures of the spatio-temporal dust storm event AI forecasting devices that integrate satellite remote sensing and meteorological data described in this embodiment belong to the same technical concept as the spatio-temporal dust storm event AI forecasting method that integrates satellite remote sensing and meteorological data described previously, and achieve the same beneficial effects through the same principle, which will not be elaborated here.

[0246] An embodiment of the present invention also provides a readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.

[0247] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above specific implementation manners, various features can be grouped together to simplify the present invention. This should not be construed as an intention that the features of an invention not claimed are necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of a specific embodiment of the invention. Thus, the following claims are incorporated herein by way of example or embodiment into the specific implementation manners, where each claim independently serves as a separate embodiment, and considering these embodiments, they can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents to which these claims are entitled.

Claims

1. An AI forecasting method for spatiotemporal sandstorm events integrating satellite remote sensing and meteorological data, characterized in that: The method comprises: Acquire a data set, and preprocess the data set to obtain a preprocessed data set; wherein the data set includes satellite image data and meteorological reanalysis data, daily PM10 data, and air quality data; Performing spatiotemporal alignment and interpolation on the preprocessed data to obtain unified input data; Constructing a dust forecast model; wherein the dust forecast model includes a spatiotemporal encoder, a feature aggregation layer, and a task-specific layer. Based on unified input data, the data is processed into an input tensor of dimension (B, C, T, H, W), wherein B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After the input passes through the dust forecast model, the output of four tasks is obtained, namely, two regression prediction tasks and two classification prediction tasks. The output shape of each task is (B, H, W), which represents the dust distribution at the next moment; For various dust event forecasting tasks, construct a multi-task loss function, and train the dust forecasting model based on the multi-task loss function; During the model training process, the comprehensive loss is used as the optimization target. The model parameters are continuously adjusted through the back propagation algorithm to gradually minimize the multi-task loss and obtain the trained dust forecast model.

2. The AI ​​prediction method for spatiotemporal sandstorm events integrating satellite remote sensing and meteorological data according to claim 1 is characterized in that: Acquiring a data set and preprocessing the data set to obtain a preprocessed data set includes: Clean the data in the data set, use statistical methods to detect and remove data points with abnormal fluctuations, eliminate outliers and noise data, and ensure the accuracy and completeness of the data; for missing values ​​caused by equipment failure or data transmission problems, use cubic spline interpolation method to fill in the missing values ​​to ensure the integrity of the data; The normalization method is used to compress the cleaned data into a preset numerical range to eliminate the dimensional differences between different data dimensions.

3. The AI ​​prediction method for spatiotemporal sandstorm events integrating satellite remote sensing and meteorological data according to claim 1 is characterized in that: The preprocessed data is subjected to time-space alignment and interpolation to obtain unified input data, including: Project all types of data in the preprocessed data onto a unified spatial grid to ensure spatial alignment of all data. For PM10 concentration data with a resolution lower than a set threshold, use bilinear interpolation or spline interpolation to adjust it to the same high-resolution grid as the satellite remote sensing data. All data were unified into hourly time intervals, and missing data at different times were filled by interpolation methods; For meteorological data and satellite remote sensing data, linear interpolation is used to smoothly fill in data with shorter time intervals to ensure the continuity and consistency of the time series; For daily PM10 data with only daily intervals, the air quality station data with hourly granularity are used to correct them to obtain gridded PM10 data with hourly granularity.

4. The method for predicting spatiotemporal sandstorm events by integrating satellite remote sensing and meteorological data according to claim 1, characterized in that: The spatiotemporal encoder comprehensively extracts spatiotemporal features through a dual-channel architecture integrating 3D convolution and Mamba modules, and is coupled through Hadamard products to ensure that local details and global context are effectively captured; wherein the dual-channel architecture includes a convolution channel and a Mamba channel; In the convolution channel, the encoder uses two layers of 3D convolution to extract local spatiotemporal features. Given a first input tensor x∈R B×C×T×H×W , R represents a set of real numbers, and the first output h obtained by the convolution channel through the following convolution operation c1 : h c1 =Conv3D2(ReLU(Conv3D1(x))) Among them, Conv3D1 and Conv3D2 both represent three-dimensional convolution modules with a convolution kernel size of 3×3×3, and ReLU is a linear rectification activation function; In the Mamba channel, the first input tensor x∈R is transformed into B×C×T×H×W Partition into non-overlapping patches of size P×P: x′=FC(Rearrange(x,P)), Where FC represents the fully connected layer, P is the patch size, Rearrange represents the patch embedding operation, x′ is the reshaped tensor, x′∈R B×N×D ,in is the number of embedded patches, D = P 2 C is the embedding dimension; After dividing into non-overlapping patches, they are processed through multiple layers of Vim blocks; In each Vim block, the reshaped tensor x′ generates the second output x″ by combining bidirectional sequence modeling and structured SSM; the reshaped tensor x′ is first normalized and linearly projected into two feature representations, namely the first feature u and the second feature v, and then a one-dimensional convolution operation is applied to the first feature u in the forward and backward directions to generate the intermediate feature u′ o , intermediate feature u′ o is converted into a set of learnable parameters A o , B o , C o and Δ o , where Δ o By applying the softplus activation function, Δ o The specific calculation formula is: Among them, Δ o is through the o After performing the linear transformation and adding the learnable bias parameter b, the softplus activation function is used to ensure that its value is positive, thereby adjusting the nonlinear transformation of the model; Potential state h o Recursive update via SSM: in, represents a matrix multiplication operation discretized by time steps; The output y for each direction o Calculated by the following formula: Forward output y forward and the backward output y backward Use v for gating and combine as: and combined =and forward ⊙SiLU(v)+y backward ⊙SiLU(v) Among them, SiLU represents the Sigmoid linear unit activation function; ⊙ represents the Hadamard product operation; Finally, the combined result is linearly transformed and added to x′ through a residual connection to generate the second output x″: x″=Linear(y combined )+x′ Among them, Linear represents the linear layer; Output y of the Mamba channel at the last moment T ∈R B×N×D , the feature map is reconstructed to the initial data dimension using the reverse embedding operation to obtain the final output h c2 ∈R B×C×T×H×W , to match the data of the convolution channel.

5. The method for predicting spatiotemporal sandstorm events by integrating satellite remote sensing and meteorological data according to claim 1 is characterized in that: The feature aggregation layer uses a GAM module to enhance the feature representation provided by the spatiotemporal encoder; wherein the GAM module integrates channel attention and spatial attention for weighted feature mapping, dynamically emphasizes channel dependency and spatial correlation, and suppresses irrelevant noise information; For a given second input tensor h c ∈R B×(C·T)×H×W , channel attention transforms it into a three-dimensional tensor h′ c ∈R B ×(H·W)×(C·T) , using a two-layer fully connected network to identify the importance of features and apply them to the original input via element-wise multiplication: h CA =h c ·σ(FC2(ReLU(FC1(h c ′)))) Among them, FC1 and FC2 represent fully connected layers, which are used to adjust the channel dimension, ReLU represents the linear rectification activation function, and σ represents the sigmoid function; h CA represents a feature map; The spatial attention module is used to capture the feature map h CA Local spatial correlation in , using convolution to emphasize key areas: h SA =h CA ·σ(Conv2(ReLU(BN(Conv1(h CA ))))) Conv1 and Conv2 represent two layers of two-dimensional convolution, which are used to adjust the channel dimension. BN represents batch normalization. SA Represents the feature map finally generated by the GAM module.

6. The method for predicting spatiotemporal sandstorm events by integrating satellite remote sensing and meteorological data according to claim 5 is characterized in that: The task-specific layer processes the aggregated features through a separate predictor or classifier to generate prediction results for various output tasks; for regression tasks that directly output numerical values, the predictor consists of two two-dimensional convolutional layers: in and represents a task-specific convolutional layer; Represents the model output result of a task; The classifier adds a binary activation function based on the predictor, expressed as:

7. The method for predicting spatiotemporal sandstorm events by integrating satellite remote sensing and meteorological data according to claim 1, characterized in that: For various dust event forecasting tasks, a multi-task loss function is constructed, and the dust forecasting model is trained based on the multi-task loss function, including: In each training cycle, the input data and the true label y i Used to calculate the prediction for the i-th task The regression task is calculated using the mean squared error loss: Among them, L i represents the loss of the i-th task; represents the prediction result of the jth sample of the i-th task; N represents the total number of label data samples; j represents the sample index; The classification task is trained using binary cross entropy loss: The total loss L of a batch is calculated by the following formula batch : where w i is the weight of the i-th task, Task is the total number of tasks; In order to dynamically balance the tasks, w is dynamically adjusted during training according to the gradient norm. i ; Gradient norm g of the i-th task i for: in, represents the gradient of the parameter θ; The average gradient norm for all tasks is Update the task weights to minimize g by the following formula i and The difference between: Where η is the learning rate of the weights; The weights are normalized to ensure that they sum to 1: where w j represents the jth weight value; In each training iteration, the total loss L is minimized batch , model parameters are updated through back-propagation to ensure balanced optimization of all tasks.

8. The method for predicting spatiotemporal sandstorm events by integrating satellite remote sensing and meteorological data according to claim 1, characterized in that: After obtaining the trained dust forecast model, the method further includes: The trained dust forecast model is evaluated using evaluation indicators; For regression tasks (prediction of PM10 and BADI), mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE) and / or coefficient of determination (R) are used. 2 To evaluate the prediction accuracy of the model; For classification tasks, accuracy, recall, precision and / or F1 score are used for evaluation.

9. An AI forecasting device for spatiotemporal sandstorm events integrating satellite remote sensing and meteorological data, characterized in that: The device comprises: A data preprocessing unit is configured to obtain a data set and preprocess the data set to obtain a preprocessed data set; wherein the data set includes satellite image data and meteorological reanalysis data, daily PM10 data and air quality data; an alignment and interpolation unit, configured to perform spatiotemporal alignment and interpolation on the preprocessed data to obtain unified input data; A model building unit is configured to build a sandstorm forecast model; wherein the sandstorm forecast model includes a spatiotemporal encoder, a feature aggregation layer, and a task-specific layer, and based on unified input data, processes it into an input tensor with a dimension of (B, C, T, H, W), wherein B is the batch size, C is the number of input channels, T is the historical time length, and H×W is the spatial resolution. After the input passes through the sandstorm forecast model, the output of four tasks is obtained, which are two regression prediction tasks and two classification prediction tasks, respectively. The output shape of each task is (B, H, W), which represents the sandstorm distribution at the next moment; A loss function determination unit is configured to construct a multi-task loss function for a plurality of dust event forecasting tasks, and train the dust forecasting model based on the multi-task loss function; The model training unit is configured to use the comprehensive loss as the optimization target during the model training process, continuously adjust the model parameters through the back propagation algorithm, gradually minimize the multi-task loss, and obtain the trained dust forecast model. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .

Citation Information

Patent Citations

  • Multi-feature fusion sand storm prediction method based on deep neural network

    CN114882373A

  • Weather forecasting method and system based on artificial intelligence

    CN117633473A

  • Strong supervision change detection method based on convolutional neural network and visual attention model

    CN119206487A

  • Hyperspectral remote sensing image classification method and device, and storage medium

    CN119313952A

  • Remote sensing image crop classification method based on Mama

    CN119418141A

Cited By

  • Sand and dust early warning method, device and system based on satellite remote sensing

    CN121580330A

  • Satellite remote sensing-based sand-dust early warning method, device and system

    CN121580330B

  • AI-based regional weather abnormity intelligent detection and judgment method and system

    CN121682197A