Sand storm prediction method based on multi-source data fusion and conditional diffusion generation
By adopting multi-source data fusion and conditional diffusion generation methods in sandstorm prediction, the problems of incomplete multi-source data fusion in the existing technology and unused deep learning technology are solved, and accurate prediction of sandstorms is achieved.
Patent Information
- Application Number
- CN202510165491.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The existing technology relies on traditional machine learning technology in sandstorm prediction, lacks the application of deep learning technology, and lacks exploration of methods for multi-source data fusion.
The sandstorm prediction method based on multi-source data fusion and conditional diffusion is adopted to clearly define the sandstorm prediction task, and multi-source data such as meteorological data and satellite remote sensing products are fused to train the multi-source data noise prediction network to achieve the training and application of the conditional diffusion model.
Accurate prediction of sandstorms is achieved, the accuracy and reliability of predictions are improved, and the problems of incomplete multi-source data fusion in traditional methods and the unused deep learning technology are solved.
Smart Images

Figure CN120198792A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sandstorm forecasting, and specifically relates to a sandstorm prediction method based on multi-source data fusion and conditional diffusion generation. Background Art
[0002] As one of the natural disasters in arid and semi-arid regions, sandstorms are characterized by rapid occurrence, high intensity, and great impact. The occurrence of sandstorms requires three basic conditions, including strong winds, drought, and sand sources. Sandstorm prediction is to predict the occurrence of sandstorms at future times using historical meteorological data including satellite data, etc. China is one of the countries most severely affected by sandstorms globally. Large-scale sandstorms cause huge economic losses to the northern regions of China every year. To prevent such weather disasters, it is very necessary to predict the occurrence range and intensity of sandstorms.
[0003] Artificial intelligence technology is still in its infancy in the field of sandstorm prediction. Although many studies have explored the application capabilities of artificial intelligence and deep learning technologies in the field of sandstorm prediction from multiple perspectives and achieved optimistic results, there are still several problems to be solved: (1) There are differences in data input, output, etc. among existing studies, and at the same time, there is a lack of a standardized sandstorm dataset and a clear definition of the sandstorm prediction problem. (2) Most studies still rely on traditional machine learning technologies and have not incorporated emerging deep learning technologies into sandstorm prediction research. (3) Existing studies do not consider the influencing factors of sandstorms comprehensively enough, fail to make full use of multi-source observation data such as meteorological data and satellite remote sensing products, and lack exploration of multi-source data fusion methods. Summary of the Invention
[0004] In order to overcome the deficiencies of the existing technology, the present invention provides a sandstorm prediction method based on multi-source data fusion and conditional diffusion generation. By means of multi-source data fusion and conditional diffusion, on the basis of clearly defining the sandstorm prediction task and data, it solves the problems that the existing technical solutions rely on traditional machine learning technologies and lack exploration of multi-source data.
[0005] The technical solution adopted by the present invention to solve its technical problems is:
[0006] A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation, comprising the following steps:
[0007] S1. Clearly define the sandstorm prediction task as a spatio-temporal sequence prediction problem, with the input being meteorological data at multiple times and the output being sandstorm change information at multiple times;
[0008] S2. Add noise to the sandstorm data according to the noise addition formula to obtain sandstorm data with noise. Train a multi-source data noise prediction network through satellite cloud images, meteorological reanalysis data, sandstorm data at historical moments, and the time step t, and then complete the training of the conditional diffusion model.
[0009] S3. Input the complete Gaussian noise map, the current time step t x , as well as satellite cloud images, meteorological reanalysis data, and sandstorm data at historical moments into the multi-source data noise prediction network to predict the noise of the current noise map. Remove the noise at the current moment through the denoising formula to obtain the Gaussian noise map at time t x-1 , that is, the sandstorm data with noise at time t x-1 . Repeat this operation until time t0 to obtain the sandstorm data with denoising completed, that is, the prediction result.
[0010] Furthermore, in S1, the spatio-temporal sequence prediction problem includes the following parts:
[0011] The input data is meteorological data collected at the previous N moments, including satellite cloud image data and meteorological reanalysis data; the output data, that is, the target is the change of sandstorms within the next M moments, including the change of the occurrence range and intensity of sandstorms; the prediction model adopted needs to accept the input of data at multiple moments and has the output of data at multiple moments.
[0012] Preferably, the meteorological reanalysis data is subdivided into (1) atmospheric motion data and (2) temperature and pressure data. The atmospheric motion data includes data such as wind speed at two meters, wind direction at two meters, wind speed at ten meters, and wind direction at ten meters; the temperature and pressure data includes data such as air temperature, dew point temperature, surface pressure, and mean sea level pressure.
[0013] Still further, in S2, the multi-source data noise prediction network integrates satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data, and promotes the prediction of noise through the method of multi-source data fusion.
[0014] First, in the downsampling stage, satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data have their own processing channels. Extract the features of various types of data through the encoder, time step embedding module, residual connection module, and downsampling operation. Each type of data extracts at least four scales of features; then, the meteorological data fusion module connects various types of data and extracts the fused data features through three-dimensional convolution, group normalization, and SILU activation function, extracting at least four scales of fused data features; immediately afterwards, in the upsampling stage, restore the feature size through the time step embedding module, residual connection module, and upsampling operation; finally, output the noise prediction result through the prediction head, that is, the fully connected layer.
[0015] Preferably, the downsampling stage is divided into four different processing channels, which separately process satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data;
[0016] First, there are certain differences in the number of data channels among satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data. When encoding, it is necessary to align the feature sizes before performing subsequent downsampling operations. Second, the goal of conditional diffusion generation is the predicted sandstorm data, which needs to be iterated from a specific time step t x to t0. Therefore, when extracting the features of sandstorm data at each scale, it is necessary to encode the time information and perform time step embedding; while satellite cloud images, atmospheric motion data, and temperature and pressure data assist the diffusion process, so when extracting the features at each scale, time step embedding operations are not required.
[0017] More preferably, the time step embedding module can embed the time step information into the sandstorm data information and use the time information to guide the implementation of conditional diffusion generation;
[0018] First, the time step information is encoded through the SILU activation function and the fully connected layer. The sandstorm data at time step t, including historical and predicted sandstorm data, is subjected to feature extraction through group normalization, the SILU activation function, and three-dimensional convolution, and the two are added bitwise. Then, the added result is further subjected to feature extraction through group normalization, the SILU activation function, the Dropout layer, and three-dimensional convolution; finally, the extracted features are subjected to residual connection with the initial sandstorm data to obtain the final time step embedding features.
[0019] Furthermore, the encoder performs the following operations:
[0020] The input data sequentially passes through the input layer and the output layer. The input layer is initially processed through group normalization, the SILU activation function, and three-dimensional convolution; the output layer obtains the encoding result through group normalization, the SILU activation function, the Dropout layer, and three-dimensional convolution to achieve data size alignment.
[0021] Furthermore, the residual connection module performs the following operations:
[0022] The input data extracts features through group normalization, the SILU activation function, the Dropout layer, and three-dimensional convolution, and the extracted features are added bitwise to the initial output data to obtain the residual connection result.
[0023] Furthermore, the downsampling operation performs the following operations:
[0024] The input features are first reduced in size through max pooling, then the data is normalized through group normalization, and finally features are extracted through two-dimensional convolution.
[0025] Furthermore, the meteorological data fusion module can fuse and extract multi-source data features, extracting at least four-dimensional multi-source data features;
[0026] First, the multi-source data features of the same scale extracted in the downsampling stage are concatenated. Then, three-dimensional convolution is used for feature extraction, and the extracted features are subjected to group normalization and SILU activation; finally, three-dimensional convolution, group normalization, and SILU activation are repeated once to obtain the fused features at this scale.
[0027] Furthermore, the upsampling stage can perform upsampling operations on the fused features extracted by the meteorological data fusion module to gradually restore the feature size.
[0028] For the fused features of each scale, the feature size is enlarged through the time step embedding module, residual module, bilinear interpolation, and two-dimensional convolution, and then concatenated with the fused features of the previous scale; continue to execute the time step embedding module, residual module, bilinear interpolation, and feature concatenation; repeat the operation until all the fused data features are processed to achieve the restoration of the feature size.
[0029] In the present invention, the data required for forecasting includes satellite data and meteorological reanalysis data related to sandstorms. The satellite data includes satellite band data such as R0.47 and R0.65 of Fengyun 4A; the meteorological reanalysis data includes two-meter height wind speed, two-meter height wind direction, etc. in ERA5 reanalysis data and MERRA-2 reanalysis data. This application clearly defines the sandstorm prediction task as a spatio-temporal sequence prediction problem, with the meteorological data collected in the previous N moments defined as the input and the changes in sandstorms within the next M moments as the output. A conditional diffusion generation model is used to gradually denoise for sandstorm prediction, and a multi-source data fusion method is used to predict the noise in the diffusion process, and finally the change information of sandstorms at future moments is obtained. This application can obtain valuable information from complex meteorological data and can accurately predict sandstorms.
[0030] The beneficial effects of the present invention are mainly manifested in: being able to accurately predict sandstorms. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart of a sandstorm prediction method based on multi-source data fusion and conditional diffusion generation.
[0032] Figure 2 It is the input and output of the sandstorm prediction task.
[0033] Figure 3It is a structural schematic diagram of a conditional diffusion model.
[0034] Figure 4 It is a structural schematic diagram of a multi-source data noise prediction network.
[0035] Figure 5 It is a structural schematic diagram of a time step embedding module. Detailed implementation manners
[0036] The present invention will be further described below with reference to the accompanying drawings.
[0037] Refer to Figures 1 to 5 , a sandstorm prediction method based on multi-source data fusion and conditional diffusion generation, trains and constructs a prediction network model on the basis of clarifying the sandstorm prediction task, and then processes multi-source data by using the trained prediction network model to obtain the change information of the sandstorm within the future time.
[0038] The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation includes the following steps:
[0039] Step S1: Clearly define the sandstorm prediction task as a spatio-temporal sequence prediction problem, with the input being meteorological data at multiple moments and the output being the change information of the sandstorm at multiple moments;
[0040] Specifically, as Figure 2 shown, the data input in sandstorm prediction is meteorological data at the historical N moments, including satellite cloud images, meteorological reanalysis data, sandstorm data, etc.; the output data is the change information of the sandstorm within the future M moments.
[0041] In a specific embodiment, for the input data, the size of the satellite cloud image is 4×5×160×480×7, the size of the meteorological reanalysis data is 4×5×160×480×12, and the size of the sandstorm data is 4×5×160×480×1. Among them, 4 represents the batch size, 5 represents the number of historical moments and future moments, the time interval is 15 minutes, 160 and 480 represent the spatial dimensions of each channel, and 7, 12, and 1 respectively represent the number of channels of their respective data.
[0042] It should be noted that the present application defines the sandstorm prediction as a spatio-temporal prediction problem, and for spatio-temporal prediction problems, different models can be used to solve them. The present application uses an advanced conditional diffusion model to realize sandstorm prediction.
[0043] Step S2: Add noise to the sandstorm data according to the noise addition formula to obtain the sandstorm data with noise; jointly train the multi-source data noise prediction network through the satellite cloud images, meteorological reanalysis data, sandstorm data at the historical moments, and the time step t, and then complete the training of the conditional diffusion model;
[0044] In this embodiment, noise is added to the sandstorm data, and a multi-source data noise prediction network is jointly trained with the time step and multi-source data at historical moments to complete the training of the conditional diffusion model.
[0045] In a specific embodiment, as Figure 3 shown, this embodiment needs to train a noise prediction network for subsequent generation of sandstorm data. First, noise is added to the sandstorm data, and the added noise is determined by the time step t to obtain sandstorm data with noise. Then, the sandstorm data with noise, the time step t, and multi-source data at historical moments are input into the multi-source data noise prediction network to minimize the gap between the noise predicted by the network and the added noise. Among them, the multi-source data at historical moments serves as a condition to promote the prediction of the noise network. The multi-source data includes satellite cloud images, meteorological reanalysis data, and sandstorm data. The satellite cloud images mainly include bands such as R0.47 and R0.65. The meteorological reanalysis data is subdivided into (1) atmospheric motion data and (2) temperature and pressure data. The atmospheric motion data mainly includes data such as wind speed at a height of two meters, wind direction at a height of two meters, wind speed at a height of ten meters, and wind direction at a height of ten meters. The temperature and pressure data mainly includes data such as air temperature, dew point temperature, surface pressure, and mean sea level pressure. The sandstorm data mainly includes the occurrence location and intensity information of sandstorms.
[0046] Specifically, the multi-source data noise prediction network is as Figure 4 shown. In the downsampling stage, the multi-source data undergoes corresponding encoders for data size alignment, and then passes through the time step embedding module, residual connection module, and downsampling operation to extract features of various types of data, extracting at least four scales of features. The sandstorm data is constrained by the time step t in the noise addition stage, so the time step embedding module is required to encode the time information, and then extract features through the residual module and downsampling operation. The time step embedding module is as Figure 5 shown. The time step information is encoded through the SILU activation and fully connected layers, and added bitwise to the features of the sandstorm data extracted through group normalization, SILU activation, and three-dimensional convolution. Then, it further extracts features through group normalization, the SILU activation function, the Dropout layer, and three-dimensional convolution. Finally, a residual link is made with the initial sandstorm data to obtain the final time step embedding features. The satellite cloud images, atmospheric motion data, and temperature and pressure data, as conditions, do not require time information encoding and can directly extract features through the residual module and downsampling operation.
[0047] In a specific embodiment, the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the satellite cloud image encoder are 7, 32, 3, 1, 1 and 32, 32, 3, 1, 1 respectively; the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the atmospheric motion data encoder are 6, 32, 3, 1, 1 and 32, 32, 3, 1, 1 respectively; the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the temperature and air pressure data encoder are 6, 32, 3, 1, 1 and 32, 32, 3, 1, 1 respectively; the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the sandstorm data encoder are 1, 32, 3, 1, 1 and 32, 32, 3, 1, 1 respectively. Therefore, after passing through the encoder, the dimensions of various types of data are all expanded to 32 channels, thereby achieving the alignment of data sizes.
[0048] The data sizes of satellite cloud images, atmospheric motion data, temperature and pressure data are the same after passing through the encoder, and the subsequent processing processes are similar. Taking satellite cloud images as an example, when extracting features of the first size, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the three-dimensional convolution in the residual module are 32, 32, 3, 1, 1; when extracting features of the second size, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the two-dimensional convolution in the downsampling operation and the three-dimensional convolution in the residual module are 32, 64, 3, 1, 1 and 64, 64, 3, 1, 1 respectively; when extracting features of the third size, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the two-dimensional convolution in the downsampling operation and the three-dimensional convolution in the residual module are 64, 128, 3, 1, 1 and 128, 128, 3, 1, 1 respectively; when extracting features of the fourth size, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the two-dimensional convolution in the downsampling operation and the three-dimensional convolution in the residual module are 128, 256, 3, 1, 1 and 256, 256, 3, 1, 1 respectively. For sandstorm data, when extracting features of the first scale, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the three-dimensional convolution in the time step embedding module are 32, 32, 3, 1, 1; when extracting features of the second scale, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the three-dimensional convolution in the time step embedding module are 64, 64, 3, 1, 1; when extracting features of the third scale, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the three-dimensional convolution in the time step embedding module are 128, 128, 3, 1, 1; when extracting features of the fourth scale, the number of input channels, output channels, convolution kernel size, convolution stride and padding width of the three-dimensional convolution in the time step embedding module are 256, 256, 3, 1, 1. The pooling operation used is max pooling, and the convolution kernel size, convolution stride and padding width are 2, 2, 0.
[0049] In this embodiment, through the encoder, downsampling operation, time step embedding module, and residual module, data features of one scale can be obtained. When data features of multiple scales are required, multiple feature extraction units are needed, which will not be elaborated here.
[0050] In the data fusion stage, the meteorological data fusion module will connect various data features at the same scale, and extract fused data features through three-dimensional convolution, group normalization, and SILU activation function, extracting at least four scales of fused features.
[0051] In a specific embodiment, the number of channels after connecting the data features of the first scale is 128, and the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution used are 128, 32, 3, 1, 1 and 32, 32, 3, 1, 1; the number of channels after connecting the data features of the second scale is 256, and the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution used are 256, 64, 3, 1, 1 and 64, 64, 3, 1, 1; the number of channels after connecting the data features of the third scale is 512, the number of channels after connecting the data features of the fourth scale is 1024, and the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution used are 512, 128, 3, 1, 1 and 128, 128, 3, 1, 1; the number of channels after connecting the data features of the fourth scale is 1024, and the number of input channels, output channels, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution used are 1024, 512, 3, 1, 1 and 512, 512, 3, 1, 1;
[0052] In the upsampling stage, starting from the fused feature with the smallest data size, the data size is restored through the time step embedding module, the residual connection module, and the upsampling operation, and is connected to the fused feature of the previous scale, and the operation is repeated until the data size is restored. Among them, the upsampling operation uses bilinear interpolation to enlarge the data size and uses two-dimensional convolution to enhance the expression ability of the feature map.
[0053] In a specific embodiment, for the features of the fourth scale, the input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the upsampling time step embedding module and the residual connection module are 512, 512, 3, 1, 1 and 512, 256, 3, 1, 1 respectively. The input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the two-dimensional convolution in the upsampling operation are 256, 128, 3, 1, 1; for the features of the third scale, the input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the upsampling time step embedding module and the residual connection module are 256, 256, 3, 1, 1 and 256, 128, 3, 1, 1 respectively. The input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the two-dimensional convolution in the upsampling operation are 128, 64, 3, 1, 1; for the features of the second scale, the input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the upsampling time step embedding module and the residual connection module are 128, 128, 3, 1, 1 and 128, 64, 3, 1, 1 respectively. The input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the two-dimensional convolution in the upsampling operation are 64, 32, 3, 1, 1; for the features of the first scale, the input channel number, output channel number, convolution kernel size, convolution stride, and padding width of the three-dimensional convolution in the upsampling time step embedding module and the residual connection module are 64, 64, 3, 1, 1 and 64, 32, 3, 1, 1 respectively.
[0054] In the prediction stage, the features are output through the prediction head, i.e., the fully connected layer, to obtain the noise prediction result.
[0055] In a specific embodiment, there are two fully connected layers in total. For the first fully connected layer, i.e., the time fully connected layer, the input dimension is 10 and the output dimension is 5; for the second fully connected layer, i.e., the channel fully connected layer, the input dimension is 32 and the output dimension is 1.
[0056] It should be noted that the goal of the conditional diffusion model generation model in the training stage is to train the noise prediction network. The multi-source data noise prediction network designed in this application uses satellite cloud images and meteorological reanalysis data as conditions to achieve accurate noise prediction.
[0057] Step S3: Input the complete Gaussian noise map, the current time step t x , as well as the satellite cloud images, meteorological reanalysis data, and sandstorm data at historical times into the multi-source data noise prediction network to predict the noise of the current noise map, and remove the noise at the current moment through the denoising formula to obtain the Gaussian noise map at time t x-1 , that is, at time t x-1Sandstorm data with noise at all times. Repeat this operation until time t0 to obtain the denoised sandstorm data, which is the prediction result.
[0058] In this embodiment, by applying the trained multi-source data noise prediction network and using the time step information and multi-source data at historical times, the completely Gaussian noise picture is gradually denoised to obtain clear sandstorm data, which is the final sandstorm prediction result.
[0059] In a specific embodiment, as Figure 3 shown, after the conditional diffusion model, i.e., the multi-source data noise prediction network, is trained, it is necessary to use the conditional diffusion model for sandstorm prediction. In a specific embodiment, the maximum time step is 1000. In the prediction stage, starting from the completely Gaussian noise map at time t 1000 input the completely Gaussian noise map, time step t 1000 and the multi-source data at historical times into the multi-source noise prediction network to obtain the noise prediction result at time t 1000 , and then obtain the sandstorm data with Gaussian noise at time t 999 according to the denoising formula. Then input the sandstorm data with Gaussian noise at time t 999 , time step t 999 and the multi-source data at historical times into the multi-source noise prediction network to obtain the noise prediction result at time t 999 , and then obtain the sandstorm data with Gaussian noise at time t 998 according to the denoising formula. Repeatedly apply the noise prediction network and the denoising formula until the sandstorm data at time t0 is obtained, which is the final prediction result.
[0060] For the conditional diffusion generation model in this application, its training process is as follows:
[0061] First, perform data set preprocessing. Divide the data set into three parts: training set, validation set, and test set according to the ratio of 7:1:2. The training set contains satellite cloud images, meteorological reanalysis data, and sandstorm data with longitude and latitude ranges of 32.4N - 52.7N, 64.3E - 129.2E from March to May in 2020 - 2021; the validation set and test set respectively contain satellite cloud images, meteorological reanalysis data, and sandstorm data in March 2022 and from April to May 2022; normalize the satellite cloud images and meteorological reanalysis data to make the data more conducive to model training; for model adaptation, all data sizes are 160×480.
[0062] During training:
[0063] The convolutional layers and other parameters in the detection network model were initialized with random weights, and the model was optimized using the model optimizer Adam at a learning rate of 0.0002 and trained for 200 cycles.
[0064] The sequence data in the training set is input into the model, and the sandstorm data is added with noise using the randomly generated time step t. Then the sandstorm data with added noise, time step t, and multi-source data at historical moments are input into the multi-source data prediction network to obtain the noise prediction result. The network prediction result is compared with the added noise to obtain the loss L. The parameters of the noise prediction network are adjusted by minimizing the loss function through AdamW to optimize the noise prediction performance.
[0065] The multi-source data noise prediction network model is trained until the loss converges, and the best performance multi-source data noise prediction network is obtained for noise prediction.
[0066] This application also provides experimental data under the data set trained by the conditional generative model, and compares the technical solution of this application with other deep learning prediction methods (U-net, 3D U-net, ConvLSTM, Earthfomer). The performance indicators used for comparison are Dice coefficient, key success index SCI, Heideck skill score HSS, Kappa coefficient and bias score Bias. Among them, the Bias index closer to 1 indicates better performance, and higher Dice, SCI, HSS and Kappa indicate better performance. The comparison results are shown in Table 1:
[0067] Methods Dice CSI HSS Kappa Bias U-net 0.2659 0.1720 0.2623 0.4012 10.2083 3D U-net 0.4371 0.3107 0.4335 0.4800 1.7637 ConvLSTM 0.4518 0.3191 0.4485 0.4999 2.0152 Earthfomer 0.4481 0.3253 0.4450 0.5817 7.5836 The method of the present invention 0.5017 0.3726 0.4971 0.5041 0.6307
[0068] Table 1
[0069] Table 1 above is the comparison results of different deep learning models in the sandstorm prediction task. The method of the present application is superior to all previous methods in Dice, CSI, HSS, and Bias, and is second only to the Earthfomer model in Kappa. It can be seen that the performance of the method of the present invention is better than other deep learning methods.
[0070] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation, characterized in that: The method comprises the following steps: S1. The sandstorm prediction task is clearly defined as a spatiotemporal series prediction problem, with the input being meteorological data at multiple moments and the output being sandstorm change information at multiple moments; S2, adding noise to the sandstorm data according to the noise addition formula to obtain sandstorm data with noise; training a multi-source data noise prediction network through satellite cloud images, meteorological reanalysis data, sandstorm data at historical moments, and time step t, thereby completing the training of the conditional diffusion model; S3, the complete Gaussian noise map, the current time step t x , and the satellite cloud images, meteorological reanalysis data, and sandstorm data at historical moments are input into the multi-source data noise prediction network to predict the noise of the current noise image. The noise at the current moment is removed by the denoising formula to obtain t x-1 Gaussian noise image at time t x-1 The sandstorm data with noise at all times; repeat this operation until time t0, and obtain the denoised sandstorm data, that is, the prediction result.
2. A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 1, characterized in that: In S1 of the method, the spatiotemporal sequence prediction problem includes the following parts: The input data is the meteorological data collected in the previous N moments, including satellite cloud image data and meteorological reanalysis data; The output data, that is, the target is the change of sandstorms in the future M moments, including the change of the scope and intensity of sandstorms; the prediction model used needs to accept data input at multiple moments and have data output at multiple moments.
3. A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 2, characterized in that: The meteorological reanalysis data are subdivided into atmospheric movement data and temperature and air pressure data. The atmospheric movement data include wind speed at two meters height, wind direction at two meters height, wind speed at ten meters height and wind direction at ten meters height; the temperature and air pressure data include air temperature, dew point temperature, surface air pressure and mean sea level air pressure data.
4. A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in any one of claims 1 to 3, characterized in that: In S2, the multi-source data noise prediction network integrates satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data, and promotes noise prediction by integrating multi-source data; First, in the downsampling stage, satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data have their own processing channels. The features of each type of data are extracted through encoders, time step embedding modules, residual link modules, and downsampling operations. At least four scales of features are extracted for each type of data. Then, the meteorological data fusion module connects various types of data, extracts fused data features through three-dimensional convolution, group normalization, and SILU activation function, and extracts fused data features of at least four scales; followed by the upsampling stage through the time step embedding module and the residual link module, and the upsampling operation restores the feature size; finally, the noise prediction result is output through the prediction head, that is, the fully connected layer.
5. The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 4, characterized in that: The downsampling stage is divided into four different processing channels, which process satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data respectively; First, there are certain differences in the number of data channels between satellite cloud images, atmospheric motion data, temperature and pressure data, and sandstorm data. When encoding, it is necessary to align the feature sizes and then perform subsequent downsampling operations. Second, the goal of conditional diffusion generation is to predict sandstorm data, which needs to be generated from a specific time step t. x Iterate to t0. Therefore, when extracting the features of sandstorm data at each scale, it is necessary to encode the time information and perform time step embedding. Satellite cloud images, atmospheric motion data, temperature and pressure data serve as auxiliary conditions to help the diffusion process. Therefore, there is no need to perform time step embedding when extracting features at each scale.
6. A sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 4, characterized in that: The time step embedding module can embed the time step information into the sandstorm data information, and use the time information to guide the realization of conditional diffusion generation; First, the time step information is encoded through the SILU activation function and the fully connected layer, and the sandstorm data at time step t, including historical and predicted sandstorm data, are extracted through group normalization, SILU activation function, and three-dimensional convolution, and the two are added bit by bit; then, the added result is further extracted through group normalization, SILU activation function, Dropout layer, and three-dimensional convolution; finally, the extracted features are residually linked with the initial sandstorm data to obtain the final time step embedding features.
7. The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 4, characterized in that: The encoder performs the following operations: The input data passes through the input layer and the output layer in turn. The input layer is initially processed through group normalization, SILU activation function, and three-dimensional convolution. The output layer obtains the encoding result through group normalization, SILU activation function, Dropout layer, and three-dimensional convolution to achieve data size alignment.
8. The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 4, characterized in that: The residual connection module performs the following operations: The input data is subjected to group normalization, SILU activation function, Dropout layer, and 3D convolution to extract features. The extracted features are bitwise added to the initial output data to obtain the residual connection result. The downsampling operation performs the following operations: The input features are first reduced in size through max pooling, then normalized through group normalization, and finally features are extracted through 2D convolution.
9. The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation as claimed in claim 4, characterized in that: The meteorological data fusion module is capable of fusing and extracting multi-source data features, and extracting multi-source data features of at least four dimensions; First, the multi-source data features of the same scale extracted in the downsampling stage are connected. Then, three-dimensional convolution is used for feature extraction, and the extracted features are group normalized and SILU activated. Finally, three-dimensional convolution, group normalization, and SILU activation are repeated to obtain the fused features of this scale.
10. The sandstorm prediction method based on multi-source data fusion and conditional diffusion generation according to claim 4, characterized in that: The upsampling stage can perform an upsampling operation on the fusion features extracted by the meteorological data fusion module, and gradually restore the feature size; For each scale of fused features, the feature size is enlarged through the time step embedding module, residual module, bilinear interpolation, and two-dimensional convolution, and then connected with the fused features of the previous scale; continue to execute the time step embedding module, residual module, bilinear interpolation and feature connection; repeat the operation until all fused data features are processed to achieve feature size recovery.
Citation Information
Patent Citations
Multi-feature fusion sand storm prediction method based on deep neural network
CN114882373A
Short-term photovoltaic power prediction method based on three-dimensional meteorological data multi-source fusion
CN117424232A
Remote sensing image fusion method based on multi-scale conditional diffusion model
CN117952843A
Sand storm prediction method based on IDSCNN-AM-LSTM
CN118278453A
Multi-modal sample generation method based on geographic information conditional diffusion model
CN119180998A
Cited By
Renewable energy extreme drought monthly scene joint generation method
CN121682466A