Multi-source data fusion image sequence generation method and system based on space-time diffusion

Through a multi-source data fusion method based on spatiotemporal diffusion, satellite remote sensing and radar reflectivity image training networks are used, combined with a spatiotemporal diffusion model, to directly learn the spatiotemporal joint distribution of lightning events, solving the problem of insufficient prediction accuracy in existing technologies and achieving lightning forecasts with higher accuracy and robustness.

CN120807279AActive Publication Date: 2025-10-17NANJING UNIV OF INFORMATION SCI & TECH +1

Patent Information

Application Number
CN202511303485.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing spatiotemporal point process models are unable to effectively capture the complex coupling correlations between time and space in systems with multiple physical processes and a large number of dynamic variables, resulting in insufficient prediction accuracy and uncertainty quantification, which affects prediction performance.

Method used

A multi-source data fusion method based on spatiotemporal diffusion is adopted. Satellite remote sensing images and radar echo reflectivity images are used to train synchronous mapping networks and prediction networks. Combined with the spatiotemporal diffusion model, the spatiotemporal joint distribution of lightning events is directly learned, the uncertainty is quantified, and the spatial distribution image of future lightning strike points is obtained through splicing.

Benefits of technology

It significantly improves the accuracy and robustness of future event predictions, can provide preliminary spatial distribution images of lightning strike points in shorter time intervals, quantify prediction uncertainties, and improve the accuracy and reliability of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807279A_ABST
    Figure CN120807279A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fusion image sequence generation method and system based on space-time diffusion, and belongs to the technical field of image sequence generation, and the method comprises the steps: taking an existing satellite remote sensing image and an existing radar echo reflectivity image as input, taking an existing lightning drop point space distribution diagram as output, and training a synchronous mapping network; training a prediction network through the existing satellite remote sensing image and the existing radar echo reflectivity image to obtain a future satellite remote sensing image and a future radar echo reflectivity image; inputting a future satellite remote sensing image and a future radar echo reflectivity image into the trained synchronous mapping network to obtain a first prediction map of a future lightning drop point spatial distribution map; inputting the existing lightning drop point spatial distribution map into the space-time diffusion model to obtain a second prediction map of the future lightning drop point spatial distribution map; and the first prediction image and the second prediction image are spliced to obtain a final prediction image sequence of the future lightning drop point spatial distribution map, so that the prediction precision and robustness are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image sequence generation, and particularly relates to a multi-source data fusion image sequence generation method and system based on space-time diffusion. BACKGROUND

[0002] Existing space-time point process models usually compromise the conditional independence assumption between time and space, or only allow spatial unilateral dependence on time (for example, predicting spatial position given time). This dependence limitation seriously weakens the complexity of the model in capturing the spatio-temporal interaction under the condition of historical events, thereby affecting the prediction performance. Secondly, the spatio-temporal joint distribution of events has a huge sample space, and it is extremely difficult to directly fit. Existing methods usually decompose the target distribution into conditional dependent distributions (for example, fitting time density and conditional spatial density respectively), but the characterization of conditional spatial density is often limited to specific model structures with weak expression ability, such as kernel density estimation (KDE) and continuous normalized flow (CNF). This structural limitation makes it difficult for the model to effectively capture the complex coupling correlation between time and space in the event occurrence process. In addition, most existing methods require complex integration operations when calculating the likelihood, or must limit the intensity function to be integrable, which sacrifices computational efficiency while pursuing accuracy.

[0003] At the level of deep fusion artificial intelligence algorithm, in-depth analysis of its mechanism shows that the most advanced space-time point prediction model is largely to produce deterministic prediction, thereby minimizing the mean square error. However, in systems involving multiple physical processes and a large number of dynamic variables, although the prediction index is improved, they lack quantification of uncertainty, which is gradually magnified over a long lead time and affects the ability of the model to describe the joint distribution, and limits the further improvement of prediction accuracy in mechanism. SUMMARY

[0004] The application develops a multi-source data fusion image sequence generation method and system based on space-time diffusion, aiming to solve the technical problem of low prediction accuracy in predicting future events in systems involving multiple physical processes and a large number of dynamic variables in the prior art.

[0005] Technical scheme: In a first aspect, the application provides a multi-source data fusion image sequence generation method based on space-time diffusion, applied to lightning fall point prediction, comprising: Using existing satellite remote sensing images and existing radar echo reflectivity images as inputs, and using an existing lightning fall point spatial distribution map as output, a synchronous mapping network is trained; A prediction network is trained through existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images; inputting the future satellite remote sensing image and the future radar echo reflectivity image into the trained synchronous mapping network to obtain a first prediction map of the future lightning landing point spatial distribution map; inputting the existing lightning landing point spatial distribution map into a spatiotemporal diffusion model to obtain a second prediction map of the future lightning landing point spatial distribution map; splicing the first prediction map and the second prediction map to obtain a final prediction image sequence of the future lightning landing point spatial distribution map.

[0006] In some embodiments, the step of inputting the existing lightning landing point spatial distribution map into the spatiotemporal diffusion model to obtain the second prediction map of the future lightning landing point spatial distribution map comprises: determining a time interval; dividing the existing lightning landing point spatial distribution map based on the time interval to obtain an existing spatiotemporal event image; obtaining an existing spatiotemporal event image sequence based on the existing spatiotemporal event image; a representation formula of the spatiotemporal event image sequence comprises: ; wherein, is the spatiotemporal event image sequence; is a matrix image containing a spatial information image used to represent a spatiotemporal event image at a corresponding time step , , is the number of time steps; is used to represent a value of the corresponding time step at a specified spatial position, indicates that a lightning event occurs at the specified spatial position at the corresponding time step , indicates that no lightning event occurs at the specified spatial position at the corresponding time step ; inputting the existing spatiotemporal event image sequence into the spatiotemporal diffusion model to predict a future spatiotemporal event image, which is the second prediction map.

[0007] In some embodiments, the spatiotemporal diffusion model comprises: a spatiotemporal self-attention encoder used to obtain a historical state representation; a spatiotemporal diffusion process module, the spatiotemporal diffusion module comprising a forward noise adding unit and a reverse noise removing unit; The existing lightning strike point spatial distribution map includes the current lightning strike point spatial distribution map and the historical lightning strike point spatial distribution map. The step of predicting the future spatiotemporal event image includes: Acquire a current spatiotemporal event image based on the current lightning strike point spatial distribution map; The current spatiotemporal event image is modeled through forward diffusion through Markov processes in the time domain and space domain, and Gaussian noise is added in each diffusion step to convert the current spatiotemporal event image into pure Gaussian noise; Obtaining a historical spatiotemporal event image based on the historical spatial distribution map of lightning strike points, and obtaining a valid representation of the historical spatiotemporal event image as a historical state representation; Based on the historical state representation, the pure Gaussian noise is subjected to inverse denoising to obtain original spatiotemporal information of the spatial distribution of the future lightning strike points; The future spatiotemporal event image is predicted based on the original spatiotemporal information.

[0008] In some embodiments, the characterization formula of the inverse denoising includes: ; ; ; in, The number of steps from inverse denoising to diffusion The subsequent spatiotemporal information image; The number of steps from inverse denoising to diffusion The spatial information image after The number of steps from inverse denoising to diffusion The subsequent time information image; For the number of diffusion steps The diffusion parameter is used to control the noise attenuation strength and signal retention ratio; For the number of diffusion steps The diffusion parameter is used to control the weight of the inverse denoising output and balance the correction amplitude; Modeling Diffusion to Diffusion Steps The spatial information image after Modeling Diffusion to Diffusion Steps The subsequent time information image; For the front Diffusion parameters in the reverse diffusion process The product of is the inverse denoising unit; Modeling Diffusion to Diffusion Steps The subsequent spatiotemporal information image; It represents the historical status; is the time step, used to represent the diffusion step number.

[0009] In some embodiments, before inverse denoising is performed, further comprising: Based on the inverse denoising result of the last diffusion step number and the historical state representation, a predicted noise is obtained, the predicted noise including spatial predicted noise and temporal predicted noise; the representation formula of the predicted noise includes: ; ; ; ; ; ; ; ; wherein, is the spatial predicted noise after inverse denoising to the diffusion step number ; is the temporal predicted noise after inverse denoising to the diffusion step number ; is the weight generated by the spatial attention mechanism; is the weight generated by the temporal attention mechanism; is the spatio-temporal information image of the current inverse denoising step; is the spatial information image of the current inverse denoising step; is the temporal information image of the current inverse denoising step; is an activation function; is a weight matrix for linear transformation of the spatial information image; is the spatial information image after inverse denoising to the diffusion step number ; is a weight matrix for linear transformation of the historical hidden spatial representation; is the historical hidden spatial representation; is a learnable parameter of linear projection; is a weight matrix for linear transformation of the temporal information image; is the temporal information image after inverse denoising to the diffusion step number ; is a weight matrix for linear transformation of the historical hidden temporal representation; is the historical hidden temporal representation; is the position encoding of the inverse denoising step ; is a sinusoidal position embedding operation; is an activation function; is a history state representation; and is a weight matrix for converting the history state representation and position encoding into attention weights; is a concatenation operation; performing reverse denoising of a next cycle based on the predicted noise.

[0010] In some embodiments, the step of obtaining the history state representation comprises: converting time information in a history spatio-temporal event image into a time embedding, and converting spatial information in the history spatio-temporal event image into a spatial embedding; obtaining a spatio-temporal embedding based on the time embedding and the spatial embedding; obtaining a plurality of time embedding sequences, spatial embedding sequences and spatio-temporal embedding sequences based on a plurality of history spatio-temporal event images; inputting the time embedding sequences into a time self-attention module of the spatio-temporal self-attention encoder to obtain a history hidden time representation; inputting the spatial embedding sequences into a spatial self-attention module of the spatio-temporal self-attention encoder to obtain a history hidden spatial representation; inputting the spatio-temporal embedding sequences into a spatio-temporal self-attention module of the spatio-temporal self-attention encoder to obtain a history hidden spatio-temporal representation; obtaining the history state representation based on the hidden time representation, the hidden spatial representation and the hidden spatio-temporal representation.

[0011] In some embodiments, the reverse denoising unit comprises a collaborative attention denoising network, and the step of training the collaborative attention denoising network comprises: obtaining a history spatio-temporal event image based on a history lightning strike spatial distribution map; performing diffusion modeling on the history spatio-temporal event image based on Markov processes in time domain and spatial domain to add Gaussian noise and obtain random noise; training the collaborative attention denoising network by taking the random noise as input data of the collaborative attention denoising network and taking the history spatio-temporal event image as target data of the collaborative attention denoising network; wherein the collaborative attention denoising network is trained by using a weighted mean square error loss function, and a representation formula of the weighted mean square error loss function comprises: ; wherein, is the weighted mean square error loss function; is an expectation, is , used to characterize the The original space-time representation of a lightning strike event, is the time step length, is the real noise, The number of times noise is added currently. Used to represent the original space-time representation under the original data distribution and real noise Calculate the average loss to ensure the model generalization ability; The inverse denoising unit is used to characterize the predicted noise; is the noise factor; Represents the historical status.

[0012] In some embodiments, the step of training a prediction network using existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images includes: determining the prediction network comprising a convolutional gating mechanism; Inputting the existing satellite remote sensing image and the existing radar echo reflectivity image into the prediction network; Based on a convolutional gating mechanism, the prediction network is trained by using the existing satellite remote sensing images and the existing radar echo reflectivity images to obtain spatiotemporal features for characterizing the spatiotemporal evolution of the weather system; The future satellite remote sensing image and the future radar echo reflectivity image are obtained based on the spatiotemporal characteristics, and the characterization formula thereof includes: ; ; ; ; ; ; in, is the input gate; For the Gate of Forgetfulness; is the output gate; is the candidate unit status; Sigmoid function is used to compress any value to between (0,1); Indicates the current time step The input feature map is used to represent satellite remote sensing images or radar echo reflectivity images; is the current time step The hidden state of The previous time step The hidden state of a cell state at a current time step; a cell state at a previous time step; a cell state at a previous time step; a cell state at a previous time step; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to a candidate cell state; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to an input gate; a convolution kernel weight corresponding to a candidate cell state; a bias parameter of an input gate; a bias parameter of a forget gate; a bias parameter of an output gate; a bias parameter of a candidate cell state; a convolution operation; a Hadamard product; a hyperbolic tangent function.

[0013] In some embodiments, the synchronous mapping network comprises a Unet encoder-decoder convolutional network, and the step of obtaining the first prediction map of the future lightning landing point spatial distribution map comprises: splicing the future satellite remote sensing image and the future radar echo reflectivity image to obtain multi-channel input data; performing down-sampling convolution on the multi-channel input data by the Unet encoder-decoder convolutional network to extract high-level semantic features and obtain a first feature map; performing up-sampling transposed convolution on the first feature map to restore spatial resolution and obtain a second feature map; processing the second feature map based on a skip connection and an activation function to obtain the first prediction map.

[0014] In a second aspect, the embodiments of the present application also provide a multi-source data fusion image sequence generation system based on space-time diffusion, applied to lightning landing point prediction, comprising: a mapping training module configured to train a synchronous mapping network by taking existing satellite remote sensing images and existing radar echo reflectivity images as input and taking an existing lightning strike spatial distribution map as output; a prediction training module configured to train a prediction network by taking existing satellite remote sensing images and existing radar echo reflectivity images as input and taking future satellite remote sensing images and future radar echo reflectivity images as output; a first prediction module configured to input the future satellite remote sensing images and the future radar echo reflectivity images into the trained synchronous mapping network and obtain a first prediction map of a future lightning strike spatial distribution map; a second prediction module configured to input the existing lightning strike spatial distribution map into a spatiotemporal diffusion model and obtain a second prediction map of the future lightning strike spatial distribution map; a splicing module configured to splice the first prediction map and the second prediction map and obtain a final prediction image sequence of the future lightning strike spatial distribution map.

[0015] Advantages: Compared with the prior art, the multi-source data fusion image sequence generation method based on spatiotemporal diffusion provided in the embodiments of the present application comprises the following steps: taking existing satellite remote sensing images and existing radar echo reflectivity images as input and taking an existing lightning strike spatial distribution map as output to train a synchronous mapping network; training a prediction network by taking existing satellite remote sensing images and existing radar echo reflectivity images as input and taking future satellite remote sensing images and future radar echo reflectivity images as output; inputting the future satellite remote sensing images and the future radar echo reflectivity images into the trained synchronous mapping network to obtain a first prediction map of a future lightning strike spatial distribution map; inputting the existing lightning strike spatial distribution map into a spatiotemporal diffusion model to obtain a second prediction map of the future lightning strike spatial distribution map; and splicing the first prediction map and the second prediction map to obtain a final prediction image sequence of the future lightning strike spatial distribution map. The present application regards lightning strike prediction as a spatiotemporal point process, directly learns the spatiotemporal joint distribution of lightning events by introducing a diffusion model, accurately quantifies the uncertainty in lightning prediction, and obtains a preliminary lightning strike spatial distribution image at a shorter time interval. The synchronous mapping network is trained by satellite remote sensing images and radar echo reflectivity images to obtain a preliminary lightning strike spatial distribution image. The positioning data of the two are fused to obtain a final positioning result. It can be seen that the method provided in the present application can significantly improve the prediction accuracy and robustness of future events in a system with multiple physical processes and a large number of dynamic variables. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0017] Figure 1 The step flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 2 The step flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 3 The step flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 4 The step chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 5 The step flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 6 The step flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 7 The module connection diagram of the multi-source data fusion image sequence generation system based on space-time diffusion provided by the embodiment of the present application; Figure 8 The program flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; Figure 9 The program flow chart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiment of the present application; The drawings are as follows: 10, mapping training module; 20, prediction training module; 30, first prediction module; 40, second prediction module; 50, splicing module. DETAILED DESCRIPTION

[0018] With reference to the drawings, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.

[0019] Timely and accurate prediction of lightning occurrence and development, quantification of prediction uncertainty (probability of lightning occurrence and development trend) is crucial for guiding relevant industries to reduce disaster impact and achieve cost-effective trade-offs, and is the basis for severe convective weather prediction, local weather process analysis, lightning disaster prevention and disaster analysis, and plays a crucial role in strengthening emergency preparedness and protective measures. Due to the mesoscale characteristics of weather systems, lightning prediction pays special attention to monitoring and nowcasting, and puts forward high requirements for the use of data with high temporal and spatial resolution such as weather radar, meteorological satellite, lightning detector, etc.

[0020] There are mainly two methods for generating image sequences for lightning, heavy precipitation and other nowcasting. One is the lightning diagnostic scheme based on NWP (Numerical Weather Prediction) output, which contains grid-based electrification parameterization or cloud microphysical parameters as proxies for lightning to obtain grid-based future lightning occurrence location image sequences. The other is based on extrapolation scheme, using satellite, weather radar, NWP model output and lightning location system measurement data, converting them into images at a certain resolution and using them separately or in combination. With the rapid increase of modern observation sources, new methods to improve convective storm and lightning nowcasting are also being explored.

[0021] The latest progress of machine learning and deep learning has greatly improved the data reasoning ability and shown great potential in weather applications. For example, various machine learning models have been developed to represent unresolved physical processes in coarse-scale climate models, improve medium-range weather forecasts, and process numerical weather model outputs. Deep learning techniques have been successfully applied to the extrapolation of convective precipitation spatiotemporal evolution based on radar images, such as by introducing ConvLSTM and U-net models to predict precipitation or cloud-to-ground (CG) lightning occurrence probability using meteorological radar, satellite or joint observation data.

[0022] However, existing methods have significant limitations in handling the spatio-temporal dependence and uncertainty quantification in lightning forecasting. Spatio-temporal point processes, as a framework to describe random events occurring with time and space, are widely applied in fields such as earthquakes, disease spread, and urban traffic. Although existing spatio-temporal point process models, such as Poisson process, Hawkes process, self-correcting process, and neural network-based recurrent marked temporal point process (RMTPP), neural Hawkes process (NHP), and Transformer Hawkes process (THP), have made progress in modeling event sequences, they face fundamental challenges in handling the highly entangled spatio-temporal phenomenon of lightning events.

[0023] Firstly, existing spatio-temporal point process models usually compromise on the conditional independence assumption between time and space, or only allow spatial dependence on time unilaterally (e.g., given time, predict spatial location). This dependence restriction severely weakens the model's ability to capture the complexity of spatio-temporal interactions under historical event conditions, thus affecting the prediction performance. Secondly, the spatio-temporal joint distribution of lightning events has a huge sample space, making direct fitting extremely difficult. Existing methods usually decompose the target distribution into conditional dependent distributions (e.g., fitting time density and conditional spatial density separately), but the characterization of conditional spatial density is often limited to specific model structures with weak expressive ability, such as kernel density estimation (KDE) and continuous normalized flow (CNF). This structural limitation makes it difficult for the model to effectively capture the complex coupling between time and space in the lightning occurrence process. In addition, most existing methods require complex integration operations when calculating the likelihood, or must restrict the intensity function to be integrable, which sacrifices computational efficiency while pursuing accuracy.

[0024] At the level of deep fusion artificial intelligence algorithms, in-depth analysis of its mechanism shows that the most advanced weather forecasting models are largely designed to produce deterministic predictions, thereby minimizing the mean square error. However, weather systems involve multiple physical processes and a large number of dynamic variables, and although the prediction indicators have improved, they lack quantification of uncertainty, a defect that is gradually amplified over longer lead times and affects the model's ability to describe the joint distribution, limiting further improvement in prediction accuracy from a mechanistic perspective.

[0025] Therefore, the application provides a multi-source data fusion image sequence generation method based on space-time diffusion, which comprises the following steps: taking existing satellite remote sensing images and existing radar echo reflectivity images as inputs, and taking an existing lightning strike spatial distribution map as an output, and training a synchronous mapping network; training a prediction network through the existing satellite remote sensing images and the existing radar echo reflectivity images, and obtaining future satellite remote sensing images and future radar echo reflectivity images; inputting the future satellite remote sensing images and the future radar echo reflectivity images into the trained synchronous mapping network, and obtaining a first prediction map of a future lightning strike spatial distribution map; inputting the existing lightning strike spatial distribution map into a space-time diffusion model, and obtaining a second prediction map of the future lightning strike spatial distribution map; and splicing the first prediction map and the second prediction map, and obtaining a final prediction image sequence of the future lightning strike spatial distribution map. The application regards lightning strike prediction as a space-time point process, directly learns the space-time joint distribution of lightning events by introducing a diffusion model, accurately quantifies the uncertainty in lightning prediction, and obtains a preliminary future lightning strike spatial distribution image at a shorter time interval. The synchronous mapping network is trained through satellite remote sensing images and radar echo reflectivity images, and a preliminary future lightning strike spatial distribution image is obtained. The positioning data of the two are fused to obtain the final positioning result. It can be seen that the method provided by the application can significantly improve the prediction accuracy and robustness of future events in a system with multiple physical processes and a large number of dynamic variables.

[0026] In some embodiments, referring to Figure 1 and Figure 8 , Figure 1 the step flowchart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiments of the application, Figure 8 the program flowchart of the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiments of the application, and the multi-source data fusion image sequence generation method based on space-time diffusion provided by the embodiments of the application is specifically implemented through steps 100 to 500: Step 100: taking existing satellite remote sensing images and existing radar echo reflectivity images as inputs, and taking an existing lightning strike spatial distribution map as an output, and training a synchronous mapping network.

[0027] Step 200: training a prediction network through the existing satellite remote sensing images and the existing radar echo reflectivity images, and obtaining future satellite remote sensing images and future radar echo reflectivity images.

[0028] In some embodiments, the method for obtaining the future satellite remote sensing images and the future radar echo reflectivity images in the application comprises: determining a prediction network containing a convolution gating mechanism.

[0029] inputting the existing satellite remote sensing images and the existing radar echo reflectivity images into the prediction network.

[0030] Based on the convolution gate mechanism, the prediction network is trained by using the existing satellite remote sensing image and the existing radar echo reflectivity image to obtain the spatiotemporal characteristics for representing the spatiotemporal evolution law of the weather system.

[0031] Based on the spatiotemporal characteristics, the future satellite remote sensing image and the future radar echo reflectivity image are obtained, and the representation formulae thereof include: ; ; ; ; ; ; wherein, is an input gate; is a forget gate; is an output gate; is a candidate cell state; is a Sigmoid function for compressing an arbitrary numerical value to (0, 1); represents an input feature map of a current time step for representing a satellite remote sensing image or a radar echo reflectivity image; is a hidden state of the current time step ; is a hidden state of a previous time step ; is a cell state of the current time step ; is a cell state of the previous time step ; is a convolution kernel weight corresponding to the output gate of the input feature map ; is a convolution kernel weight corresponding to the candidate cell state of the input feature map ; is a convolution kernel weight corresponding to the input gate of the input feature map ; is a convolution kernel weight corresponding to the forget gate of the input feature map ; is a convolution kernel weight corresponding to the input gate of the hidden state ; is a convolution kernel weight corresponding to the forget gate of the hidden state ; is a convolution kernel weight corresponding to the output gate of the hidden state ; is a convolution kernel weight corresponding to the candidate cell state of the hidden state ; is a bias parameter for the input gate; is a bias parameter for the forget gate; is a bias parameter for the output gate; is a bias parameter for the candidate cell state; is a convolution operation; is a Hadamard product; is a hyperbolic tangent function.

[0032] Specifically, the application extracts spatio-temporal features through a convolution gating mechanism based on a ConvLSTM convolutional long short-term memory network.

[0033] Step 300: input the future satellite remote sensing image and the future radar echo reflectivity image into the trained synchronous mapping network to obtain a first prediction map of the future lightning landing point spatial distribution map.

[0034] In some embodiments, the synchronous mapping network includes a Unet encoder-decoder convolutional network, please refer to Figure 6 , Figure 6 The step flow chart for obtaining the first prediction map of the future lightning landing point spatial distribution map in the method for generating a multi-source data fusion image sequence based on spatio-temporal diffusion provided by the embodiments of the application, the method for obtaining the first prediction map of the future lightning landing point spatial distribution map in the application is specifically implemented through steps 310 to 340: Step 310: splice the future satellite remote sensing image and the future radar echo reflectivity image to obtain multi-channel input data.

[0035] Step 320: perform down-sampling convolution on the multi-channel input data through the Unet encoder-decoder convolutional network to extract high-level semantic features and obtain a first feature map.

[0036] Step 330: perform up-sampling transposed convolution on the first feature map to restore the spatial resolution and obtain a second feature map.

[0037] Step 340: process the second feature map based on the skip connection and the activation function to obtain the first prediction map.

[0038] Specifically, the application splices the two features output by the ConvLSTM into multi-channel input, extracts high-level semantic features through the down-sampling convolution layer of the Unet with the encoder-decoder structure, restores the spatial resolution through the up-sampling transposed convolution, retains the details in combination with the skip connection, and finally uses the Sigmoid activation function in the output layer to generate the lightning probability map.

[0039] Step 400: input the existing lightning landing point spatial distribution map into the spatio-temporal diffusion model to obtain a second prediction map of the future lightning landing point spatial distribution map.

[0040] In some embodiments, the spatiotemporal diffusion model includes: Spatiotemporal self-attention encoder, which is used to obtain historical state representation; The spatiotemporal diffusion process module includes a forward denoising unit and a reverse denoising unit; In some embodiments, the reverse denoising unit includes a collaborative attention denoising network, and the existing lightning strike spatial distribution map includes a current lightning strike spatial distribution map and a historical lightning strike spatial distribution map. In some embodiments, refer to Figure 5 , Figure 5 This is a flowchart of the steps for training a collaborative attention denoising network in the method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion provided in an embodiment of the present application. The method for training a collaborative attention denoising network in the present application is specifically implemented through steps 401 to 403: Step 401: Acquire a historical spatiotemporal event image based on a historical lightning strike point spatial distribution map.

[0041] Specifically, first collect historical original lightning events, and obtain lightning spatiotemporal event image sequences based on historical original lightning events. sequence The time interval of the spatiotemporal event image is 1 minute. , An image containing spatial location information Matrix diagram of the corresponding time step The space-time event image, , is the number of time steps; For characterization Corresponding time step The value at the specified spatial location, express Corresponding time step A lightning event occurs at a specified spatial location. express Corresponding time step There are no lightning events at the specified spatial location.

[0042] Step 402: Perform diffusion modeling on the historical spatiotemporal event image based on the Markov process in the time domain and the space domain to add Gaussian noise to obtain random noise.

[0043] Specifically, the diffusion of lightning events is modeled as a Markov process on the image pixel grid through the forward diffusion process. The diffusion of is modeled as a Markov process in the spatial and temporal domains ,in, is the number of diffusion steps, is the total number of diffusion steps. In this process, the spatiotemporal diffusion model gradually adds a small amount of Gaussian noise to the time interval and spatial position of the lightning event independently until the noise becomes pure Gaussian noise .

[0044] Step 403: Take the random noise as input data of the collaborative attention denoising network, take the historical spatiotemporal event image sequence as target data of the collaborative attention denoising network, and train the collaborative attention denoising network.

[0045] Specifically, the random noise is taken as input data of the collaborative attention denoising network, the original lightning event spatiotemporal image sequence is taken as target data, and the collaborative attention denoising network is trained. The training adopts a weighted mean square error loss function: ; wherein, is the weighted mean square error loss function; is the expectation, is , used to represent the original spatiotemporal representation of the th lightning strike event, is the step length of the time step, is the real noise, is the number of times of adding noise at present, is used to represent the original spatiotemporal representation and the real noise under the original data distribution; the average loss is calculated to ensure the model generalization ability; is the inverse denoising unit, used to represent the predicted noise; is the noise coefficient; is the historical state representation.

[0046] In some embodiments, please refer to Figure 2 , Figure 2 is a step flowchart for acquiring a second prediction map in the spatiotemporal diffusion-based multi-source data fusion image sequence generation method provided by the embodiments of the present application. The method for acquiring the second prediction map in the present application is specifically implemented through steps 410 to 440: Step 410: Determine the time interval.

[0047] Specifically, the time interval in the present application is set to 1 min.

[0048] Step 420: Divide the existing lightning strike spatial distribution map based on the time interval to obtain an existing spatiotemporal event image.

[0049] Specifically, a current raw lightning event is acquired, and the current raw lightning event is divided based on a 1-minute time interval to acquire an existing spatiotemporal event image.

[0050] Step 430: Based on the existing spatiotemporal event image, an existing spatiotemporal event image sequence is acquired.

[0051] In some embodiments, the representation formula of the spatiotemporal event image sequence comprises: ; Wherein, is the spatiotemporal event image sequence; is a matrix image containing a spatial information image for representing the spatiotemporal event image of the corresponding time step , , is the number of time steps; is used to represent the value of the corresponding time step at a specified spatial position, indicates that a lightning event occurs at the specified spatial position at the corresponding time step , indicates that there is no lightning event at the specified spatial position at the corresponding time step .

[0052] Step 440: The existing spatiotemporal event image sequence is input into a spatiotemporal diffusion model to predict a future spatiotemporal event image, which is a second prediction image.

[0053] Specifically, the existing lightning landing point spatial distribution map includes a current lightning landing point spatial distribution map and a historical lightning landing point spatial distribution map, please refer to Figure 3 and Figure 9 , Figure 3 is a step flow chart for predicting a future spatiotemporal event image in the multi-source data fusion image sequence generation method based on spatiotemporal diffusion provided by the embodiments of the present application, Figure 9 is a program flow chart for reverse denoising in the multi-source data fusion image sequence generation method based on spatiotemporal diffusion provided by the embodiments of the present application, the method for predicting a future spatiotemporal event image in the present application is specifically implemented by steps 441 to 445: Step 441: A current spatiotemporal event image is acquired based on a current lightning landing point spatial distribution map.

[0054] Step 442: A forward diffusion modeling is performed on the current spatiotemporal event image through a Markov process in the time domain and the spatial domain, and Gaussian noise is added in each diffusion step to convert the current spatiotemporal event image into pure Gaussian noise.

[0055] Step 443: Obtain a historical spatio-temporal event image based on a historical lightning landing point spatial distribution map, and obtain an effective representation of the historical spatio-temporal event image, as a historical state representation.

[0056] In some embodiments, referring to Figure 4 , Figure 4 The step chart for obtaining a historical state representation in the multi-source data fusion image sequence generation method based on spatio-temporal diffusion provided by the embodiments of the present application, in the present application, the method for obtaining a historical state representation is specifically implemented through steps 4431 to 4437: Step 4431: convert the time information in the historical spatio-temporal event image into a time embedding, and convert the spatial information in the historical spatio-temporal event image into a spatial embedding.

[0057] Specifically, the timestamp and the spatial position in the historical spatio-temporal event image are converted into a time embedding and a spatial embedding .

[0058] Step 4432: obtain a spatio-temporal embedding based on the time embedding and the spatial embedding.

[0059] Specifically, the time embedding and the spatial embedding of each lightning event are added to obtain a spatio-temporal embedding .

[0060] Step 4433: obtain a plurality of time embedding sequences, spatial embedding sequences and spatio-temporal embedding sequences based on a plurality of historical spatio-temporal event images.

[0061] Specifically, the time embedding sequence , the spatial embedding sequence , and the spatio-temporal embedding sequence .

[0062] Step 4434: input the time embedding sequence into a time self-attention module of a spatio-temporal self-attention encoder to obtain a historical hidden time representation.

[0063] Step 4435: input the spatial embedding sequence into a spatial self-attention module of the spatio-temporal self-attention encoder to obtain a historical hidden spatial representation.

[0064] Step 4436: input the spatio-temporal embedding sequence into a spatio-temporal self-attention module of the spatio-temporal self-attention encoder to obtain a historical hidden spatio-temporal representation.

[0065] Step 4437: obtain a historical state representation based on the hidden time representation, the hidden spatial representation and the hidden spatio-temporal representation.

[0066] In particular, the historical state representation , is a historical hidden space representation, is a historical hidden time representation, is a historical hidden space-time representation.

[0067] Step 444: reverse denoising of pure Gaussian noise is performed conditioned on the historical state representation, to obtain the original space-time information of the future lightning strike distribution.

[0068] In some embodiments, before reverse denoising, further comprising: based on the reverse denoising result of the last diffusion step and the historical state representation, obtaining a predicted noise, the predicted noise including a spatial predicted noise and a temporal predicted noise; the representation formula of the predicted noise includes: ; ; ; ; ; ; ; ; wherein, is the spatial predicted noise after reverse denoising to the diffusion step ; is the temporal predicted noise after reverse denoising to the diffusion step ; is the weight generated by the spatial attention mechanism; is the weight generated by the temporal attention mechanism; is the space-time information image of the current reverse denoising step; is the spatial information image of the current reverse denoising step; is the temporal information image of the current reverse denoising step; is an activation function; is a weight matrix for linear transformation of the spatial information image; is the spatial information image after reverse denoising to the diffusion step ; is a weight matrix for linear transformation of the historical hidden space representation; is the historical hidden space representation; is a learnable parameter of linear projection; is a weight matrix for linear transformation of the temporal information image; is the temporal information image after reverse denoising to the diffusion step The subsequent time information image; is the weight matrix for linear transformation of historical hidden time representation; Hiding time representation for history; is the reverse denoising step Positional encoding; It is a sinusoidal position embedding operation; is the activation function; It represents the historical status; and is the weight matrix, which is used to represent the historical state and positional encoding Convert to attention weight; For splicing operation.

[0069] The next cycle of inverse denoising is performed based on the predicted noise.

[0070] It is understandable that the core component of inverse denoising is the collaborative attention denoising network, which is a denoising neural network whose key purpose is to capture the complex interdependencies between spatial and temporal domains, thereby significantly promoting the learning of spatiotemporal joint distribution. The collaborative attention denoising network shares the same network structure in each denoising step and receives the value generated by the denoising result of the previous step. and , denoising step with position encoding and the hidden representation from the spatiotemporal self-attention encoder As input, conditional denoising is achieved. The collaborative attention denoising network quantifies the dependencies between time and space by calculating mutual attention weights. Finally, the collaborative attention denoising network includes spatial attention and temporal attention, and the outputs of spatial attention and temporal attention are the predicted noise and Based on these predicted noise and , the spatiotemporal diffusion model can obtain the The predicted value of the step and , and feed it back into the collaborative attention denoising network to iteratively predict and gradually approximate the clean values ​​of the spatial and temporal dimensions. Through this iterative refinement process, the interdependence between time and space is adaptively and dynamically captured.

[0071] In some embodiments, the characterization formula of the inverse denoising process includes: ; ; ; in, The number of steps from inverse denoising to diffusion The subsequent spatiotemporal information image; The number of steps from inverse denoising to diffusion The spatial information image after The number of steps from inverse denoising to diffusion The subsequent time information image; For the number of diffusion steps The diffusion parameter is used to control the noise attenuation strength and signal retention ratio; For the number of diffusion steps The diffusion parameter is used to control the weight of the inverse denoising output and balance the correction amplitude; Modeling Diffusion to Diffusion Steps The spatial information image after Modeling Diffusion to Diffusion Steps The subsequent time information image; For the front Diffusion parameters in the reverse diffusion process The product of It is the reverse denoising unit; Modeling Diffusion to Diffusion Steps The subsequent spatiotemporal information image; It represents the historical status; is the time step, which is used to represent the number of diffusion steps.

[0072] Understandably, the trained spatiotemporal self-attention encoder is used to obtain its hidden representation based on the past event image sequence. Subsequently, the diffusion model starts with pure Gaussian noise and, conditioned on the hidden representation, it iteratively predicts the spatiotemporal information of the next lightning event.

[0073] Cooperative Attention Denoising Network As the core denoiser, it receives the denoising result from the previous step. , Hidden Representation of Historical Image Sequences and diffusion steps As input. By repeating the denoising process, the diffusion model can generate a series of trajectory sets of future lightning events using different initial noise samples. This inherent probability generation ability is the key to the diffusion model's quantification of uncertainty, which can provide a comprehensive view of the possible distribution of lightning strike points rather than just a single deterministic prediction. The raw spatiotemporal information of these predicted lightning events is then used to generate weather state data for future time steps. , weather status data This includes predictions of future lightning strike locations or lightning approach trends. This rich probabilistic output is of significant value in guiding critical decisions such as disaster warning and emergency response.

[0074] Step 445: Predict future spatiotemporal event images based on the original spatiotemporal information.

[0075] Step 500: Splice the first prediction image and the second prediction image to obtain a final prediction image sequence of the spatial distribution map of future lightning strike points.

[0076] Specifically, the present application inputs the spliced ​​image into a CNN network, which contains multiple convolutional layers, each of which uses a 3x3 convolution kernel and combines the ReLU activation function for feature extraction. Next, a pooling layer is added to reduce the spatial dimension and retain important features. After that, high-level semantic features are gradually extracted by stacking multiple convolution and pooling layers. Finally, a fully connected layer is used to map the features to the final prediction results, and the output layer uses a softmax function to generate a probability distribution. The entire CNN network is optimized by backpropagation and stochastic gradient descent, and the loss function uses cross-entropy loss to ensure that the model can effectively learn the features of multi-source data fusion during training, so as to perform feature fusion through the CNN network and obtain the final prediction results.

[0077] It can be understood that the multi-source data fusion image sequence generation method based on spatiotemporal diffusion provided in the embodiment of the present application includes training a synchronous mapping network with existing satellite remote sensing images and existing radar echo reflectivity images as input and an existing lightning strike point spatial distribution map as output; training a prediction network with existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images; inputting the future satellite remote sensing images and future radar echo reflectivity images into the trained synchronous mapping network to obtain a first prediction map of the future lightning strike point spatial distribution map; inputting the existing lightning strike point spatial distribution map into the spatiotemporal diffusion model to obtain a second prediction map of the future lightning strike point spatial distribution map; and splicing the first prediction map and the second prediction map to obtain a final prediction image sequence of the future lightning strike point spatial distribution map. The present application regards lightning strike point prediction as a spatiotemporal point process. By introducing a diffusion model, the spatiotemporal joint distribution of lightning events is directly learned, the uncertainty in lightning forecasting is accurately quantified, and a preliminary future lightning strike point spatial distribution image is obtained at a shorter time interval. A synchronous mapping network is trained using satellite remote sensing images and radar echo reflectivity images to obtain a preliminary spatial distribution image of future lightning strike points. The positioning data from both is then spliced ​​and fused to obtain the final positioning result. This demonstrates that the method provided by this application can significantly improve the accuracy and robustness of future event predictions in systems with multiple physical processes and a large number of dynamic variables.

[0078] Accordingly, the embodiment of the present application also provides a multi-source data fusion image sequence generation system based on spatiotemporal diffusion, which is applied to lightning strike prediction. Figure 7 , Figure 7A module connection diagram of a multi-source data fusion image sequence generation system based on space-time diffusion provided by an embodiment of the present application, the multi-source data fusion image sequence generation system based on space-time diffusion provided by the present application comprises: A mapping training module 10, the mapping training module 10 is configured to take an existing satellite remote sensing image and an existing radar echo reflectivity image as input, and train a synchronous mapping network with an existing lightning fall point spatial distribution map as output; A prediction training module 20, the prediction training module 20 is configured to train a prediction network by using the existing satellite remote sensing image and the existing radar echo reflectivity image, and obtain a future satellite remote sensing image and a future radar echo reflectivity image; A first prediction module 30, the first prediction module 30 is configured to input the future satellite remote sensing image and the future radar echo reflectivity image into the trained synchronous mapping network, and obtain a first prediction map of a future lightning fall point spatial distribution map; A second prediction module 40, the second prediction module 40 is configured to input the existing lightning fall point spatial distribution map into a space-time diffusion model, and obtain a second prediction map of the future lightning fall point spatial distribution map; A splicing module 50, the splicing module 50 is configured to splice the first prediction map and the second prediction map, and obtain a final prediction image sequence of the future lightning fall point spatial distribution map.

[0079] In some embodiments, the second prediction module 40 is specifically configured to: Determine a time interval; Divide the existing lightning fall point spatial distribution map based on the time interval, and obtain an existing space-time event image; Obtain an existing space-time event image sequence based on the existing space-time event image; a representation formula of the space-time event image sequence comprises: ; Wherein, is the space-time event image sequence; is a matrix image containing a spatial information image , used to represent a space-time event image at a corresponding time step , , is the number of time steps; is used to represent a value of the corresponding time step at a specified spatial position, represents a lightning event occurring at the specified spatial position at the corresponding time step , represents no lightning event at the specified spatial position at the corresponding time step ; The existing spatiotemporal event image sequence is input into the spatiotemporal diffusion model to predict a future spatiotemporal event image as a second prediction image.

[0080] In some embodiments, the second prediction module 40 is specifically configured to: obtain a current spatiotemporal event image based on a current lightning landing point spatial distribution map; model forward diffusion of the current spatiotemporal event image through Markov processes in the time domain and the spatial domain, and add Gaussian noise at each diffusion step to convert the current spatiotemporal event image into pure Gaussian noise; obtain a historical spatiotemporal event image based on a historical lightning landing point spatial distribution map, and obtain an effective representation of the historical spatiotemporal event image as a historical state representation; obtain original spatiotemporal information of a future lightning landing point spatial distribution by reverse denoising the pure Gaussian noise conditioned on the historical state representation; predict a future spatiotemporal event image based on the original spatiotemporal information.

[0081] In some embodiments, the spatiotemporal diffusion multi-source data fusion image sequence generation system is specifically configured to: obtain prediction noise based on reverse denoising results of the last diffusion step and the historical state representation, the prediction noise including spatial prediction noise and temporal prediction noise; and a representation formula of the prediction noise includes: ; ; ; ; ; ; ; ; wherein, is the spatial prediction noise after reverse denoising to the diffusion step ; is the temporal prediction noise after reverse denoising to the diffusion step ; is a weight generated by a spatial attention mechanism; is a weight generated by a temporal attention mechanism; is a spatiotemporal information image of a current reverse denoising step; is a spatial information image of the current reverse denoising step; is a temporal information image of the current reverse denoising step; is an activation function; a weight matrix for linearly transforming the spatial information image; a spatial information image after inverse denoising to a diffusion step number ; a weight matrix for linearly transforming the historical hidden spatial representation; a historical hidden spatial representation; a learnable parameter for linear projection; a weight matrix for linearly transforming the temporal information image; a temporal information image after inverse denoising to a diffusion step number ; a weight matrix for linearly transforming the historical hidden temporal representation; a historical hidden temporal representation; a position encoding for the inverse denoising step ; a sinusoidal position embedding operation; an activation function; a historical state representation; and a weight matrix for converting the historical state representation and the position encoding into attention weights; a concatenation operation; performing inverse denoising for the next cycle based on the predicted noise.

[0082] In some embodiments, the second prediction module 40 is specifically configured to: convert the temporal information in the historical spatio-temporal event image into a temporal embedding, and convert the spatial information in the historical spatio-temporal event image into a spatial embedding; obtain a spatio-temporal embedding based on the temporal embedding and the spatial embedding; obtain a plurality of temporal embedding sequences, spatial embedding sequences and spatio-temporal embedding sequences based on a plurality of historical spatio-temporal event images; input the temporal embedding sequences into a temporal self-attention module of the spatio-temporal self-attention encoder to obtain a historical hidden temporal representation; input the spatial embedding sequences into a spatial self-attention module of the spatio-temporal self-attention encoder to obtain a historical hidden spatial representation; input the spatio-temporal embedding sequences into a spatio-temporal self-attention module of the spatio-temporal self-attention encoder to obtain a historical hidden spatio-temporal representation; obtain a historical state representation based on the hidden temporal representation, the hidden spatial representation and the hidden spatio-temporal representation.

[0083] The above provides a detailed introduction to the method and system for generating a multi-source data fusion image sequence based on space-time diffusion provided by the embodiments of the present application. The principles and implementation modes of the present application are described in this paper by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A method for generating multi-source data fusion image sequences based on spatiotemporal diffusion, characterized in that: Applied to lightning strike prediction, including: The synchronous mapping network is trained using existing satellite remote sensing images and existing radar echo reflectivity images as input and existing lightning strike spatial distribution maps as output; Training a prediction network using existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images; Inputting the future satellite remote sensing image and the future radar echo reflectivity image into the trained synchronous mapping network to obtain a first predicted image of the future spatial distribution map of lightning strike points; Inputting the existing spatial distribution map of lightning strike points into a spatiotemporal diffusion model to obtain a second predicted map of the spatial distribution map of lightning strike points in the future; The first prediction image and the second prediction image are spliced ​​together to obtain a final prediction image sequence of the spatial distribution map of the future lightning strike points.

2. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 1, characterized in that: The step of inputting the existing spatial distribution map of lightning strike points into the spatiotemporal diffusion model to obtain a second predicted map of the spatial distribution map of lightning strike points in the future comprises: Determine the time interval; Divide the existing spatial distribution map of lightning strike points based on the time interval to obtain an existing spatiotemporal event image; Based on the existing spatiotemporal event image, an existing spatiotemporal event image sequence is obtained; the representation formula of the spatiotemporal event image sequence includes: ; in, is the spatiotemporal event image sequence; For images containing spatial information Matrix diagram of the corresponding time step The space-time event image, , is the number of time steps; For characterization Corresponding time step The value at the specified spatial location, express Corresponding time step A lightning event occurs at a specified spatial location. express Corresponding time step There are no lightning events at the specified spatial location; The existing spatiotemporal event image sequence is input into the spatiotemporal diffusion model to predict future spatiotemporal event images, which are the second prediction images.

3. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 2, characterized in that: The spatiotemporal diffusion model includes: A spatiotemporal self-attention encoder, which is used to obtain historical state representation; A spatiotemporal diffusion process module, wherein the spatiotemporal diffusion module includes a forward denoising unit and a reverse denoising unit; The existing lightning strike point spatial distribution map includes the current lightning strike point spatial distribution map and the historical lightning strike point spatial distribution map. The step of predicting the future spatiotemporal event image includes: Acquire a current spatiotemporal event image based on the current lightning strike point spatial distribution map; The current spatiotemporal event image is modeled through forward diffusion through Markov processes in the time domain and space domain, and Gaussian noise is added in each diffusion step to convert the current spatiotemporal event image into pure Gaussian noise; Obtaining a historical spatiotemporal event image based on the historical spatial distribution map of lightning strike points, and obtaining a valid representation of the historical spatiotemporal event image as a historical state representation; Based on the historical state representation, the pure Gaussian noise is subjected to inverse denoising to obtain original spatiotemporal information of the spatial distribution of the future lightning strike points; The future spatiotemporal event image is predicted based on the original spatiotemporal information.

4. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 3, characterized in that: The characterization formula of the inverse denoising includes: ; ; ; in, The number of steps from inverse denoising to diffusion The subsequent spatiotemporal information image; The number of steps from inverse denoising to diffusion The spatial information image after The number of steps from inverse denoising to diffusion The subsequent time information image; For the number of diffusion steps The diffusion parameter is used to control the noise attenuation strength and signal retention ratio; For the number of diffusion steps The diffusion parameter is used to control the weight of the inverse denoising output and balance the correction amplitude; Modeling Diffusion to Diffusion Steps The spatial information image after Modeling Diffusion to Diffusion Steps The subsequent time information image; For the front Diffusion parameters in the reverse diffusion process The product of is the inverse denoising unit; Modeling Diffusion to Diffusion Steps The subsequent spatiotemporal information image; It represents the historical status; is the time step, which is used to represent the number of diffusion steps.

5. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 3, characterized in that: Before inverse denoising, it also includes: Based on the inverse denoising result of the previous diffusion step and the historical state representation, the prediction noise is obtained, and the prediction noise includes spatial prediction noise and temporal prediction noise; the characterization formula of the prediction noise includes: ; ; ; ; ; ; ; ; in, The number of steps from inverse denoising to diffusion The spatial prediction noise after The number of steps from inverse denoising to diffusion The subsequent time prediction noise; Weights generated for the spatial attention mechanism; Weights generated for the temporal attention mechanism; is the spatiotemporal information image of the current reverse denoising step; is the spatial information image of the current reverse denoising step; is the time information image of the current reverse denoising step; is the activation function; is the weight matrix for linear transformation of spatial information image; The number of steps from inverse denoising to diffusion The spatial information image after is the weight matrix for linear transformation of the historical latent space representation; Hidden spatial representation for history; is the learnable parameter of the linear projection; is the weight matrix for linear transformation of the time information image; The number of steps from inverse denoising to diffusion The subsequent time information image; is the weight matrix for linear transformation of historical hidden time representation; Hiding time representation for history; is the reverse denoising step Positional encoding; It is a sinusoidal position embedding operation; is the activation function; It represents the historical status; and is the weight matrix, which is used to represent the historical state and positional encoding Convert to attention weight; For splicing operation; The next cycle of reverse denoising is performed based on the predicted noise.

6. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 3, characterized in that: The step of obtaining the historical status representation includes: Convert the temporal information in the historical spatiotemporal event images into temporal embeddings, and convert the spatial information in the historical spatiotemporal event images into spatial embeddings; Obtaining a spatiotemporal embedding based on the temporal embedding and the spatial embedding; Based on multiple historical spatiotemporal event images, multiple time embedding sequences, spatial embedding sequences, and spatiotemporal embedding sequences are obtained; Inputting the time embedding sequence into the temporal self-attention module of the spatiotemporal self-attention encoder to obtain a historical hidden time representation; Inputting the spatial embedding sequence into the spatial self-attention module of the spatiotemporal self-attention encoder to obtain a historical latent space representation; Inputting the spatiotemporal embedding sequence into the spatiotemporal self-attention module of the spatiotemporal self-attention encoder to obtain a historical hidden spatiotemporal representation; The historical state representation is obtained based on the latent time representation, the latent space representation and the latent spatiotemporal representation.

7. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 3, characterized in that: The inverse denoising unit includes a collaborative attention denoising network, and the steps of training the collaborative attention denoising network include: Acquire historical spatiotemporal event images based on the historical spatial distribution map of lightning strike points; Performing diffusion modeling on the historical spatiotemporal event image based on Markov processes in the time domain and the space domain to add Gaussian noise to obtain random noise; The random noise is used as input data of the collaborative attention denoising network, and the historical spatiotemporal event image is used as target data of the collaborative attention denoising network, and the collaborative attention denoising network is trained; wherein, the collaborative attention denoising network is trained using a weighted mean square error loss function, and the representation formula of the weighted mean square error loss function includes: ; in, is the weighted mean square error loss function; For expectations, for , used to characterize the The original space-time representation of a lightning strike event, is the time step length, is the real noise, The number of times noise is added currently. Used to represent the original space-time representation under the original data distribution and real noise Calculate the average loss to ensure the model generalization ability; The inverse denoising unit is used to characterize the predicted noise; is the noise factor; Represents the historical status.

8. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 1, characterized in that: The step of training a prediction network by using existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images includes: determining the prediction network comprising a convolutional gating mechanism; Inputting the existing satellite remote sensing image and the existing radar echo reflectivity image into the prediction network; Based on a convolutional gating mechanism, the prediction network is trained by using the existing satellite remote sensing images and the existing radar echo reflectivity images to obtain spatiotemporal features for characterizing the spatiotemporal evolution of the weather system; The future satellite remote sensing image and the future radar echo reflectivity image are obtained based on the spatiotemporal characteristics, and the characterization formula thereof includes: ; ; ; ; ; ; in, is the input gate; For the Gate of Forgetfulness; is the output gate; is the candidate unit status; is the Sigmoid function, which is used to compress any value to between (0,1); Indicates the current time step The input feature map is used to represent satellite remote sensing images or radar echo reflectivity images; is the current time step The hidden state of The previous time step The hidden state of is the current time step The unit status; The previous time step The unit status; Input feature map The convolution kernel weight corresponding to the output gate; Input feature map The convolution kernel weights corresponding to the candidate unit states; Input feature map The convolution kernel weight corresponding to the input gate; Input feature map The convolution kernel weight corresponding to the forget gate; Hidden The convolution kernel weight corresponding to the input gate; Hidden The convolution kernel weight corresponding to the forget gate; Hidden The convolution kernel weight corresponding to the output gate; Hidden The convolution kernel weights corresponding to the candidate unit states; is the bias parameter of the input gate; is the bias parameter of the forget gate; is the bias parameter of the output gate; is the bias parameter of the candidate unit state; is the convolution operation; For Hadamard; is the hyperbolic tangent function.

9. The method for generating a multi-source data fusion image sequence based on spatiotemporal diffusion according to claim 1, characterized in that: The synchronous mapping network includes a Unet encoder-decoder convolutional network, and the step of obtaining a first prediction map of the spatial distribution map of future lightning strike points includes: splicing the future satellite remote sensing image and the future radar echo reflectivity image to obtain multi-channel input data; Downsampling and convolving the multi-channel input data through the Unet encoder-decoder convolutional network to extract high-level semantic features and obtain a first feature map; Performing upsampling and transposed convolution on the first feature map to restore spatial resolution and obtain a second feature map; The second feature map is processed based on a skip connection and an activation function to obtain a first prediction map.

10. A multi-source data fusion image sequence generation system based on spatiotemporal diffusion, characterized in that: Applied to lightning strike prediction, including: A mapping training module (10) is used to train a synchronous mapping network using existing satellite remote sensing images and existing radar echo reflectivity images as input and existing lightning strike point spatial distribution maps as output; A prediction training module (20), the prediction training module (20) is used to train a prediction network through existing satellite remote sensing images and existing radar echo reflectivity images to obtain future satellite remote sensing images and future radar echo reflectivity images; A first prediction module (30) is configured to input the future satellite remote sensing image and the future radar echo reflectivity image into the trained synchronous mapping network to obtain a first prediction map of the future lightning strike point spatial distribution map; A second prediction module (40), the second prediction module (40) is used to input the existing spatial distribution map of lightning strike points into the spatiotemporal diffusion model to obtain a second prediction map of the spatial distribution map of lightning strike points in the future; A splicing module (50) is used to splice the first prediction image and the second prediction image to obtain a final prediction image sequence of the future lightning strike point spatial distribution map.

Citation Information

Patent Citations

  • Short temporary rainfall prediction method based on space-time attention and data fusion

    CN114764602A

  • Multi-source data fused long-sequence radar image prediction method and system

    CN117368881A

  • Lightning nowcasting method and device based on space-time attention gating fusion network

    CN117555049A

  • Space-time event prediction method and device based on deep learning

    CN118378734A

  • Event sequence prediction method based on space-time diffusion generation network

    CN119691544A

Cited By

  • Radar echo extrapolation method based on space-time attention mechanism

    CN121049908A

  • Meteorological radar missing frame reconstruction method fusing spatio-temporal context information

    CN121613459A

  • Short-critical extreme rainfall prediction method based on abnormal driving residual dynamic diffusion

    CN122131269A