A motion feature driven image sequence data prediction method

CN122434978BActive Publication Date: 2026-08-21HANGZHOU DIANZI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610903941.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-08-21
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0004]当前遥感云图预测技术仍存在诸多亟待解决的技术痛点,其一,云运动具有高度不确定性,既包括运动轨迹的随机波动,也包括云团生成与消散过程的不规则性,现有时空预测模型多采用确定性建模策略,预测结果多为多种可能状态的叠加均值,易产生预测图像模糊、伪影等问题,即便引入生成式模型优化图像清晰度,整体预测效果仍未达到实际应用标准;其二,遥感云图数据蕴含海量时空信息,现有模型多直接对原始数据进行处理与预测,不仅导致云运动特征提取不充分,还会消耗大量计算资源,难以适配先进预测与生成模型的技术需求;其三,云运动特征的有效表征仍是核心技术难点,现有部分方法采用端到端预测模式,虽能在简单场景下提升预测精度,但在复杂云运动场景中存在明显不足,预测细节精度较低;另有部分方法尝试设计运动特征提取模块,却忽视了云的复杂生消演变规律及不同云类的差异化运动特性,无法实现全场景精准预测

Benefits of technology

[0026] This invention proposes a motion feature-driven image sequence data prediction method based on a diffusion model framework and the idea of ​​decoupling motion and content, achieving accurate prediction and generation of future frames of remote sensing cloud image sequences. The method, based on the idea of ​​decoupling motion and content and coordinating prediction and generation, designs a three-stage framework: First, a spatiotemporal dynamic coding module extracts multi-scale and multi-type cloud motion features; second, a future motion prediction module based on stochastic differential equations is constructed to model the uncertain evolution process of future cloud motion features; finally, an autoregressive diffusion generation module is used, fusing motion features and the predicted remote sensing cloud image from the previous frame as guiding conditions to generate clear and temporally coherent future remote sensing cloud image sequences frame by frame, effectively improving the accuracy, clarity, and temporal consistency of remote sensing cloud image sequence prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434978B_ABST
    Figure CN122434978B_ABST
Patent Text Reader

Abstract

The application discloses a motion feature driven image sequence data prediction method, and belongs to the field of remote sensing big data. The method is based on the idea of motion and content decoupling and prediction and generation collaboration, and a three-stage framework is designed. Firstly, a space-time dynamic coding module is used to extract multi-scale and multi-type cloud motion features. Secondly, a future motion prediction module based on a stochastic differential equation is constructed to model the uncertainty evolution process of future cloud motion features. Finally, an autoregressive diffusion generation module is used to generate a clear and time-sequential future remote sensing cloud image sequence frame by frame, effectively improving the accuracy, clarity and time sequence consistency of remote sensing cloud image sequence prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing big data, and in particular relates to a motion feature-driven image sequence data prediction method. Background Technology

[0002] Remote sensing cloud images, with their core advantages of full-area coverage and real-time dynamic monitoring, play an irreplaceable supporting role in key technology areas such as ecological environment assessment, meteorological and climate research, extreme weather early warning, and safety. In the field of meteorological forecasting, by analyzing the morphology, structure, and movement trajectory characteristics of cloud systems, accurate predictions of severe weather such as rainfall and storms can be made, providing a basis for the implementation of disaster prevention and mitigation technologies. In the field of aviation navigation, remote sensing cloud images can accurately identify dangerous cloud areas such as cumulonimbus clouds and thunderstorm clouds, enabling dynamic optimization of flight routes, ensuring flight safety, and improving operational efficiency. In the field of safety, by analyzing meteorological elements such as cloud cover and visibility in battlefield airspace through cloud images, key technical assistance can be provided for tactical action decision-making.

[0003] Remote sensing cloud image sequence prediction is essentially a spatiotemporal prediction problem of a series of images. Remote sensing cloud image sequences contain rich spatiotemporal dynamic information, encompassing the entire process of cloud formation, dissipation, drift, and evolution. The core technical challenge lies in accurately capturing the spatiotemporal dynamic characteristics of cloud movement and effectively modeling its uncertainties. Traditional remote sensing cloud image prediction methods are based on meteorological observation data and historical statistical analysis. Early methods relied on visual interpretation and experience-based inference by meteorological experts, resulting in limitations such as strong subjectivity, low prediction efficiency, and limited accuracy. Subsequent methods, such as spectral analysis, numerical analysis, and mathematical statistics, can only achieve simple cloud image trend prediction, failing to meet the demands of high-precision applications. In recent years, the rapid development of machine learning and deep learning technologies has provided new technical pathways for cloud image prediction. From convolutional neural networks and recurrent neural networks to spatiotemporal prediction models that integrate both, the limitations of traditional methods have been effectively overcome, driving the iterative upgrade of remote sensing cloud image prediction technology towards higher precision, longer timeliness, and wider coverage.

[0004] Current remote sensing cloud image prediction technology still faces several critical technical challenges. First, cloud motion is highly uncertain, including random fluctuations in trajectory and irregularities in cloud formation and dissipation. Existing spatiotemporal prediction models often employ deterministic modeling strategies, resulting in predictions that are merely the average of multiple possible states, easily leading to issues such as blurred images and artifacts. Even with the introduction of generative models to optimize image clarity, the overall prediction performance still falls short of practical application standards. Second, remote sensing cloud image data contains massive amounts of spatiotemporal information. Existing models often directly process and predict the raw data, resulting in insufficient extraction of cloud motion features and consuming significant computational resources, making it difficult to meet the technical requirements of advanced prediction and generative models. Third, the effective representation of cloud motion features remains a core technical challenge. Some existing methods employ end-to-end prediction models, which can improve prediction accuracy in simple scenarios but are significantly insufficient in complex cloud motion scenarios, resulting in low accuracy in predicting details. Other methods attempt to design motion feature extraction modules but neglect the complex formation and dissipation patterns of clouds and the differentiated motion characteristics of different cloud types, failing to achieve accurate prediction across all scenarios. Summary of the Invention

[0005] Cloud image sequence prediction provides crucial information for cloud motion analysis and meteorological situation assessment, playing a key role in atmospheric system research. However, most existing methods employ content-based end-to-end deterministic modeling, which can easily lead to blurred and artifact-prone predicted images, and makes it difficult to depict the entire process of cloud formation, dissipation, drift, and evolution. To address these issues, this invention delves into the theoretical research of cloud motion characterization and modeling, proposing a motion feature-driven image sequence data prediction method aimed at improving the accuracy and clarity of remote sensing cloud image sequence prediction.

[0006] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:

[0007] In a first aspect, the present invention provides a motion feature-driven image sequence data prediction method, comprising the following steps: inputting a historical remote sensing cloud image sequence into a prediction model comprising a spatiotemporal dynamic coding module, a future motion prediction module, and an autoregressive diffusion generation module; firstly, the spatiotemporal dynamic coding module extracts multi-scale, multi-type spatiotemporal dynamic contextual features containing cloud system motion patterns from the historical remote sensing cloud image sequence, providing high-quality initial contextual information for the future motion prediction module; then, the future motion prediction module, based on the spatiotemporal dynamic contextual features, performs evolution prediction of cloud system motion features at each future moment, models the time dependence and uncertainty of cloud motion, and outputs the motion features at each future moment; finally, the autoregressive diffusion generation module takes the last frame of the historical remote sensing cloud image sequence as the initial input, adopts autoregressive logic guided by the preceding frame, combines the motion features at each future moment to construct guiding conditions, and generates remote sensing cloud images at each future moment frame by frame, thereby predicting the future remote sensing cloud image sequence.

[0008] Based on the above scheme, each step can be implemented in the following preferred manner.

[0009] As a preferred embodiment of the first aspect mentioned above, the internal processing flow of the spatiotemporal dynamic coding module is as follows:

[0010] S11. By performing differential processing on the input historical remote sensing cloud image sequence, the difference image between adjacent frames is calculated to enhance the salience of motion information;

[0011] S12. The obtained difference image is then processed through a multi-branch convolutional structure to extract spatial features in order to capture motion features at different scales;

[0012] S13. Then, through the attention mechanism, the adaptive weighted fusion of multi-scale motion features is achieved to obtain the initial fusion features that integrate cloud morphology and motion details from micro to macro.

[0013] S14. Next, obtain the cloud type label sequence corresponding to the historical remote sensing cloud image sequence, and convert the discrete labels in the cloud type label sequence into continuous feature vectors through the embedding layer to obtain cloud type features; then fuse the cloud type features with the initial fusion features through feature concatenation to obtain the final fusion features;

[0014] S15. Finally, convolutional gated recurrent units are used to temporally encode the final fused features at each historical moment to model the temporal dependency of cloud motion, thereby obtaining the motion features after temporal fusion, and using the set of motion features as spatiotemporal dynamic context features.

[0015] As a preferred embodiment of the first aspect mentioned above, the internal processing flow of the future motion prediction module is as follows: a stochastic differential equation is used to model the continuous evolution process of motion characteristics at future moments, and the Euler numerical integration method is used for discretization calculation to solve the stochastic differential equation. Finally, the motion characteristics at each future moment are obtained through iterative solution.

[0016] As a preferred embodiment of the first aspect above, the stochastic differential equation is in the form that the differential of the motion feature modeled at time t is obtained by adding two parts: the first part is formed by multiplying the drift term and the differential at that time, and the second part is formed by multiplying the diffusion term and the increment of the Wiener process.

[0017] As a preferred embodiment of the first aspect mentioned above, both the drift term and the diffusion term are implemented using a parameterized network. Specifically, the parameterized network for the drift term adopts the UNet architecture, taking the concatenation result of the current motion features and the temporal embedding as input to achieve joint modeling of cloud motion from local details to global trends. The parameterized network for the diffusion term adopts a convolutional network of the same type as the parameterized network for the drift term, also taking the concatenation result of the current motion features and the temporal embedding as input, and adapting to the time-varying uncertainty of cloud motion by learning the perturbation intensity of different time intervals.

[0018] As a preferred embodiment of the first aspect mentioned above, in the autoregressive diffusion generation module, the future time step... The diffusion generation process of the frame remote sensing cloud image is as follows: The future time frame... The remote sensing cloud image of the first frame, as a preceding frame already generated, is stitched together with a random noise image to form the initial input of the trained diffusion network, which will then be used to generate the next frame. Motion characteristics at each time step are used as guiding conditions to generate remote sensing cloud images through reverse diffusion. During the diffusion process, the guiding conditions are injected into each layer of UNet through a cross-layer fusion mechanism. Finally, by progressively removing noise, the next time step is generated. Frame-by-frame remote sensing cloud image.

[0019] As a preferred option in the first aspect mentioned above, the specific process of injecting guiding conditions through a cross-layer fusion mechanism is as follows: first, for the future... Motion features at each time step are transformed into features of the same dimension as the output of each downsampling layer of UNet through multiple convolutions. The transformed features are then passed through a multilayer perceptron to obtain conditional features. In each downsampling and upsampling layer of UNet, the conditional features and the original features of that layer are fused through cross-layer adaptive group normalization.

[0020] In a second aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement a motion feature-driven image sequence data prediction method as described in any of the solutions in the first aspect above.

[0021] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a motion feature-driven image sequence data prediction method as described in any of the solutions of the first aspect above.

[0022] Fourthly, the present invention provides a computer electronic device, which includes a memory and a processor;

[0023] The memory is used to store computer programs;

[0024] The processor is configured to, when executing the computer program, implement a motion feature-driven image sequence data prediction method as described in any of the first aspects above.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] This invention proposes a motion feature-driven image sequence data prediction method based on a diffusion model framework and the idea of ​​decoupling motion and content, achieving accurate prediction and generation of future frames of remote sensing cloud image sequences. The method, based on the idea of ​​decoupling motion and content and coordinating prediction and generation, designs a three-stage framework: First, a spatiotemporal dynamic coding module extracts multi-scale and multi-type cloud motion features; second, a future motion prediction module based on stochastic differential equations is constructed to model the uncertain evolution process of future cloud motion features; finally, an autoregressive diffusion generation module is used, fusing motion features and the predicted remote sensing cloud image from the previous frame as guiding conditions to generate clear and temporally coherent future remote sensing cloud image sequences frame by frame, effectively improving the accuracy, clarity, and temporal consistency of remote sensing cloud image sequence prediction. Attached Figure Description

[0027] Figure 1 This is an overall flowchart of the method of the present invention;

[0028] Figure 2 This is a diagram of the spatiotemporal dynamic coding module architecture of the present invention;

[0029] Figure 3 This is a schematic diagram of the multi-branch convolution structure of the present invention;

[0030] Figure 4 This is a diagram of the future motion prediction module architecture of the present invention;

[0031] Figure 5 This is a schematic diagram of the parameterized stochastic differential equation of the present invention;

[0032] Figure 6 This is a schematic diagram of the predicted denoised remote sensing cloud image in the autoregressive diffusion generation module of the present invention;

[0033] Figure 7 This is a schematic diagram of a computer electronic device provided by the present invention. Detailed Implementation

[0034] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0035] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.

[0036] This invention, based on a diffusion model framework and the idea of ​​decoupling motion and content, proposes a motion feature-driven image sequence data prediction method to achieve accurate prediction and generation of future frames in remote sensing cloud image sequences. For example... Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned motion feature-driven image sequence data prediction method includes the following steps:

[0037] The historical remote sensing cloud image sequence is input into a prediction model comprising a Spatial Temporal Dynamics Encoder, a Future Motion Prediction module, and an Autoregressive Prediction module. First, the Spatial Temporal Dynamics Encoder extracts multi-scale, multi-type spatiotemporal dynamic contextual features containing cloud system movement patterns from the historical remote sensing cloud image sequence, providing high-quality initial contextual information for the Future Motion Prediction module. Then, based on the spatiotemporal dynamic contextual features, the Future Motion Prediction module predicts the evolution of cloud system movement characteristics at future times, models the temporal dependence and uncertainty of cloud movement, and outputs the movement characteristics at each future time. Finally, the Autoregressive Prediction module uses the last frame of the historical remote sensing cloud image sequence as initial input, employs autoregressive logic guided by previous frames, and combines the movement characteristics at each future time to construct guiding conditions, generating remote sensing cloud images frame by frame for each future time, thereby predicting the future remote sensing cloud image sequence.

[0038] The specific implementation process of each module of the prediction model will be described in detail below.

[0039] It should be noted that in this invention, the spatiotemporal dynamic coding module achieves joint coding of spatial morphological features and temporal evolution patterns in cloud image sequences through multi-scale feature extraction and dynamic temporal modeling. For example... Figure 2 As shown, its internal processing flow is as follows:

[0040] S11. By performing differential processing on the input historical remote sensing cloud image sequence, the differential image between adjacent frames is calculated to enhance the saliency of motion information.

[0041] In this embodiment S11, the historical remote sensing cloud image sequence Depend on The image consists of frames, the first The frame images are represented as follows: , respectively corresponding to the first time, The length of the historical remote sensing cloud image sequence. (The last part is a series of numbers and doesn't need a direct translation.) Frame Image For example, its corresponding difference image It can be represented as:

[0042]

[0043] in, These represent the image height, width, and number of channels, respectively. This step, by calculating pixel value differences, directly reflects abstract motion information such as the displacement trend, morphological changes, and degree of change of the cloud system at adjacent time points.

[0044] S12. The obtained difference image is then processed through a multi-branch convolutional structure to extract spatial features in order to capture motion features at different scales.

[0045] In this embodiment S12, let the first... The kernel size of each branch is The feature extraction process can be represented as follows:

[0046]

[0047] in, For the first The output of the branch Motion characteristics at any given moment; respectively motion characteristics Height and width, For feature dimensions; Indicates adoption Convolution operation with different kernel sizes; For the first Learnable parameters for each branch.

[0048] At each scale, the impact of each frame in the historical remote sensing cloud image sequence on remote sensing cloud image prediction varies. Therefore, as... Figure 3 As shown, the present invention designs the above-mentioned multi-branch convolutional structure, which enables the model to learn a wider range of cloud motion information: small-sized convolutional kernels focus on cloud edge details and local turbulent motion, while large-sized convolutional kernels capture cloud band distribution and global migration trends.

[0049] S13. Then, through the attention mechanism, the adaptive weighted fusion of multi-scale motion features is achieved to obtain the initial fusion features that integrate cloud morphology and motion details from micro to macro.

[0050] In this embodiment S13, most studies use fixed weights to concatenate multi-scale motion features, which is difficult to adapt to motion characteristics at different scales. The method of this invention achieves adaptive weighted fusion of multi-scale motion features through an attention mechanism, so as to simultaneously retain micro and macro motion features and highlight motion information at key scales. In this process, the model learns relevant motion features at different levels of abstraction. Small-scale local motion information and large-scale global motion trends are fused through the attention mechanism, which can assign different weights to motion feature sequences at each scale and emphasize or weaken specific parts according to the prediction task, thereby improving the model's performance in global motion prediction and local detail restoration in cloud motion prediction tasks.

[0051] Specifically, this embodiment first uses convolution operations to map motion features at each scale into query vectors, key vectors, and value vectors, respectively, achieving feature dimension unification and spatial information transformation; then, it calculates the attention scores of the query vectors and key vectors, and performs scale normalization, fully connected layers, and... The function obtains attention weights to quantify the correlation and importance of features at different scales; finally, the value vectors are weighted and summed based on the attention weights to complete the fusion of multi-scale features.

[0052] Furthermore, the calculation process for the attention weights mentioned above is as follows:

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059] in, , , respectively motion characteristics The mapped query vector, key vector, and value vector; For attention feature dimensions; , , These are convolution operations used to map query vectors, key vectors, and value vectors, respectively. , and For the corresponding learnable parameters; For transpose; For the first The query vector of the branch and the first branch query vector Branch key vectors Attention score; For the first Learnable attention weights for each branch; This represents the number of branches in a multi-branch convolutional structure. For the first The query vector of the branch and the first branch query vector Attention score of each branch key vector; For the first Initial fusion features at time step; For the first Each branch value vector.

[0060] S14. Next, obtain the cloud type label sequence corresponding to the historical remote sensing cloud image sequence, and convert the discrete labels in the cloud type label sequence into continuous feature vectors through the embedding layer to obtain cloud type features; then fuse the cloud type features with the initial fusion features through feature concatenation to obtain the final fusion features.

[0061] In this embodiment S14, to address the diversity of cloud types, a cloud type perception mechanism is introduced to enhance the model's regional association and motion perception of cloud pixels. The aforementioned cloud type label sequence is denoted as... The sequence is also of length . , The first Cloud type label for the frame image. For the first... Individual cloud type tags its interior value at pixel position This represents the cloud type at that pixel location. , This represents the total number of cloud types. Based on the aforementioned cloud type label sequence, this embodiment further selects the top [number] from it. A cloud type tag is used for embedding encoding. The encoding process of the above embedding layer can be represented as:

[0062]

[0063] in, For the first Cloud type characteristics at any given time; Encoding operations for the embedding layer; For embedding layer parameters, For the embedded dimension.

[0064] After embedding and encoding, this embodiment further concatenates the cloud type features with the initial fused features:

[0065]

[0066] in, For the first The final fusion feature at any given moment includes joint information on cloud type and spatial morphology; This indicates that the cloud type feature is copied to a feature of the same size as the initial fused feature; This indicates a feature splicing operation.

[0067] As can be seen from the above process, the method proposed in this invention introduces cloud classification information. This information integrates inter-frame change information of remote sensing cloud images during use, and can effectively extract cloud motion features while fusing cloud type information, modeling specific motion features for each type of cloud pixel cluster. The aforementioned feature stitching fusion method enables the model to adaptively adjust feature representations for the motion characteristics of different cloud types.

[0068] S15. Finally, Convolutional Gated Recurrent Unit (ConvGRU) is used to temporally encode the final fused features at each historical moment to model the temporal dependency of cloud motion, thereby obtaining the temporally fused motion features, and using the set of motion features as spatiotemporal dynamic context features.

[0069] In this embodiment S15, the state update process of the convolution gated recurrent unit is defined as follows:

[0070]

[0071]

[0072]

[0073]

[0074] in, These are the update door and the reset door, respectively. It is the sigmoid activation function; , , , , , Both represent convolution operations in ConvGRU; , , , , , These are the learnable parameters of ConvGRU; No. The hidden state at any given moment; For the first The hidden state at all times For the hidden layer dimension; For the first Candidate hidden states used for residual updates at any time; For activation functions; This indicates element-wise multiplication.

[0075] By iterating every moment, ConvGRU will store the history. The final fusion feature at each moment Gradually integrated into a temporally coherent motion characteristic, denoted as , representing from the previous moment to The motion trend at any given moment can effectively capture the long-term motion trend of cloud systems. Ultimately, the above set of motion features is used as the spatiotemporal dynamic context features. It includes not only the spatial distribution and morphological details of cloud systems, but also integrates their dynamic evolution over time, providing comprehensive prior information for predicting the evolution of future motion characteristics.

[0076] It should be noted that in this invention, the future motion prediction module adopts neural stochastic differential equations (SDE) as its core framework, realizing the continuous evolution from historical motion patterns to future motion states through numerical integration. Since the proposed model requires spatiotemporal dynamic context features as input to predict future motion features and serve as conditional inputs for the diffusion model, this invention introduces stochastic differential equations to model the evolution of sequence features. This equation can model the evolution of cloud motion features as a stochastic process that changes continuously over time, thereby modeling the uncertainty of the probabilistic process of spatiotemporal dynamic context features. This is more conducive to capturing the generation and decay changes in cloud motion and is consistent with the diffusion model framework, facilitating training.

[0077] In this invention, such as Figure 4As shown, the internal processing flow of the future motion prediction module is as follows: stochastic differential equations are used to model the continuous evolution process of motion characteristics at future moments, and Euler numerical integration method is used for discretization calculation to solve the stochastic differential equations. Finally, the motion characteristics at each future moment are obtained through iterative solution.

[0078] Furthermore, the form of the stochastic differential equation is: The first... The differential of the motion feature modeled at each time step is calculated by adding two parts: the first part is formed by multiplying the drift term and the differential at that time step, and the second part is formed by multiplying the diffusion term and the increment of the Wiener process.

[0079] In this embodiment, the above stochastic differential equation can be expressed as:

[0080]

[0081] in, For the first Motion features modeled at any given moment; To determine the motion characteristics The differential; The drift term describes the deterministic evolution trend of motion features, and is defined by the first parameterized network. Modeling; The diffusion term, used to control the intensity of random perturbations, is determined by the second parameterized network. Decide; For time The differential; As an increment of the Wiener Process, its randomness closely matches the uncontrollable factors in cloud motion, and can characterize the random uncertainty in motion, such as the random morphological changes of cloud systems and sudden changes in wind fields.

[0082] In this embodiment, when solving the above stochastic differential equation, an initial state is given. The discretization calculation process of future motion characteristics can be expressed as:

[0083]

[0084] in, and These are two adjacent moments; and Motion features modeled for two adjacent time points; For time step; Let be a Gaussian random variable. Standard Gaussian noise is used to ensure that the prediction results can capture the randomness and diversity of cloud motion; It is the identity matrix; It follows a Gaussian distribution.

[0085] Furthermore, the first parameterized network adopts the UNet architecture, using the concatenation result of the current motion features and temporal embedding as input to achieve joint modeling of cloud motion from local details to global trends; the second parameterized network adopts a convolutional network of the same type as the first parameterized network, also using the concatenation result of the current motion features and temporal embedding as input, and adapts to the time-varying uncertainty of cloud motion by learning the perturbation intensity of different time intervals.

[0086] In this embodiment, to effectively parameterize the stochastic differential equation, neural network structures are designed for the drift term and the diffusion term respectively. For example... Figure 5 As shown, the parameterization network for the drift term is implemented using the UNet architecture, based on the current motion features. and time embedding Using concatenation as input, an encoder-decoder structure employing multi-scale downsampling encoding and upsampling decoding captures the multi-scale spatiotemporal dependencies of motion features. It can also capture the multi-scale spatial distribution of motion features (from pixel-level details to region-level trends) and their dynamic temporal dependencies (through temporal embedding and fusion with current motion features), ultimately achieving joint modeling of cloud motion from "local details to global trends."

[0087]

[0088] in, Based on the UNet architecture; The sine and cosine embedding of the time steps transforms the discrete time steps into continuous feature vectors, enhancing the network's ability to perceive time information.

[0089]

[0090] in, For time scale parameters; For the embedded dimension; This is the time scaling factor; These are the sine and cosine functions, respectively. For dimensional indexing.

[0091] The parameterized network for the drift term employs the same type of convolutional network as the drift term. This data-driven diffusion term design allows the model to autonomously learn the intensity of disturbances under different cloud scenarios, ensuring enhanced randomness during periods of intense random motion (such as the development phase of convection) and reduced disturbances during periods of stable motion (such as clear-sky areas).

[0092]

[0093] After the above process, the time series formed by the future moments is And satisfy At that time, the future motion prediction module can output the future motion. The motion characteristics at each moment form a motion characteristic sequence. :

[0094]

[0095] in, That is, the end time of the historical remote sensing cloud image sequence; Each of these is a future moment in the time series. This is the length of the time series; These represent the motion characteristics corresponding to each future moment.

[0096] This motion feature sequence possesses both the rationality of prior knowledge—ensuring that cloud movement conforms to the laws of continuous migration and gradual morphological change through deterministic constraints of the drift term—and random diversity—preserving the diverse variations in local cloud details through dynamic perturbations of the diffusion term. This provides precise dynamic constraints for the subsequent autoregressive diffusion generation module, ensuring that the final generated cloud map conforms to both prior motion laws and possesses natural detail diversity.

[0097] It should be noted that in this invention, the autoregressive diffusion generation module adopts a collaborative design of "autoregressive temporal dependence" and "conditional diffusion generation" to ensure that cloud system movement conforms to the physical laws of continuous migration and gradual morphological change. Through the iterative process of "frame-by-frame generation" and "dynamic guidance," the high-fidelity generation capability of the diffusion model is utilized to ensure the authenticity of cloud image details, while the autoregressive mechanism constrains the continuity of motion between frames, effectively avoiding the ambiguity problem caused by the superposition of multiple results in deterministic prediction. The core idea of ​​this design is to jointly distribute the generated future remote sensing cloud image sequences. It is decomposed into a product of frame-by-frame conditional distributions. By modeling the conditions frame by frame, the complexity of generating long sequences is reduced, while the inter-frame correlation is strengthened.

[0098]

[0099] in, For future remote sensing cloud image sequences; For the future Remote sensing cloud images at any given time, that is, co-generated Remote sensing cloud images of future moments; Indexes of remotely sensed cloud images generated for future moments; For the reason and generate The conditional probability; The first generated for future moments Frame-by-frame remote sensing cloud image; The first generated for future moments Frame remote sensing cloud image, that is, the preceding frame that has already been generated; For the future Motion characteristics at any given moment.

[0100] In the autoregressive diffusion generation module of the present invention, the future time step... The diffusion generation process of the frame remote sensing cloud image is as follows: The future time frame... Frame remote sensing cloud image As a previously generated preceding frame, and with random noise map The data is spliced ​​together to form the initial input of the trained diffusion network, which will then be used to generate the future... Motion characteristics at each time step are used as guiding conditions to generate remote sensing cloud images through reverse diffusion. During the diffusion process, the guiding conditions are injected into each layer of UNet through a cross-layer fusion mechanism. Finally, by progressively removing noise, the next time step is generated. Frame remote sensing cloud image .

[0101] In this invention, the generated preceding frames are spliced ​​with random noise maps, so that the diffusion network is continuously guided by the appearance features of historical frames during the denoising process, avoiding the generation of cloud map structures that are disconnected from the generated preceding frames, such as large-scale shape changes and position changes of the same cloud cluster.

[0102] Furthermore, this invention injects guiding conditions into each layer of UNet through a cross-layer fusion mechanism during the diffusion process, rather than simply concatenating them at the input layer, to ensure that the guiding conditions are integrated throughout the entire feature extraction and reconstruction process. Further, the specific process of injecting guiding conditions through the cross-layer fusion mechanism is as follows: first, for the future... Motion features at each time step are transformed into features of the same dimension as the output of each downsampling layer of UNet through multiple convolutions. This ensures that motion information at different scales matches the features of the corresponding layers in UNet; then, the dimensionality-transformed features are passed through a multilayer perceptron to obtain conditional features. In each downsampling and upsampling layer of UNet, the conditional features are compared with the original features of that layer. Conditional fusion is achieved through cross-layer adaptive group normalization.

[0103] In this embodiment, the specific process of injecting guiding conditions through the aforementioned cross-layer fusion mechanism can be represented as follows:

[0104]

[0105]

[0106]

[0107] in, The number of diffusion steps is Features following time-fusion guidance conditions; For UNet hierarchical indexes; For Adaptive Group Normalization (AdaGN); This is a convolution operation; It is a multilayer perceptron.

[0108] This design allows conditions such as motion trends and appearance content to guide the entire feature extraction and reconstruction process. It guides the model's direction for generating clear cloud maps based on deep global features, constrains the morphological changes in the cloud evolution process based on shallow detailed features, and allows motion features to specifically adjust the cloud generation mode.

[0109] Ultimately, through The next iteration generates a complete sequence of future remote sensing cloud images. This sequence ensures inter-frame continuity through autoregressive dependency, and the generation result strictly follows the guidance of motion features and appearance content by using the cross-layer fusion mechanism of conditional diffusion. At the same time, it retains the uncertainty of cloud map by relying on the randomness of the diffusion process, and realizes the accurate mapping from abstract motion features and previous frame content to concrete cloud map, thus satisfying the physical rationality of motion evolution.

[0110] Additionally, it should be noted that during the training process of the diffusion network, such as... Figure 6 As shown, this is achieved iteratively through two processes: forward diffusion (noise addition) and reverse diffusion (denoising), which integrate external guiding conditions. During the noise addition process, the noise coefficient at future time steps is adjusted according to a preset noise level. Frame-by-frame real remote sensing cloud image Gradually add Gaussian noise plot The number of diffusion steps generated is Noisy remote sensing cloud image :

[0111]

[0112] in, The number of diffusion steps is The cumulative noise figure at that time; The number of diffusion steps is The single-step noise figure at that time; , This represents the maximum number of diffusion steps in the noise-adding process.

[0113] Then, the above noisy remote sensing cloud images With the future moment Frame-by-frame real remote sensing cloud image Concatenation serves as the basic input for each layer. Features after combining fusion guidance conditions Perform single-step noise prediction to obtain the predicted noise. And update the noisy remote sensing cloud image using the reverse diffusion formula:

[0114]

[0115] in, The number of diffusion steps is Noisy remote sensing cloud images at that time; The sampled Gaussian noise; To ensure that a moderate degree of randomness is retained during the generation process to match the uncertainty of cloud motion, the noise standard deviation is used. This is the noise dispatch coefficient.

[0116] Therefore, by integrating the spatiotemporal dynamic coding module, the future motion prediction module, and the autoregressive diffusion generation module, this invention can better understand the spatiotemporal characteristics and motion characteristics of remote sensing cloud image sequences, thereby improving prediction accuracy and clarity.

[0117] The present invention will now demonstrate the application effect of the motion feature-driven image sequence data prediction method described in the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.

[0118] Example

[0119] The specific implementation process of the motion feature-driven image sequence data prediction method used in this embodiment is as described above and will not be repeated here. The following section demonstrates some of the implementation process and results:

[0120] This embodiment uses data from the Advanced Himawari Imager (AHI) of the Himawari-8 / 9 geostationary meteorological satellites of a certain regional meteorological bureau as the data source. The target region is East Asia and the Northwest Pacific, with longitudes ranging from 110°E to 175°E and latitudes from 35°S to 35°N. All raw satellite data were uniformly reprojected to an equal latitude and longitude grid (EPSG:4326), with a spatial resolution of approximately 0.05°.

[0121] To construct data suitable for training the deep learning model, the target region was spatially divided into 256×256 pixel blocks, uniformly divided into 5×5 grids along the latitude and longitude directions, generating a total of 25 sampling areas. Each sampling area covers a geographical range of approximately 12.8°×12.8° (approximately 1300km×1300km). This embodiment uses the visible light band of AHI as the primary input channel, with the band values ​​of the original L1-level data as the main reference, combined with cloud type (CLT) detection data from L2-level cloud products as an auxiliary data source, to perform cloud masking processing. This removes cloudless or low-cloud areas, ensuring that the model learning focuses on the evolution of real cloud systems.

[0122] This embodiment uses a 30-minute time interval for downsampling. Each training sequence contains 12 frames of images, corresponding to a 6-hour time window. All methods are performed under a standard experimental setting of 12-frame sequences and 4 randomly missing frames. All data are z-score normalized. The dataset spans from January 1 to July 31, 2025, and is randomly divided into training and test sets in an 8:2 ratio.

[0123] This embodiment uses mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual image patch similarity (LPIPS) as evaluation metrics to assess the invention.

[0124] Mean squared error is used to calculate the error between predicted image pixels and actual image pixels. Its mathematical formula is as follows:

[0125]

[0126] in, Number of images; This represents the total number of pixels in the image. These are the image's height, width, and number of channels, respectively. For the first The first real image in One channel, Pixel value at the location; For the first Zhang predicted image in the first One channel, The pixel value at the location.

[0127] Peak signal-to-noise ratio (PSNR) is calculated based on mean squared error and is used to measure the overall reconstruction quality of the predicted image. It is defined as follows:

[0128]

[0129] in, This represents the maximum pixel value predicted for the image.

[0130] The structural similarity index comprehensively considers the brightness, contrast, and structural information of an image, and is defined as follows:

[0131]

[0132] in, Indicates the structural similarity between the real image and the predicted image; These are the average brightness values ​​of the real image and the predicted image, respectively. These represent the variances of the real image and the predicted image, respectively. and It is the stability constant; Let be the covariance between the real image and the predicted image.

[0133] Learning to perceive image patch similarity better reflects the subjective perceptual quality of human eyes in terms of details, textures, and boundaries; it is defined as:

[0134]

[0135] in, The learned perceptual image patch similarity between the real image and the predicted image; The first The height and width of the layer feature map; For network layer index; For the first Learnable channel weights for layers are used to balance the importance of features at different levels; This indicates a channel-by-channel weighted average. This indicates that the network extracts the first [image] from the real image. Layer features; Representation of features exist Pixel value at the location; This indicates that the network extracts the first [image] from the predicted image. Layer features; Representation of features exist Pixel value at the location; This represents the square of the L2 norm.

[0136] To verify the effectiveness of the method of this invention, five methods were selected for comparative implementation examples. The implementation methods of these methods all belong to the prior art. Among them, ConvLSTM is a model combining Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) for processing spatiotemporal sequence data; PredRNN is a spatiotemporal sequence prediction model based on Convolutional Recurrent Neural Network (ConvRNN), which builds a deep structure by stacking multiple ConvRNN units to enhance the model's expressive power and prediction performance; E3DLSTM is a deep learning model that integrates 3D convolution and LSTM, using gating mechanisms and self-attention modules to effectively access historical memory; SimVP is a model with a lightweight and efficient pure CNN end-to-end architecture at its core, achieving accurate modeling of inter-frame temporal features through a three-level structure of encoder, temporal translator, and decoder; TAU is an attention unit focused on temporal feature modeling, enhancing the model's ability to capture temporal dependencies in sequence data.

[0137] Table 1. Comparison of experimental results using different methods

[0138] Table 1 shows the performance comparison of different methods. Experiments show that the present invention outperforms the best-performing control model on the dataset in all evaluation metrics. At the pixel accuracy level, MSE, as the core objective metric for pixel-level prediction error, demonstrates that the present invention possesses more accurate pixel-level reconstruction capabilities in remote sensing cloud image sequence prediction. At the visual clarity level, SSIM, as the core metric for characterizing image structural similarity, reflects that the cloud image frames generated by the present invention are more consistent with the real frames in terms of core structural information such as texture details and edge contours. LPIPS, as a similarity metric that aligns with human visual perception, accurately measures the perceptual difference between the generated frame and the real frame at the human visual level; its decrease further verifies that the visual effect of the cloud image frames generated by the present invention is closer to the real scene. Although the present invention shows a slight improvement in PSNR, a comprehensive metric for pixel accuracy and visual clarity, it still maintains an excellent level comparable to the best control model. This result indicates that the present invention achieves synergistic optimization of pixel accuracy and visual clarity, balancing the numerical accuracy of cloud image prediction with visual presentation effects. This invention achieves a more accurate characterization of the dynamic changes in remote sensing cloud image sequences by optimizing the modeling method of time series features and the uncertainty estimation strategy. Ultimately, it surpasses existing benchmark models in various indicators, verifying the effectiveness and advancement of the proposed method.

[0139] In summary, accurate prediction of future remote sensing cloud image sequences has significant application value in many fields such as meteorology and aviation. However, the complexity and uncertainty of cloud motion, including non-stationary motion, generation and dissipation changes, and different evolutionary patterns of various cloud types, often results in traditional deterministic prediction models generating blurred, superimposed images, making it difficult to simultaneously guarantee pixel accuracy and visual clarity. To address these challenges, this invention proposes a novel prediction model based on multi-scale, multi-type cloud motion and stochastic differential equation evolution. This model innovatively integrates three core modules: spatiotemporal dynamic coding, future motion prediction, and autoregressive diffusion generation. The proposed method not only provides a new approach for remote sensing cloud image prediction with higher accuracy and clarity, but more importantly, it demonstrates an effective paradigm that combines motion knowledge-guided feature engineering, uncertainty-aware temporal modeling, and the powerful synthetic capabilities of generative models. This paradigm has universal reference value for predicting Earth observation data with complex spatiotemporal dynamics and inherent randomness.

[0140] It is understood that the motion feature-driven image sequence data prediction method described in the above embodiments can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the motion feature-driven image sequence data prediction method provided in the above embodiments, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, they can implement the motion feature-driven image sequence data prediction method as described in the above embodiments.

[0141] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the motion feature-driven image sequence data prediction method provided in the above embodiments, such as... Figure 7 As shown, it includes a memory and a processor;

[0142] The memory is used to store computer programs;

[0143] The processor is configured to implement a motion feature-driven image sequence data prediction method according to the above embodiments when executing the computer program.

[0144] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0145] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the motion feature-driven image sequence data prediction method provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can implement the motion feature-driven image sequence data prediction method in the above embodiments.

[0146] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0147] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0148] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A motion feature-driven image sequence data prediction method, characterized in that, The process includes the following steps: Inputting a historical remote sensing cloud image sequence into a prediction model comprising a spatiotemporal dynamic coding module, a future motion prediction module, and an autoregressive diffusion generation module: First, the spatiotemporal dynamic coding module extracts multi-scale, multi-type spatiotemporal dynamic contextual features containing cloud system movement patterns from the historical remote sensing cloud image sequence, providing high-quality initial contextual information for the future motion prediction module; then, the future motion prediction module, based on the spatiotemporal dynamic contextual features, predicts the evolution of cloud system movement characteristics at future times, models the time dependence and uncertainty of cloud movement, and outputs the motion characteristics at future times; finally, the autoregressive diffusion generation module uses the last frame of the historical remote sensing cloud image sequence as initial input, employs autoregressive logic guided by previous frames, combines the motion characteristics at future times to construct guiding conditions, and generates remote sensing cloud images for each future time frame by frame, thereby predicting the future remote sensing cloud image sequence. The internal processing flow of the spatiotemporal dynamic coding module is as follows: S11. By performing differential processing on the input historical remote sensing cloud image sequence, the difference image between adjacent frames is calculated to enhance the saliency of motion information; S12. The obtained difference image is then processed through a multi-branch convolutional structure to extract spatial features in order to capture motion features at different scales; S13. Adaptive weighted fusion of multi-scale motion features is achieved through an attention mechanism to obtain initial fusion features that integrate cloud morphology and motion details from micro to macro. S14. Obtain the cloud type label sequence corresponding to the historical remote sensing cloud image sequence, and convert the discrete labels in the cloud type label sequence into continuous feature vectors through the embedding layer to obtain cloud type features; The cloud type features are then merged with the initial fusion features through feature concatenation to obtain the final fusion features; S15. Convolutional gated recurrent units are used to temporally encode the final fused features at each historical moment to model the temporal dependency of cloud motion, thereby obtaining the motion features after temporal fusion, and using the set of motion features as spatiotemporal dynamic context features.

2. The motion feature-driven image sequence data prediction method as described in claim 1, characterized in that, The internal processing flow of the future motion prediction module is as follows: stochastic differential equations are used to model the continuous evolution process of motion characteristics at future moments, and Euler numerical integration method is used for discretization calculation to solve the stochastic differential equations. Finally, the motion characteristics at each future moment are obtained through iterative solution.

3. The motion feature-driven image sequence data prediction method as described in claim 2, characterized in that, The form of the stochastic differential equation is as follows: the differential of the motion feature modeled at time t is obtained by adding two parts. The first part is formed by multiplying the drift term and the differential at that time, and the second part is formed by multiplying the diffusion term and the increment of the Wiener process.

4. The motion feature-driven image sequence data prediction method as described in claim 3, characterized in that, Both the drift term and the diffusion term are implemented using parametric networks. The parametric network for the drift term adopts the UNet architecture, taking the concatenation result of the current motion features and the temporal embedding as input to achieve joint modeling of cloud motion from local details to global trends. The parametric network for the diffusion term adopts a convolutional network of the same type as the parametric network for the drift term, also taking the concatenation result of the current motion features and the temporal embedding as input, and adapting to the time-varying uncertainty of cloud motion by learning the perturbation intensity of different time intervals.

5. The motion feature-driven image sequence data prediction method as described in claim 1, characterized in that, In the autoregressive diffusion generation module, the future time step... The diffusion generation process of the frame remote sensing cloud image is as follows: The future time frame... The remote sensing cloud image of the first frame, as a preceding frame already generated, is stitched together with a random noise image to form the initial input of the trained diffusion network, which will then be used to generate the next frame. Motion characteristics at each time step are used as guiding conditions to generate remote sensing cloud images through reverse diffusion. During the diffusion process, the guiding conditions are injected into each layer of UNet through a cross-layer fusion mechanism. Finally, by progressively removing noise, the next time step is generated. Frame-by-frame remote sensing cloud image.

6. The motion feature-driven image sequence data prediction method as described in claim 5, characterized in that, The specific process of injecting guiding conditions through the cross-layer fusion mechanism is as follows: First, for the future... The motion features at each moment are transformed into features of the same dimension as the output of each downsampling layer of UNet through multiple convolutions; The dimensionality-transformed features are then passed through a multilayer perceptron to obtain conditional features. In each downsampling and upsampling layer of UNet, the conditional features and the original features of that layer are fused through cross-layer adaptive group normalization.

7. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they can implement the motion feature-driven image sequence data prediction method as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the motion feature-driven image sequence data prediction method as described in any one of claims 1 to 6.

9. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the motion feature-driven image sequence data prediction method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Agricultural meteorological prediction method based on multi-scale convolutional network and diffusion system

    CN119047661A