Generative AI-based high-frequency long-time-sequence short-imminent forecasting method and system
By employing a generative AI-based high-frequency, long-term short-term forecasting method, and utilizing causal variational autoencoders and diffusion transformers to process radar reflectivity data, the predictability barrier in short-term forecasting is overcome, enabling accurate precipitation forecasts with high frequency and long lead times, and making it suitable for digital twins of the Earth system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI TYPHOON INST OF CHINA METEOROLOGICAL ADMINISTRATION (SHANGHAI INST OF METEOROLOGICAL SCI)
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies face predictive limitations in short-term forecasting of extreme precipitation within 2-6 hours. Traditional methods fail due to the accumulation of nonlinear errors or are limited by computational efficiency, while AI methods are prone to model collapse or generating false echoes, making it impossible to accurately predict the structural evolution of storm systems.
We employ a high-frequency, sub-long-term short-term forecasting method based on generative AI. We use a pre-trained causal variational autoencoder to compress the radar reflectivity sequence into a latent variable token sequence, and generate the predicted radar reflectivity sequence for future time periods through a diffusion Transformer. We utilize the diffusion generative model to learn the kinematic and dynamic causal chain of the atmosphere in the latent space, avoiding mode collapse and maintaining the sharpness of the forecast image.
It achieves high-resolution, long-term rainfall forecasting, improves the accuracy of short-term forecasts, can stably explore data distribution within the gray zone, spontaneously learns atmospheric physical laws, maintains the sharpness of extreme weather characteristics, and is suitable for building digital twins of the Earth system.
Smart Images

Figure CN121980183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of precipitation prediction technology, and in particular to a high-frequency, long-term, short-term forecasting method and system based on generative AI. Background Technology
[0002] Accurate short-term forecasts and early warnings for extreme precipitation are crucial for global disaster prevention and mitigation. However, within the 2-6 hour forecast lead timeframe, a formidable "predictability barrier"—the so-called "short-term forecast gray zone"—has long existed. Within this time window, traditional observation-based extrapolation methods fail due to the accumulation of nonlinear errors, while numerical weather prediction (NWP), limited by computational efficiency, often cannot resolve the storm-scale dynamic characteristics before the storm occurs.
[0003] Despite recent advancements in AI short-term methods, they still fall short of fully addressing the problem. Deterministic models, aiming to minimize the mean error (MSE), are systematically constrained by the "regression-to-mean" phenomenon, leading to blurred forecast images and loss of extrema. Conversely, Generative Adversarial Networks (GANs), while preserving image sharpness, are prone to "mode collapse" and produce operationally unacceptable "hallucinatory echoes." To mitigate these issues, recent decomposition paradigms (such as NowcastNet, DiffCast, Cascast, and AlphaPre) attempt to explicitly separate weather evolution into additive components: deterministic optical flow for advection and stochastic residuals for detail. However, this artificial separation is fundamentally flawed because it severs the tightly intertwined causal chains of atmospheric processes. In the real atmosphere, kinematics (such as movement) and dynamics (such as growth) are coupled together through thermodynamic feedback loops; for example, the kinematic advection at the outflow boundary formed by rain cooling is often the direct dynamic mechanism that triggers the generation of new convection. A model that structurally separates "motion" from "growth" is blind to this mechanism. If extrapolation is performed within the "gray zone," this lack of ability to model coupled evolution will directly lead to the structural disintegration of the storm system.
[0004] Therefore, it is necessary to explore a precise and efficient method for short-term precipitation forecasting. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for high-frequency, sub-long-term, short-term forecasting based on generative AI.
[0006] To achieve the above objectives, in a first aspect, this invention provides a high-frequency, long-term time-series short-term forecasting method based on generative AI. The method includes the following steps: compressing the original radar reflectivity sequence into a compact latent variable token sequence using a pre-trained variational autoencoder; within the compressed latent space, using the latent variable token sequence as a condition, progressively generating a predicted latent variable token sequence for future time periods through probabilistic generation; and reconstructing the predicted latent variable token sequence into a predicted radar reflectivity sequence for future time periods using the decoder of the variational autoencoder. This invention achieves high-resolution, long-term radar nowcasting, improves the accuracy of rainfall forecasting within the short-term forecast gray zone time window, and lays a solid foundation for constructing a digital twin of the Earth system.
[0007] Optionally, the variational autoencoder is a causal variational autoencoder.
[0008] Optionally, the causal variational autoencoder uses causal 3D convolution.
[0009] Optionally, the predicted latent variable token sequence is generated using a diffusion generation model based on the noisy token sequence, with the latent variable token sequence as a condition.
[0010] Optionally, the noise token sequence is obtained by random sampling from Gaussian noise.
[0011] Optionally, the number of frames in the noise token sequence is consistent with the length and time step of the predicted radar reflectivity sequence.
[0012] Optionally, the diffusion generation model is based on the Rectified Flow framework and is constructed using a diffusion Transformer.
[0013] Optionally, the diffusion Transformer uses a 3D causal self-attention mechanism to simultaneously capture local advection and non-local convection triggers.
[0014] Optionally, the diffusion Transformer injects time coordinates into each layer of the network through adaptive layer normalization.
[0015] Secondly, the present invention provides a high-frequency sub-long-term time series short-term forecasting system based on generative AI. The high-frequency sub-long-term time series short-term forecasting system based on generative AI includes: a data input device, a data output device, a processor, and a storage device. The storage device includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor, cause the processor to implement the high-frequency sub-long-term time series short-term forecasting method based on generative AI provided by the present invention.
[0016] In summary, this method has at least the following beneficial effects: 1. To address the problem that numerical weather prediction models fail in the gray region due to the accumulation of nonlinear errors and are plagued by the "cold start" effect, based on observational extrapolation, this method maps the forecasting problem onto a compressed latent space manifold. By processing the data through a diffusion Transformer, it can better handle the complex atmospheric conditions in the gray region and achieve effective forecasts for long time series and high frequency.
[0017] 2. Unlike deterministic models that aim to minimize mean error, resulting in "mean regression" and image blurring, this method's hidden manifold processing can better maintain the details and structure of the data, enabling the forecast image to clearly present relevant information such as precipitation and retain the sharpness of key features such as extreme weather.
[0018] 3. This method adopts a probabilistic diffusion process. This diffusion generation method has its own stability and diversity characteristics. Unlike GAN, which is prone to getting stuck in local optima during training and causing pattern collapse, the diffusion generation model generates data by gradually adding and removing noise, which can explore the data distribution more stably, thereby avoiding pattern collapse and ensuring generation quality.
[0019] 5. Unlike the decomposition paradigm that explicitly separates weather evolution into additive components, this method employs a global attention mechanism in the Diffusion Transformer, which allows the model to focus on long-distance dependencies in the data, thereby learning the nonlocal causal rules of convection. It does not rely on artificial decomposition assumptions, but spontaneously learns the physical laws governing the atmosphere, fully models the causal chain in which kinematics and dynamics are closely intertwined, and improves the accuracy of prediction.
[0020] 6. The model provided by this method, after being trained on large-scale data, can spontaneously learn the physical laws governing the atmosphere. It has both the flexibility of data-driven methods and the ability to simulate physical processes. Therefore, it can become an effective tool for connecting data-driven extrapolation and physical simulation, laying a solid foundation for building a digital twin of the Earth system.
[0021] 7. A system adapted to the method is provided, which not only improves the practicality of the method, but also facilitates its promotion. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the high-frequency, long-term, short-term forecasting method based on generative AI according to an embodiment of the present invention. Figure 2 This is a phase space visualization diagram of an embodiment of the present invention; Figure 3 These are the experimental results of the StormDiT stability verification experiment according to an embodiment of the present invention; Figure 4 This presents experimental results of the rapid generation and organization of squall lines using StormDiT, as described in this embodiment of the invention. Figure 5 This is an experimental result of StormDiT analyzing the dissipation of typhoon spiral rainbands in an embodiment of the present invention; Figure 6 The above are the prediction results of different prediction models of the present invention for the generation of multi-monitor convection at different forecast lead times; Figure 7 This is a comparison chart of the ensemble forecast results of StormDiT and DiffCast in an embodiment of the present invention; Figure 8 This is a comparison chart of the prediction results of StormDiT and DiffCast in a 3-hour forecasting task according to an embodiment of the present invention; Figure 9 The prediction results of different prediction models for squall line evolution in embodiments of the present invention are shown. Figure 10 The following is an example of the performance of DiffCast and StormDiT in a structure ablation experiment during a self-sustaining squall line event according to embodiments of the present invention. Figure 11 This is a schematic diagram of the framework of a high-frequency, long-term, short-term forecasting system based on generative AI according to an embodiment of the present invention. Detailed Implementation
[0024] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0025] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0026] It should be noted in advance that, in one alternative embodiment, except for independent descriptions, the same symbols or letters appearing in all formulas have the same meaning.
[0027] In one optional embodiment, please refer to Figure 1 This invention provides a high-frequency, sub-long-term short-term forecasting method based on generative AI. The method is implemented using a generative AI prediction model called StormDiT, which includes an encoder, a diffusion generation part, and a decoder. StormDiT is a unified generative framework designed to model the entangled kinematics and dynamics of precipitation as a holistic video generation problem. Its core design concept abandons traditional pixel-space regression, instead operating entirely within a compressed latent manifold. The method includes the following steps: S1. The original radar reflectivity sequence is compressed into a compact latent variable token sequence using a pre-trained variational autoencoder.
[0028] Specifically, in this embodiment, this step is implemented through the encoder portion of StormDiT. The original radar reflectivity data sequence is a sequence of visualized images of the original radar reflectivity data, where each visualized image has a width of W and a height of H. The variational autoencoder is specifically a causal variational autoencoder (Causal VAE), and the ordinary convolutional layers in the Causal VAE are replaced with causal 3D convolutions to strictly enforce the temporal directionality and prevent future information from leaking into the historical context token.
[0029] Raw radar reflectivity data is sparse, meaning that statistically, the vast majority of it consists of non-precipitation areas. This dilutes the model's learning signal for severe weather. Directly modeling dynamics within the spatial representation of raw radar reflectivity data forces the model to waste computational resources characterizing redundant zeros and background noise, rather than focusing on the structural evolution of high-impact convective cores. Therefore, this embodiment inputs the raw radar reflectivity sequence into a pre-trained Causal VAE, compressing the raw radar reflectivity sequence into a compact latent variable token sequence. This maps the forecasting problem onto the compressed latent space manifold, laying a solid foundation for achieving effective long-term and high-frequency forecasts. The Causal VAE uses an encoder pre-trained with a large number of video and image samples from the field of computer video generation.
[0030] S2. Within the compressed latent space, using the latent variable Token sequence as a condition, a predictive latent variable Token sequence for future time periods is gradually generated through probability generation.
[0031] Specifically, in this embodiment, this step is implemented through the diffusion generation part of StormDiT, specifically through a diffusion generation model.
[0032] A single historical precipitation pattern can evolve into multiple plausible future precipitation patterns. Traditional deterministic models are trained by minimizing the mean-variance (MSE) to predict future precipitation patterns. Essentially, they predict the conditional expectations of these future precipitation patterns, leading to the phenomenon of "fuzzy means," such as... Figure 2 As shown in the blue halo, extreme values are suppressed, and sharp convective structures are smoothed out; this is known as the mean reversion limitation. Furthermore, in precipitation forecasting, traditional diffusion methods typically traverse to winding, random paths, such as... Figure 2 As shown by the red dashed arrow in the image, this requires a large number of sampling steps to resolve.
[0033] To fundamentally address this limitation, this embodiment builds a diffusion-generative model based on the Rectified Flow framework and using Diffusion Transformer (DiT). This model then generates a predicted latent variable token sequence based on the noisy token sequence, using the latent variable token sequence as a condition. Specifically, this model employs the Rectified Flow framework, no longer predicting a single point estimate, but instead learning a velocity field in the latent space. This velocity field represents a simple Gaussian noise distribution. (representing initial uncertainty) Figure 2 The gray dots in the image are mapped to a complex distribution of actual precipitation patterns. (Physical weather conditions, Figure 2(The blue dots in the diagram). During model training, conditioned on the historical latent variable token sequence, Rectified Flow enforces a straight "Optimal Transport" trajectory, such as... Figure 2 As shown by the green straight arrow, the DiT backbone network learns the linear transmission trajectory from Gaussian noise to the future predicted latent variable token sequence. This linear coupling enables the model to capture the complete multimodal distribution, retain the sharpness of extreme weather, and significantly improve the sampling efficiency and stability during inference. When using the model for prediction, it starts from the noisy token sequence, using the historical latent variable token sequence as a condition, and gradually generates the predicted latent variable token sequence for future time periods by solving the ordinary differential equation (ODE) of the Rectified Flow. This diffusion generation method enables the model to explore the data distribution more stably, thereby avoiding mode collapse and ensuring generation quality.
[0034] More specifically, the noise token sequence is randomly sampled from Gaussian noise, and the number of frames in the noise token sequence is consistent with the length and time step of the predicted radar reflectivity sequence. DiT employs a 3D causal self-attention mechanism to simultaneously capture local advection and non-local convection triggers, allowing the model to dynamically adjust its attention mode: on the one hand, it focuses on local gradients to simulate short-term advection (kinematics), and on the other hand, it accesses non-local historical precursors to resolve long-term convection triggers (dynamics). This enables the diffusion-generative model to learn the non-local causal rules of convection, without relying on artificial decomposition assumptions, spontaneously learning the physical laws governing the atmosphere, and fully modeling the causal chain in which kinematics and dynamics are closely intertwined in the atmosphere, thus improving the accuracy of predictions. To rigorously execute the continuous-time dynamics required by the Rectified Flow ordinary differential equations, DiT employs an adaptive layer normalization (adaLN) mechanism to inject time coordinates into each layer of the network for regressing dimensional scaling and offset parameters, thereby effectively modulating the normalized feature statistics and guiding the latent state to evolve along a precise flow trajectory from noise to coherent forecasts.
[0035] S3. The decoder of the variational autoencoder is used to reconstruct the predicted latent variable Token sequence into a predicted radar reflectivity sequence for a future time period.
[0036] Specifically, in this embodiment, this step is implemented through the decoder part of StormDiT. The predicted latent variable token sequence output by the diffusion generation model is input into the decoder of Causal VAE to obtain and output the predicted radar reflectivity sequence for future time periods, reconstructed from the predicted latent variable token sequence.
[0037] The advantages of this method compared to existing precipitation forecasting methods will be verified through specific experimental data in the following sections.
[0038] StormDiT was trained on a Chinese radar composite reflectivity dataset from August 2022 to the end of December 2024. The temporal resolution was 6 minutes. The training process involved 40,000 iterations.
[0039] Specifically, in this embodiment, a large-scale radar reflectivity dataset collected throughout 2025 is used to evaluate the stability of StormDiT in long-term forecasts. This dataset contains 2624 consecutive precipitation sequences sampled from different meteorological regions. The final experimental results are as follows: Figure 3 As shown.
[0040] Figure 3 Figure (b) shows the evolution curve of the Critical Success Index (CSI) with forecast lead time. It can be seen that StormDiT exhibits stable performance curves at radar reflectivities of 5 dBZ, 15 dBZ, 25 dBZ, 35 dBZ, and 45 dBZ, without the catastrophic collapse common in long-lead-time forecasts. Crucially, for strong convection at 35 dBZ, StormDiT maintains a CSI above 0.1 throughout the entire window (0–6 h), effectively extending the reliable lead time for early warning. Furthermore, Figure 3 (a) shows that even at extreme convection intensities (45 dBZ), the average CSI of StormDiT is 0.03, indicating that StormDiT remains robust at this point. This suggests that StormDiT successfully preserves rare and high-value precipitation events when uncertainty increases.
[0041] Specifically, in this embodiment, to rigorously evaluate StormDiT's ability to resolve complex structural evolution, a detailed qualitative analysis was conducted on specific high-impact events. Two distinct meteorological flow patterns representing the two ends of the convective lifecycle were selected to test the versatility of StormDiT: the rapid formation and organization of squall lines, and the dissipation of typhoon spiral rainbands. The final test results are as follows: Figure 4 and Figure 5 As shown.
[0042] For the generation process of a severe squall line on March 13, 2025, based on the StormDiT prediction results, the following were obtained sequentially: spatial scale fidelity map (power spectral density PSD as a function of wavelength), reflectivity distribution map (distribution of observed and predicted radar reflectivity), forecast skill map (the critical success index as a function of forecast lead time at Z values of 15 dBZ, 25 dBZ, and 35 dBZ), observation-prediction comparison map at different times (visualization of observed and predicted Z values at T+1h, T+2h, T+3h, T+4h, T+5h, and T+6h, where T is the forecast start time), and magnified views of the observation-prediction comparison map at T+2h and T+6h. These are as follows: Figure 4 (a) in Figure 4 As shown in (e) in the diagram. From Figure 4 As can be seen in (a), the power spectral density of the predicted result closely tracks the power spectral density of the observed result, confirming that StormDiT preserves high-frequency information. From Figure 4 As can be seen in (b), the distribution of observed and predicted radar reflectivity is highly similar, proving the high accuracy of the StormDiT prediction results. Figure 4 As can be seen from (c), the critical success index of StormDiT exhibits a stable performance curve under different radar reflectivities, without the catastrophic collapse commonly seen in long-term predictions. From Figure 4 As shown in (d), StormDiT successfully predicted the nonlinear transition of the gust system into an integrated arc-shaped structure over the following 3 hours. At T+6h, StormDiT accurately delineated the sharp reflectivity gradient at the leading edge of the gust system, a hallmark of a mature "bow echo" structure. This structural fidelity is clearly emphasized in the magnified details at T+3h and T+6h, specifically as follows... Figure 4 As shown in (e), unlike deterministic baselines that typically smooth high-frequency features to minimize errors, StormDiT preserves the structural integrity of the “hook” and “bow” regions that are crucial for identifying potential strong gusts.
[0043] Furthermore, for the dissipation phase of Typhoon Co-May, based on the StormDiT forecast results, spatial scale fidelity maps, reflectance distribution maps, forecasting skill maps, comparison maps of observations and forecasts at different times, and locally enlarged images of the comparison maps of observations and forecasts at T+2h and T+6h times were sequentially obtained. Forecasting this flow pattern faces a dual challenge: resolving the kinematic cyclonic rotation while simultaneously capturing the thermodynamic dissipation of the spiral rainbands. From Figure 5As can be seen from (d) in the figure, the forecast results accurately reproduce the observed process of the outer rainband breaking into discrete single structures as the typhoon system moved northwest. Figure 5 (e) shows that StormDiT avoids artificial retention or blurring artifacts, correctly reduces the area of convection, and preserves the fine-grained texture of the residual echo. Meanwhile, Figure 5 (b) also confirms that StormDiT correctly simulated the intensity decay of Typhoon Co-May and did not suffer from "model collapse" because the distribution of observed and predicted radar reflectivity closely matches above 50 dBZ. Furthermore, Figure 5 (a) shows that the power spectral density of the predicted results closely follows the power spectral density of the observed results, confirming that StormDiT preserves high-frequency information; Figure 5 (c) shows that StormDiT’s critical success index exhibits a stable performance curve under different radar reflectivities, without the catastrophic collapse commonly seen in long-term predictions.
[0044] Specifically, in this embodiment, in order to verify the generalization ability of StormDiT, CSI, structural similarity index (SSIM) and mean squared error (MSE) were used as evaluation metrics. On the standard SEVIR (Storm EVent ImagRy) dataset, it was benchmarked against a series of leading models such as ConvGRU, Earthformer, DiffCast and AlphaPre. The evaluation metric values of each model under the standard 1-hour forecast are shown in Table 1.
[0045] Table 1 Performance Comparison of Different Models In this text, "↑" indicates that a larger value indicates better model performance, and "↓" indicates that a smaller value indicates better model performance. CSI-M is the average CSI, and CSI-181 is the vertical integral liquid water content greater than 181 kg / m³. 2 The CSI of the time model, CSI-219, is for a vertical integral liquid water content greater than 219 kg / m³. 2 The CSI of the time-integrated model is shown in Table 1. It is clear from Table 1 that StormDiT's overall performance is significantly better than other models. While AlphaPre achieved a slightly higher CSI-M, likely due to its optimization for the average operating condition, StormDiT achieved a transformative performance leap in the long-tail region, specifically for vertical integral liquid water content greater than 219 kg / m³. 2For this extreme precipitation event, StormDiT achieved a CSI of 0.1301, a 140% improvement over AlphaPre's 0.0545, and more than double that of DiffCastd. This demonstrates that StormDiT is inherently better suited for capturing the rapid intensification of high-impact storms than methods limited to decomposition or regression smoothing. Furthermore, this embodiment showcases this advantage in a complex case of multicell convective growth, specifically as follows... Figure 6 As shown. Figure 6 The top row of images shows the actual evolution of multicell convection at different times. Using the images corresponding to -60min, -40min, -20min, and 0min as input, ConvLSTM, PredRNN, Earthformer, DiffCast, and StormDiT were used sequentially to predict radar reflectivity data corresponding to 10min, 20min, 30min, 40min, 50min, and 60min. It is easy to see that the prediction results of ConvLSTM, PredRNN, and Earthformer exhibit typical "spectral smoothing," correctly pinpointing the location of the storm system but systematically erasing the high-intensity core. While DiffCast improves the prediction of core intensity, it suffers from "mode dropping," completely ignoring the rapid development of a primary secondary cell in the upper right quadrant. In stark contrast, StormDiT accurately resolved the complete multi-cell dynamics, not only preserving the intensity of the main storm but also successfully predicting the triggering and growth of secondary cells. This demonstrates that DiT's global attention mechanism effectively captures the multi-scale interactions required to resolve the development of synchronous convection.
[0046] Specifically, in this embodiment, precipitation nowcasting has inherent randomness, especially for high-impact events governed by chaotic dynamics. A core advantage of the diffusion Transformer architecture is its ability to inherently simulate a complete conditional probability distribution, rather than simply outputting a deterministic point estimate. To quantify this capability, the entire SEVIR test set was sorted according to peak vertical integral liquid water content (VIL) intensity, and the top 200 strongest precipitation events were selected as the basis for generating 1-hour ensemble forecasts for StormDiT and DiffCast (containing predictions at 12 time points with a 5-minute interval between adjacent time points). Then, based on the ensemble forecasts from StormDiT and DiffCast, the variation curves of the dispersion-skill ratio (SSR), the continuous graded probability score (CRPS) of VIL, and the mean absolute error (MAE) of VIL with forecast lead time were obtained, as shown in the figure below. Figure 7 As shown in (a) to (c). Meanwhile, for a cyclone system with strong rotational characteristics among the 200 strongest precipitation events, the ensemble mean of its prediction results was calculated at forecast lead times of 5 min, 10 min, 20 min, 30 min, 40 min, 50 min, 55 min, and 60 min. MAE and the standard deviation of the prediction results The statistical gains, such as those mentioned above, were visualized to diagnose how the model characterizes dynamic risk. The results are as follows: Figure 7 As shown in (d) in the figure.
[0047] A key requirement for operational forecast trust is calibration, which this embodiment quantifies using the dispersion-skill ratio (SSR). Dispersion refers to the degree of dispersion among ensemble forecast members. An SSR value of 1.0 indicates that the dispersion accurately reflects the forecast error. Figure 7 As shown in (a), DiffCast exhibits a dispersion much greater than 1 from 10 min to 60 min, indicating that it tends to generate an excessively wide range of possibilities to encompass the true value. In contrast, StormDiT maintains an SSR (mean SSR) very close to 1 throughout the entire forecast lead time. 0.96). This demonstrates that StormDiT captures the inherent random uncertainties of the atmosphere with high fidelity, achieving a balance between diversity and accuracy. This advantage of StormDiT lies in... Figure 7 This is further confirmed by the CRPS variation curve shown in (b) of the diagram, where StormDiT consistently improves probability accuracy by 10-15% compared to the baseline. Furthermore, from... Figure 7 As shown in (c), the MAE curves indicate that the MAE of both StormDiT and DiffCast increases with time. However, the MAE curve of StormDiT is lower than that of DiffCast for most forecast lead times. This suggests that StormDiT has a smaller average difference between the predicted and observed results and higher prediction accuracy under different forecast lead times.
[0048] like Figure 7 As shown in (d) of the diagram, the physical validity of the statistical gain becomes apparent when the spatial structure of the forecast ensemble is visualized. It is easy to see that the ensemble average of DiffCast ( After 40 minutes, the spectrum smoothed, leading to a gradual attenuation of peak intensity within the high-intensity echo core, a problem significantly mitigated by StormDiT. Secondly, regarding the spatial distribution of standard deviation, DiffCast generated a relatively broad uncertainty field, roughly covering the convective region. In contrast, StormDiT exhibited a highly structured uncertainty distribution, with its standard deviation precisely distributed along the system's dynamic gradient, concentrated at the edges of the spiral rainband and the convective core. This flow-pattern-dependent model indicates that StormDiT correlates uncertainty with specific precipitation structures, providing a refined risk assessment consistent with underlying fluid dynamics. Finally, it can be observed that the magnitude of MAE varies across different precipitation regions. For example, in some convective core regions, the MAE value may be relatively high, indicating greater difficulty in prediction and larger discrepancies between predicted and observed results; while in some relatively stable precipitation regions, the MAE value may be lower. Comparing the MAE visualizations of StormDiT and DiffCast reveals that StormDiT has a relatively smaller MAE value in most areas, especially in key convection regions, further demonstrating StormDiT's advantage in precipitation forecast accuracy.
[0049] Specifically, in this embodiment, while the 1-hour forecast mission demonstrated the superior performance of StormDiT, the "gray zone" (2-6 hours) forecast is the main challenge. To test this capability, a new and more challenging 3-hour (13-hour) forecast was established on SEVIR using StormDiT's long lead time capability. A 36-frame forecast baseline was established and compared with the forecast results from DiffCast. Ultimately, the VIL was obtained as 16 kg / m³ at resolutions of 1 km, 4 km, and 16 km. 2 74kg / m 2 133kg / m 2 160kg / m 2 181kg / m 2 and 219kg / m 2 CSI for StormDiT and DiffCast, and plotted as follows Figure 8 The bar charts shown in (a), (b), and (c) are illustrated. Simultaneously, the CSI curves of StormDiT and DiffCast over time were obtained at resolutions of 1 km, 4 km, and 16 km, with radar reflectivities of 74 dBZ, 133 dBZ, and 181 dBZ, respectively, as shown in the figures below. Figure 8As shown in (d), (e), and (f), the curves of the score skill score (FSS) of StormDiT and DiffCast (primarily used to measure the reliability of spatial alignment between predictions and observations; a higher value indicates better spatial alignment) over time were also obtained, as shown in the figures below. Figure 8 As shown in (g), (h) and (i). Figure 8 In this context, images in the same column have the same resolution.
[0050] from Figure 8 As shown in (a), StormDiT achieved the highest CSI across all VILs at a resolution of 1 km. This means that at a relatively fine spatial resolution, StormDiT can predict precipitation more accurately, and from... Figure 8 As can be seen from (b) and (c), this advantage is further consolidated at larger resolutions (4 km and 16 km), indicating that StormDiT is not only able to handle fine-scale precipitation forecasts, but also has a stronger ability to resolve the characteristics of weather systems at different scales in meteorology at larger scales, which translates into significant index gains, demonstrating that StormDiT has structural superiority at different spatial scales.
[0051] from Figure 8 As can be seen from (d) to (i), both StormDiT and DiffCast exhibit the expected performance degradation over time; that is, as the forecast lead time lengthens, both CSI and FSS generally show a downward trend. This is because the difficulty and uncertainty of forecasting increase over time. However, StormDiT consistently maintains higher CSI and FSS than DiffCast throughout the entire 180-minute window. This difference is particularly pronounced in the FSS index. Even during high-intensity events (133 kg / m²), the difference remains significant. 2 At the 3-hour (180-minute) node, StormDiT still retained considerable predictive power. This indicates that the generation process of StormDiT can effectively maintain the structural integrity of the precipitation field, without the problems of rapid dissipation or displacement of the precipitation field as seen in common long-term forecasts, further confirming the robustness of StormDiT.
[0052] Specifically, in this embodiment, Figure 8 This quantitative robustness demonstrates StormDiT's exceptional mastery of entanglement physics and dynamics, just as... Figure 9 As clearly demonstrated by the qualitative contrast. Figure 9The top row of images illustrates the actual evolution of a typical self-sustaining squall line from -60 min to +180 min. This is a system whose evolution is constrained by nonlocal, entangled physics defining the "gray zone": new, intense convection (the dynamic core) is constantly triggered by the system's own propagating outward flow boundary (kinematic origin). Using the images corresponding to -60 min, -40 min, -20 min, and 0 min as input, ConvLSTM, PredRNN, and Earthforme are used to predict radar reflectivity data at 5 min, 20 min, 40 min, and 60 min, respectively. DiffCast and StormDiT are used to predict radar reflectivity data at 5 min, 20 min, 40 min, 60 min, 80 min, 100 min, 120 min, 140 min, 160 min, and 180 min.
[0053] from Figure 9 It is evident that ConvLSTM, PredRNN, and Earthforme succumb to "deterministic failure" (fuzziness). Within the first 60 minutes, they average the sharp linear convective boundary into non-physical, dissipating clumps, becoming blurred and losing their original linear characteristics. This results in predictions that differ significantly from the clear squall line structure observed in actual observations. This means that ConvLSTM, PredRNN, and Earthforme cannot accurately capture the dynamic changes of weather systems with clear boundary structures like squall lines, leading to predictions that do not conform to actual physical laws. DiffCast, on the other hand, suffers from "causal discontinuity failure." While attempting to transport linear structures via advection, it fails to capture the dynamic formation process along the boundary. After 80 minutes, the squall line structure fatally breaks down and dissipates; the originally continuous linear structure becomes fragmented and gradually dissipates, unable to maintain the complete shape of the squall line. This indicates that while DiffCast can handle the advection motion of weather systems to some extent, it cannot accurately simulate the crucial dynamic process of new convective cell formation in squall line systems. Compared to ConvLSTM, PredRNN, Earthforme, and DiffCast, StormDiT's predicted squall line maintains a clear linear structure throughout the entire 180-minute prediction period, with new high-intensity convective cells continuously forming along the squall line front, closely matching the actual squall line evolution. This demonstrates that StormDiT can effectively capture the dynamic evolution of squall line systems, including propagation and the formation of new convective cells. It proves that StormDiT can spontaneously learn the complex, non-local physical rules governing convective evolution, successfully overcoming the 3-hour forecast challenge that other architectures have failed to achieve.
[0054] Specifically, in this embodiment, the performance of DiffCast and StormDiT in self-sustaining squall line events was compared through structural ablation experiments at prediction lead times of 20 min, 50 min, 70 min, 80 min, 90 min, and 100 min. The results are as follows: Figure 10 As shown. The deterministic backbone of DiffCast ( The DiffCast algorithm is responsible for simulating physical advection, but it exhibits "spectral smoothing" during forecasting. This means it erases the high-frequency gradient information necessary to maintain convective organization, much like how over-smoothing in image processing leads to loss of detail, thus destroying key features of the convective structure. The residual components of DiffCast ( Due to the lack of a coherent structure to guide it, it degenerates into unanchored random noise. This indicates that the residual component in DiffCast cannot reasonably supplement and correct the changes in the squall line system when the deterministic backbone loses key information. When the prediction results of the DiffCast deterministic backbone and residual components are recombined, the squall line system undergoes structural disintegration after 80 minutes. Figure 10 (Image from top to bottom, 4th row) This is because DiffCast failed to capture the nonlinear feedback loop required for storm sustaining, and thus could not accurately simulate the complex dynamic evolution of the squall line system.
[0055] In stark contrast to DiffCast, StormDiT maintained the structural integrity of the squall line throughout the entire 100-minute forecast and accurately predicted the new convective cells continuously forming along its leading edge. This demonstrates the significant advantage of the StormDiT model in handling long-term forecasts of complex weather systems like squall lines.
[0056] To verify that StormDiT's ability stems from genuine learning of entanglement physics rather than pattern memorization, this embodiment visualizes the attention weights of the 21st DiT block (wireframe). Figure 10 The image (bottom row) focuses on severe weather signals driven by complex multi-scale dynamics, namely the "bow echo" region. The image shows that StormDiT does not merely extrapolate the current location of the storm; instead, it significantly focuses on the non-local environment surrounding the convective core. Physically, the high-attention region directly maps to known dynamic precursors: the model focuses on the immediate front of the storm, corresponding to the propagating outflow boundary (or gust front), where mechanical lifting triggers new cell formation. Simultaneously, it focuses on the trailing stratiform cloud region, consistent with the back-side inflow jet dynamics that enhance system momentum. This interpretable evidence suggests that StormDiT has spontaneously learned to identify and utilize non-local causal precursors to drive its formation process, effectively addressing the "gray zone" physics that deterministic and decompositional models fail to capture.
[0057] It should be noted that in some cases, the actions described in the specification can be performed in different orders and still achieve the desired results. In this embodiment, the order of steps is given only to make the embodiment clearer and easier to explain, and not to limit it.
[0058] In one optional embodiment, please refer to Figure 11 To improve the practicality and facilitate the promotion of this method, this invention provides a high-frequency, second-long-term time-series short-term forecasting system based on generative AI. The high-frequency, second-long-term time-series short-term forecasting system based on generative AI includes: a data input device 1, a data output device 2, a processor 3, and a storage device 4. The storage device 4 includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor 3, cause the processor 3 to implement the high-frequency, second-long-term time-series short-term forecasting method based on generative AI provided by this invention.
[0059] In summary, this method offers at least the following advantages: Addressing the issue of observation-based extrapolation models failing in gray regions due to accumulated nonlinear errors, and numerical weather prediction models suffering from the "cold start" effect, this method maps the forecasting problem to a compressed latent manifold. By processing data through a diffusion Transformer, it can better handle complex atmospheric conditions in gray regions, achieving effective forecasts over long periods and at high frequencies. Unlike deterministic models that aim to minimize mean error, leading to "mean regression" and image blurring, this method's latent manifold processing better preserves data details and structure, ensuring that forecast images clearly present precipitation and other relevant information, while retaining the sharpness of key features such as extreme weather. Furthermore, this method employs a probabilistic diffusion process. This diffusion generation method possesses inherent stability and diversity. Unlike GANs, which are prone to getting trapped in local optima during training, leading to model collapse, the diffusion generation model generates data by progressively adding and removing noise. This method allows for a more stable exploration of data distribution, thus avoiding model collapse and ensuring generation quality. Unlike decomposition paradigms that explicitly separate weather evolution into additive components, this method employs a global attention mechanism in the Diffusion Transformer, allowing the model to focus on long-range dependencies in the data. This enables the model to learn nonlocal causal rules of convection, learn the physical laws governing the atmosphere spontaneously without relying on artificial decomposition assumptions, and fully model the causal chains in the atmosphere where kinematics and dynamics are closely intertwined, improving prediction accuracy. After training on large-scale data, the model provided by this method can spontaneously learn the physical laws governing the atmosphere, possessing both the flexibility of data-driven methods and the ability to simulate physical processes. Therefore, it can serve as an effective tool connecting data-driven extrapolation and physical simulation, laying a solid foundation for building digital twins of the Earth system. A system adapted to this method is provided, which not only enhances the practicality of this method but also facilitates its promotion.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A high-frequency, long-term, short-term forecasting method based on generative AI, characterized in that, Includes the following steps: The original radar reflectivity sequence is compressed into a compact latent variable token sequence using a pre-trained variational autoencoder. Within the compressed latent space, using the latent variable Token sequence as a condition, a predictive latent variable Token sequence for future time periods is gradually generated through probability generation. The decoder using the variational autoencoder reconstructs the predicted latent variable token sequence into a predicted radar reflectivity sequence for future time periods.
2. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 1, characterized in that: The variational autoencoder is a causal variational autoencoder.
3. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 2, characterized in that: The causal variational autoencoder uses causal 3D convolution.
4. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 1, characterized in that: Using the latent variable token sequence as a condition, and based on the noisy token sequence, the predicted latent variable token sequence is generated using a diffusion generation model.
5. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 4, characterized in that: The noise token sequence is obtained by random sampling from Gaussian noise.
6. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 5, characterized in that: The number of frames in the noise token sequence is consistent with the length and time step of the predicted radar reflectivity sequence.
7. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 6, characterized in that: The diffusion generation model is based on the Rectified Flow framework and is built using the Diffusion Transformer.
8. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 7, characterized in that: The diffusion Transformer uses a 3D causal self-attention mechanism to simultaneously capture local advection and non-local convection triggers.
9. The high-frequency sub-long-term short-term forecasting method based on generative AI according to claim 8, characterized in that: The diffusion Transformer injects time coordinates into each layer of the network through adaptive layer normalization.
10. A high-frequency, long-term time-series short-term forecasting system based on generative AI, characterized in that, The high-frequency sub-long-term short-term forecasting system based on generative AI includes: a data input device, a data output device, a processor, and a storage device. The storage device includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor, cause the processor to implement the high-frequency sub-long-term short-term forecasting method based on generative AI as described in any one of claims 1-9.
Citation Information
Cited By
Extreme precipitation prediction method based on high-resolution statistical extreme value recovery
CN122283982A