A multi-climate zone memory codebook pre-training large-area severe convective weather forecasting method and system

CN122388708BActive Publication Date: 2026-08-28OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610857722.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-08-28
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

[0007]为了解决现有方法在大区域强对流临近预报中因地域性差异和泛化能力不足的问题,本发明提供一种多气候带记忆码本预训练的大区域强对流预报方法及系统

Benefits of technology

[0026]针对于现有的深度学习方法主要适用于几百公里范围内的局部区域,对大范围、尤其是多气候区的强对流天气预报仍存在特征提取能力不足、模型泛化性较差,以及预报结果缺乏细节等问题。本发明通过构建适应多气候区域的预报模型,实现了全中美范围的强对流天气预报,本发明所带来的有益效果,主要体现在以下几个方面:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122388708B_ABST
    Figure CN122388708B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of severe convective weather forecasting, in particular to a multi-climate zone memory codebook pre-training large-area severe convective weather forecasting method and system. The method comprises data preprocessing based on obtained radar weather data; constructing a multi-climate zone general codebook representation based on the Cobbett climate classification method; generating a general climate representation based on feature extraction of the multi-climate zone general codebook representation; obtaining a spatiotemporal dynamic feature by spatiotemporal feature decoupling of the general climate representation based on an improved Transformer; and generating a severe convective weather forecasting result by double-path representation enhancement and gradual decoding based on the spatiotemporal dynamic feature. The present application effectively solves the problems of insufficient feature extraction and poor generalization ability caused by climate differences when a traditional deep learning forecasting model is applied across regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of severe convective weather forecasting technology, and in particular to a large-area severe convective weather forecasting method and system with multi-climate zone memory codebook pre-training. Background Technology

[0002] Severe convective weather typically forms within convective cloud clusters or individual convective clouds. It is a weather phenomenon triggered by changes in vertical airflow, generally occurring within hours in various extreme forms, such as short-duration heavy rainfall, thunderstorms, strong winds, hail, tornadoes, or squall lines. It may also be accompanied by secondary disasters such as mudslides, landslides, and floods, posing a serious threat to human life and social production. It is a highly destructive weather phenomenon. Its main characteristics include suddenness, localization, and a short lifespan, making accurate prediction a global challenge.

[0003] As the greenhouse effect intensifies, atmospheric water vapor content continues to increase, leading to a significant rise in the frequency of severe convective weather. Improving the accuracy of severe convective weather forecasts in terms of timing, location, and intensity has become a critical technical issue that urgently needs to be addressed to optimize disaster prevention measures and reduce socio-economic losses.

[0004] Traditional methods for nowcasting severe convective weather mainly include radar extrapolation, numerical models, and conceptual models. While these methods can provide forecast information to some extent, the small spatial scale and short duration of severe convective weather phenomena, coupled with the current inability to comprehensively and effectively analyze and utilize observational data, result in problems such as short forecast lead times and low accuracy. Furthermore, traditional methods often rely on the experience of meteorologists, which can easily lead to a high false alarm rate. How to accurately forecast severe convective weather within a short timeframe remains a research challenge.

[0005] With the significant improvement in computing power, deep learning technology has achieved breakthroughs in many fields, including natural language processing and computer vision. These technologies are data-driven and end-to-end trained, enabling them to automatically extract useful features from massive amounts of data. In recent years, the application of deep learning technology in weather forecasting has gradually increased, demonstrating enormous potential in severe convective weather forecasting. Deep learning-based radar echo extrapolation technology, by combining radar monitoring data with deep learning techniques, can predict the short-term movement direction and intensity of radar echoes, opening a new path for nowcasting severe convective weather.

[0006] While deep learning methods have made some progress in the field of severe convective weather nowcasting, the occurrence of severe convective weather exhibits significant regional characteristics, with substantial differences in frequency, intensity, and manifestations across different climate types. This regionality limits the predictive power of unified models and increases the complexity of large-scale, multi-climate-region forecasting. Existing methods are often trained based on data from specific regions, resulting in poor cross-regional generalization ability and difficulty in adapting to the severe convective weather forecasting needs of different climate zones. Therefore, the core challenge of current research is how to accurately capture and characterize the characteristics of severe convective weather under different climate types over a large scale to improve the cross-regional predictive power of models, and further optimize the details and dynamic evolution of forecasts based on this. Summary of the Invention

[0007] To address the problems of regional differences and insufficient generalization ability in existing methods for large-area severe convective weather nowcasting, this invention provides a large-area severe convective weather forecasting method and system pre-trained with multi-climate zone memory codebooks.

[0008] Firstly, the present invention provides a large-area severe convection forecasting method using multi-climate zone memory codebook pre-training, which employs the following technical solution:

[0009] A large-area severe convection forecasting method pre-trained with multi-climate zone memory codebooks includes:

[0010] Acquire radar weather data;

[0011] Data preprocessing is performed based on the acquired radar weather data;

[0012] A universal codebook representation for multiple climate regions is constructed based on the Köppen climate classification method;

[0013] Feature extraction based on a universal codebook representation for multiple climate regions to generate a universal climate representation;

[0014] Spatiotemporal dynamic features are obtained by decoupling the spatiotemporal characteristics of general climate representation based on the improved Transformer.

[0015] Strong convection forecast results are generated by dual-path characterization enhancement and progressive decoding based on spatiotemporal dynamic features.

[0016] Secondly, a large-area severe convection forecasting system pre-trained with multi-climate zone memory codebooks includes:

[0017] The data acquisition module is configured to acquire radar weather data;

[0018] The preprocessing module is configured to perform data preprocessing based on the acquired radar weather data;

[0019] The codebook representation module is configured to construct a universal codebook representation for multiple climate regions based on the Köppen climate classification method.

[0020] The general characterization module is configured to extract features based on the general codebook characterization of multiple climate regions to generate a general climate characterization.

[0021] The dynamic feature module is configured to obtain spatiotemporal dynamic features by decoupling the spatiotemporal features of the general climate representation based on the improved Transformer.

[0022] The decoding module is configured to perform dual-path characterization enhancement and progressive decoding based on spatiotemporal dynamic features to generate strong convection forecast results.

[0023] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the large-area severe convection forecasting method pre-trained with a multi-climate-zone memory codebook.

[0024] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a large-area strong convection forecasting method pre-trained with a multi-climate zone memory codebook.

[0025] In summary, the present invention has the following beneficial technical effects:

[0026] Existing deep learning methods are primarily applicable to localized areas within a few hundred kilometers. However, they suffer from insufficient feature extraction capabilities, poor model generalization, and a lack of detail in forecasts for large-scale, especially multi-climate, severe convective weather. This invention constructs a forecasting model adapted to multiple climate regions, achieving severe convective weather forecasting across Central and Eastern China. The beneficial effects of this invention are mainly reflected in the following aspects:

[0027] (1) By introducing deep learning algorithms and combining meteorological knowledge with deep learning methods, a series of technical means are used to achieve accurate nowcasting of severe convective weather in a wide range of climate regions. These include S1 constructing a multi-region codebook representation based on Köppen climate classification, S2 designing a codebook dynamic calling and fusion module, S3 designing a spatiotemporal decoupling module, and S4 constructing a dual-path fusion and progressive decoding architecture.

[0028] (2) Using the multi-climate region codebook representation pre-training method based on Köppen climate classification proposed in S1, the entire CREF dataset is divided into 8 major climate zones, and a dedicated climate codebook is pre-trained independently for each climate zone, constructing a large-scale external climate codebook library containing 160,000 feature vectors. This method stores and represents climate prior knowledge in the form of discrete codebooks, and replaces the traditional nearest neighbor static matching method with a dynamic codebook matching mechanism, enabling the codebook to more accurately capture the typical features of severe convective weather under different climate conditions. This effectively solves the problem of insufficient feature extraction and poor generalization ability caused by climate differences when traditional deep learning forecast models are applied across regions.

[0029] (3) Through the multi-climate codebook representation extraction module designed in S2 and the spatiotemporal feature decoupling module designed in S3, complementary synergy between local spatiotemporal dynamic features and global general climate representation is achieved. Specifically, S2 deploys the eight pre-trained regional codebooks as external memory modules, adaptively retrieves and fuses them during the forecasting stage to generate a global general climate representation, providing cross-regional climate prior guidance for the model; S3, through the parallel decoupling mechanism of temporal attention and spatial attention, splits multi-head attention into independent temporal and spatial branches, which can accurately capture the temporal evolution and spatial distribution characteristics of severe convective weather. The two paths complement each other in terms of information sources and functional positioning, together forming a complete feature extraction system.

[0030] (4) Through the dual-path representation enhancement and progressive decoding architecture proposed in S4, the global climate general representation extracted in S2 and the local spatiotemporal dynamic features extracted in S3 are cross-attention fused, enabling the model to adaptively select and enhance the most relevant climate prior information according to the actual characteristics of the current input weather system. The enhanced features after fusion are restored layer by layer through the progressive upsampling decoder to restore the spatial and temporal resolution, effectively avoiding the image stitching traces and spatial detail loss problems caused by traditional single-step restoration. At the same time, the weight-based differential loss function proposed in S4 assigns higher loss weights to high echo value areas according to radar echo intensity, forcing the model to pay more attention to the core area of ​​strong convection during training, which significantly improves the model's ability to identify and forecast high-intensity echo areas.

[0031] (5) Through the overall technical solution from S1 to S4, this invention constructs a complete technical chain of "database construction - database usage - decoupling - fusion". S1 is the offline pre-training stage, which encodes the strong convection features of multiple climate regions into independent codebooks; S2 and S3 are two parallel paths in the online inference stage, which extract complementary features from external priors and input data, respectively; S4 deeply fuses the two sets of features and generates the final forecast result. The steps in this technical chain have clear division of labor and synergistic efficiency, which makes the model significantly improve in key indicators such as critical success index and Heidelberg skill score in cross-climate region strong convection forecasting tasks compared with existing methods. In addition, the codebook scalability update strategy proposed in this invention allows for incremental training of new codebooks and expansion to the existing codebook library when adding new climate regions, without retraining the entire model, ensuring the long-term availability and maintainability of the model in the context of continuously growing meteorological data. Attached Figure Description

[0032] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention.

[0033] Figure 2 This is a schematic diagram (BWk) of the prediction results of the comparative model at different times in Embodiment 1 of the present invention.

[0034] Figure 3 This is a schematic diagram of the prediction results (Cfa) of the comparative model at different times in Embodiment 1 of the present invention.

[0035] Figure 4 This is a schematic diagram of the visualization (Dwa) of the prediction results of the comparative model at different times in Embodiment 1 of the present invention.

[0036] Figure 5 This is a schematic diagram of the prediction results (ET) of the comparative model at different times in Embodiment 1 of the present invention.

[0037] Figure 6 This is a line graph showing the changes of various indicators of the comparative model in Embodiment 1 of the present invention at different time frames of 35dBZ. Detailed Implementation

[0038] The present invention will be further described in detail below with reference to the accompanying drawings.

[0039] Example 1

[0040] Reference Figure 1 This embodiment presents a large-area severe convection forecasting method pre-trained with a multi-climate-zone codebook. Its overall technical framework includes a large-area severe convection forecasting pre-training method and a large-area severe convection forecasting model (Trans-CodeNet). The specific scheme is as follows:

[0041] S1: Based on the Köppen climate classification method, construct a universal codebook representation for multiple climate regions to form a priori knowledge base.

[0042] This step aims to provide reusable climate prior knowledge for subsequent cross-regional forecasts, addressing the problem of insufficient model generalization ability caused by training in a single climate region. The physical mechanisms and manifestations of severe convective weather differ significantly across different climate regions, and the performance of traditional single-region-trained forecast models deteriorates markedly when applied across regions. This step introduces climate classification prior knowledge and constructs a multi-regional codebook library to provide forecast models with a universal and efficient representation of climate characteristics.

[0043] S1.1: Target Area Data Retrieval

[0044] Retrieves historical archives and real-time updated weather radar reflectivity factor (dBZ) data for a specified area.

[0045] Data sources include, but are not limited to, national operational radar networks and specialized experimental radars. The radar reflectivity factor (CREF) data used in this invention integrates advanced technologies from S-band and C-band Doppler weather radars, enabling real-time capture of the evolution of extreme weather events. The geographical coordinates of the study area range from North latitude... to East longitude to The total coverage area is approximately 4200km × 6200km. The temporal resolution is 10 minutes per frame, and the spatial resolution is 1km × 1km.

[0046] S1.2: Data preprocessing methods.

[0047] The radar data preprocessing in this study comprises four stages. First, a sliding window is used to extract a 384km × 384km local area and radar images from 19 consecutive time points to construct a complete spatiotemporal sequence sample. Second, outlier samples without storms are removed, missing values ​​(e.g., -999dBZ) are replaced with 0, and the data is truncated to the effective interval [0, 80]dBZ to eliminate weak signal interference. Next, storm points are screened using a threshold of 35dBZ. When there are 1900 storm points in the sequence, they are defined as positive samples; otherwise, they are negative samples. Negative samples are randomly undersampled to alleviate data imbalance caused by long-tail distribution. Finally, the Min-Max method is used to linearly map the data to the interval [0, 1] and perform boundary truncation, ultimately generating a standardized model training set.

[0048] Specifically,

[0049] (1) Spatial sampling and temporal sample construction

[0050] In the spatial dimension, a fixed step size is used to perform sliding sampling on the global radar image, extracting a local area image with a preset scale of 384km × 384km to reduce computational overhead and ensure that the sampling results cover different climate zones. In the temporal dimension, based on a sliding window strategy, radar images at 19 consecutive time points are extracted along the time dimension to construct a spatiotemporal sequence sample, and samples with missing time points are removed to ensure the continuity and integrity of the spatiotemporal sequence data. The formula for this process is expressed as follows: Let the original global radar data sequence be... ,in .in, This represents the raw spatial height and width of the global radar image. The step size of the spatial sliding window is set to... and The size of the cropped local image is In the embodiments, that is... If because the spatial resolution is ,so . No. Spatial starting coordinates of a local region in the global image It can be represented as: The extracted spatial sequence is represented as The sliding window size along the time dimension is set to... In the embodiments (Frame). The first frame constructed. Spatiotemporal sequence samples Represented as: .

[0051] (2) Outlier correction and numerical truncation

[0052] Outlier filtering is performed on the radar images in the spatiotemporal sequence samples, directly removing radar images that do not contain storm information and contain outliers; for the retained radar images, pixel values ​​marked as outliers or missing (e.g., -999dBZ) are replaced with 0dBZ; simultaneously, an effective numerical range for radar reflectivity is set (e.g., [0, 80]dBZ), and the data is truncated: values ​​below the lower limit of the range are corrected to the lower limit value, and values ​​above the upper limit of the range are corrected to the upper limit value, in order to eliminate interference from low-value weak signals. The formula for this process is expressed as follows: Let... This represents any pixel value in the original radar image. Outlier correction function. Defined as:

[0053] ,

[0054] in For missing markers (e.g.) After correction, numerical truncation is performed, and the truncation function is used. Defined as:

[0055] ,

[0056] in, This is the lower limit of the effective range of radar reflectivity, which is set to 0 dBZ in this embodiment. It is the upper limit of the effective range of radar reflectivity, which is set to 80 dBZ in the embodiment.

[0057] (3) Separation of positive and negative samples and data balancing

[0058] A storm intensity threshold of 35 dBZ is set, and pixels in the radar image with values ​​greater than or equal to this threshold are defined as storm information points. The total number of storm information points contained in the spatiotemporal sequence sample (19 consecutive time points) is counted. If the total number is greater than or equal to a preset threshold n (set to n=1900), the spatiotemporal sequence sample is classified as a positive sample; otherwise, it is classified as a negative sample. To alleviate the data imbalance problem caused by the long-tail distribution of severe convective weather, the negative samples are randomly undersampled (i.e., some negative samples are randomly discarded) to increase the proportion of positive samples in the dataset. The formula for this process is as follows: Let a single spatiotemporal sequence sample be... The pixel index is Define the storm intensity indicator function. :

[0059] ,

[0060] in, The storm intensity threshold is set at 35 dBZ in this example. The total number of storm points within the spatiotemporal sequence is counted. The formula is . The first in the spatiotemporal sequence sample Frame, coordinates are Radar echo value at the location. Positive and negative sample classification labels. for:

[0061] ,

[0062] in, The storm information points contained within a single spatiotemporal sequence of 19 consecutive moments (i.e. The total number of pixels in dBZ. The threshold for the number of positive samples is set; in this example, this value is set to 1900. For the sample pool with label 0, the sampling rate is... Perform random undersampling, retaining a sample size equal to a certain proportion of the positive sample size.

[0063] (4) Data normalization processing

[0064] The Min-Max normalization method is used to normalize the processed spatiotemporal sequence samples. This study utilizes predetermined range boundaries (maximum and minimum values) of the radar data to linearly map the pixel values ​​in the samples to the interval [0, 1]. Values ​​exceeding this range after mapping are then truncated, with values ​​less than 0 set to 0 and values ​​greater than 1 set to 1. The formula is as follows: This ultimately yields the normalized radar sequence dataset used for model training. The formula for this process is expressed as follows: Given the previously processed pixel values... The Min-Max normalization mapping process is as follows: Then, boundary truncation is performed to ensure the final normalized value. Strictly fall into Inside: .in, This is the predetermined theoretical boundary minimum value of the radar data, which is 0 dBZ in this embodiment. The predetermined theoretical maximum value for radar data is 80 dBZ in this embodiment. It is the intermediate value after linear mapping. This is the final output of the network model for training, and its value range is [value range missing]. .

[0065] S1.3: Based on the Köppen climate classification method and combined with temperature and precipitation indicators, the entire target area is divided into 8 major climate zones.

[0066] The Köppen climate classification system categorizes climate types based on the distribution characteristics of temperature and precipitation. Its unique feature lies in its focus on seasonal temperature variations closely related to the occurrence and evolution of severe convective weather. The dataset is largely covered by eight Köppen climate types: Bsh (tropical savanna), BWk (cold desert), Cwa (humid subtropical monsoon), Cfa (humid subtropical), Dwa (temperate continental monsoon), Dwb (temperate continental monsoon), Dwc (subarctic monsoon), and ET (tundra). During data preprocessing, climate zone assignments were determined based on latitude and longitude information from radar echo data, resulting in eight independent climate region datasets to ensure the inherent consistency of severe convective characteristics within each climate zone.

[0067] S1.4: In the codebook pre-training stage, a codebook pre-training network structure with multiple temporal inputs and outputs was designed to fully capture spatiotemporal dynamic features.

[0068] Traditional representation learning models typically employ a single-time-input-output ("one in, one out") mechanism. This invention improves upon this by using a multi-time-series input-output mechanism consistent with downstream forecasting models, i.e., inputting data frames from the previous hour (7 frames) and predicting data frames from the next 2 hours (12 frames). Input data It is a five-dimensional tensor, in which Where is the number of samples, and T is the time dimension. For the number of channels, and These represent the spatial dimensions. The encoder uses a 3D convolutional network to compress the features of the input data, and its calculation formula is as follows:

[0069] ,

[0070] in, This represents the encoded query vector. Represents the number of query vectors. For the feature dimensions of the query vector, This represents the encoder's operation. The query vector 𝑄 output by the encoder contains key spatiotemporal information about the input data.

[0071] S1.5: A set of learnable discrete feature vectors is introduced between the encoder and decoder. An independent codebook is constructed for each climate zone, with each codebook having a size of 20,000, resulting in a total of 160,000 codebook vectors for the eight climate zones. The codebooks are initialized randomly and their representational capabilities are dynamically updated during training. The pre-trained set of eight climate codebooks can be represented as:

[0072] ,

[0073] The codebook for each climate region is as follows:

[0074] ,

[0075] in, The feature dimension of the codebook vector. In traditional image representation learning, the construction of latent space vectorized codebooks mainly relies on nearest neighbor search, finding the nearest point based on the Euclidean distance between features. This static matching method ignores the contextual information and potential distribution characteristics of the input data, easily leading to a lack of sensitivity to subtle changes in the codebook constructed in complex climatic regions or high-dimensional data. To solve this problem, this invention introduces a dynamic codebook matching mechanism. Its core idea is to use a Query-Key-Value structure for dynamic feature matching to improve the model's adaptability to feature distributions under different climatic conditions. Specifically, the codebook is first... Mapping to the key space and value space through two independent linear transformations:

[0076] ,in, It is pre-trained, the first A set of discrete feature vectors for each climate region. It stores the prior features of typical severe convective weather in that climate region. From the source codebook Multiply by weight The calculation yields the result. In the subsequent attention mechanism, it is used to calculate the dot product similarity with the input query vector, thereby evaluating which features in the codebook best match the current weather input. From the source codebook Multiply by weight It is calculated. It carries the actual feature information. When the model passes... Once the matching weights are calculated, the matching will be performed according to these weights. The vectors in the dataset are weighted and summed to extract useful representational features. The weight matrix is ​​a learnable matrix. This provides a unified feature dimension for both keys and values. Then, the query vector output by the encoder is used. As a query, representative features are extracted from the current climate codebook using an attention mechanism. This mechanism calculates the similarity between the query and the key vector, selects the value vector most relevant to the input, and then extracts features from the codebook.

[0077] ,

[0078] in, For from the first Feature representations extracted from climate region codebooks. This design significantly enhances the dynamic representation capability of the codebook, enabling it to more accurately capture the spatiotemporal features of the input data, thereby constructing a higher-quality climate codebook.

[0079] S1.6: Finally, the decoder section restores the radar data at the original resolution from the features retrieved from the codebook by upsampling layer by layer:

[0080] ,in, This represents the reconstructed radar timing data. For from the first Feature representations extracted from codebooks of climate regions. This represents the layer-by-layer upsampling operation of the decoder. During the decoding process, layer-by-layer convolution operations can refine spatial structure information, preserving and restoring complex spatial details in severe convective weather.

[0081] S2: Design a multi-climate codebook representation extraction module to dynamically call pre-trained codebooks and fuse them to generate a general climate representation during the forecasting stage.

[0082] This step addresses how to effectively utilize the eight external climate codebooks pre-trained in S1 during the inference stage of severe convection nowcasting to accurately extract and integrate cross-regional climate prior information from the current input data. Unlike S1, which focuses on codebook pre-training and construction, S2 focuses on the deployment and invocation strategies of the codebooks in the forecasting model, and is part of the Trans-CodeNet model inference process.

[0083] S2.1: In the forecast model, the codebooks for the eight climate regions pre-trained in S1 are used. These codebooks are deployed as external memory modules. Their parameters remain fixed during the forecast model training and inference phases and are not updated; they serve only as a read-only prior knowledge base to support climate characteristics. When expansion to new forecast regions or climate types is required, only the corresponding codebooks for the new regions need to be incrementally trained according to the above steps, and then added to the existing codebook library; the entire model does not need to be retrained. To achieve effective integration between the codebooks and the forecast model, this module uses a 3D convolutional encoder with the same structure as the S1 pre-training phase to process the radar echo sequences input to the forecast model. Feature compression is performed. This design ensures that the input feature distribution remains aligned with the feature space during the codebook pre-training stage, enabling efficient and accurate subsequent codebook retrieval. The encoder outputs a set of query vectors. This serves as the basis for subsequent cross-regional feature retrieval.

[0084] S2.2: The query vector generated by the encoder Simultaneously, it is sent to eight pre-trained climate region codebooks, and an independent dynamic matching retrieval operation is performed on each codebook. For the first... Codebook for each climate zone Using the dynamic codebook matching mechanism defined in S1 (formula same as S1.4), the local representations most relevant to the climate characteristics of the region are retrieved and extracted. :

[0085] ,

[0086] Here, `CodebookRetrieve(⋅)` represents a dynamic codebook retrieval operation based on an attention mechanism. The specific calculation method is consistent with the pre-training stage in S1.4, but the codebook parameters are fixed in this module and do not participate in gradient updates. Through parallel retrieval, a set of eight local climate feature representations is obtained:

[0087] , This represents the final collection of local climate feature representations compiled after parallel retrieval. Each local representation... It contains information on strong convection patterns unique to the corresponding climate region, providing diverse candidate features for subsequent global fusion.

[0088] S2.3: Due to the significant differences in the relevance of different climate regions to the current input data (for example, a weather system that mainly occurs in a humid region should have a correspondingly lower feature contribution from arid region codebooks), this module designs an adaptive fusion strategy based on an attention mechanism to dynamically evaluate and weightedly integrate eight local climate representations. Specifically, the eight local features are first concatenated along the feature dimension:

[0089] Subsequently, the concatenated features are mapped to the attention space through a learnable linear transformation:

[0090] ,in, These are learnable parameters unique to this module and are independent of S1. This represents the unified feature dimension of the multi-head attention fusion space. This is a tensor formed by splicing together the local features of eight climate regions along the feature dimension. The concatenated features are linearly mapped to a query, key, and value matrix that enters the attention space. Finally, an attention mechanism adaptively assigns fusion weights to the representations of different climate regions, generating a unified global climate representation.

[0091] Where N represents the number of query vectors, As a global climate general characterization output, it integrates prior information from eight climate regions, providing cross-regional climate background guidance for subsequent forecasts.

[0092] S3: Design a spatiotemporal feature decoupling module to extract local spatiotemporal dynamic features from the input data in a refined manner.

[0093] This step focuses on uncovering the spatiotemporal evolution patterns of the input radar echo data itself, complementing the static climate background representation extracted by S2 from the external codebook. The nowcasting of severe convective weather is essentially a spatiotemporal sequence prediction problem, and radar data exhibits significant temporal evolution and spatial dependence. While S2's codebook representation extraction can capture global features across multiple climate regions, it may still overlook the dynamic evolution details of the input data itself, leading to insufficient modeling capabilities for local features. This module, based on an improved Transformer architecture, achieves fine decoupling of temporal and spatial dimensions through an attention mechanism, thereby capturing key patterns in the weather evolution process.

[0094] S3.1: To more effectively extract temporal features, this module designs a frame-level patch segmentation strategy, treating the time dimension T as an independent dimension and performing patch segmentation on each radar image frame separately. Unlike the traditional method of merging the time dimension into channel number before unified segmentation, this design can capture the temporal evolution pattern more meticulously, while avoiding interference between the time dimension and spatial features. The input data is a radar sequence containing the time dimension. ,in Where is the number of samples, and T is the time dimension. For the number of channels, and These are the spatial dimensions. For each frame in the temporal dimension... Divide it into fixed-size patches and map them to a high-dimensional feature space: Specifically, set the space size of the patch to be... This can be achieved through a sliding window or an equivalent two-dimensional convolution operation (where both kernel size and stride are...). ), dividing spatial dimensions into The first non-overlapping local image patches. For the first... The first frame Image blocks It is flattened and mapped to the feature dimension through linear projection. :

[0095] ,

[0096] in, Indicates the flattening operation. It is a learnable linear projection matrix. This is the bias vector. Each frame is thus transformed into a serialized patch feature. To preserve temporal and spatial location information, temporal and spatial location codes were subsequently injected into the Patch features. Considering the spatiotemporal evolution characteristics of severe convective weather, spatial location coding... Used to label the absolute spatial coordinates and temporal location encoding of each patch in the original 2D image. Used to distinguish the chronological order of different time frames. The injection and fusion process of the two is represented as:

[0097] ,

[0098] in, This can be achieved using sine and cosine periodic functions or constructed through learnable parameters to ensure the model is aware of the absolute and relative distances at time steps. Subsequently, all time dimensions... The feature sequences of the frames are concatenated to obtain the complete input representation:

[0099] ,

[0100] S3.2: The encoded data is processed using a global attention mechanism to construct a global representation of the input data. To more fully capture the spatiotemporal dependencies of severe convective weather on a global scale, the input features are... Perform multi-head self-attention computation. Specifically, firstly, the input features are processed through different linear mappings. Projected into query matrices respectively Key matrix Sum matrix :

[0101] ,in, For the total number of attention heads, For the first The learnable weight matrix corresponding to each attention head The feature dimensions for each attention head.

[0102] Subsequently, scaled dot product attention is computed independently for each attention head to dynamically evaluate the interaction and correlation strength between different spatiotemporal patches:

[0103] ,in, For the first A query, key, and value matrix for each attention head. For the first The output result after dynamically evaluating the spatiotemporal patch relevance by an attention head. The feature dimension for each attention head is also the scaling factor. Finally, the outputs of all attention heads are concatenated along the feature dimension and mapped back to the original dimension using the output projection matrix to generate an intermediate feature representation containing global spatiotemporal dependencies. :

[0104] ,

[0105] in, This is the learnable weight matrix for the global attention fusion stage.

[0106] S3.3: To address the heterogeneity of temporal and spatial patterns in severe convective weather, two decoupling mechanisms, temporal attention and spatial attention, are further introduced. To balance computational complexity, a parallel temporal and spatial attention mechanism based on head allocation is constructed. This integrates the multi-head attention mechanism... Each attention head is split into two groups: Size allocation is used for spatial attention, focusing on modeling dependencies between different spatial locations within a single frame; the rest... Size allocation is used for temporal attention, focusing on modeling the dynamic evolution of the same spatial location across different time frames. For the spatial attention branch, the input features are reshaped into sequences of spatial patches, allowing attention computation to be performed only in the spatial dimension. Specifically, for the input features of the spatial attention branch... First, a query matrix is ​​generated through a linear transformation. Key matrix Sum matrix :

[0107] ,

[0108] in, This is the learnable weight matrix corresponding to the spatial branch. Subsequently, the spatial self-attention weights are calculated, and the value matrix is ​​weighted and aggregated:

[0109] ,in, This is used to capture the spatial branch output of local structural features at different geographical locations within the same timeframe. In this calculation, the attention weight matrix has a dimension of [missing information]. It is specifically designed to capture the spatial correlation and local structural features between radar echoes from different geographic locations (patches) within the same timeframe. Similarly, the temporal attention branch output is calculated. Finally, the outputs of the two branches are merged through a concatenation operation:

[0110] ,

[0111] in, It is a learnable linear projection matrix used to remap the stitched spatiotemporal features to a unified feature space.

[0112] S3.4: After completing the decoupling computation of temporal and spatial attention, this module introduces another layer of global attention mechanism to further refine and integrate the decoupled and fused spatiotemporal features. The purpose of this design is to optimize the features from a global perspective while retaining the refined spatiotemporal information after decoupling, forming complete spatiotemporal dynamic features that contain both local details and global semantics.

[0113] ,

[0114] Final output This refers to the local spatiotemporal dynamic feature representation extracted by this module. It includes the dynamic evolution information of the input radar echo data in the time dimension and the distribution structure information in the spatial dimension. It will serve as one input branch of the subsequent S4 dual-path fusion module, together with the global climate general representation output by S2. To achieve deep integration.

[0115] S4: Construct a dual-path representation enhancement and progressive decoding architecture to integrate climate priors and spatiotemporal dynamic features to generate severe convection forecast results.

[0116] This step proposes a dual-path representation enhancement mechanism, which effectively fuses the "global climate general representation" extracted by S2 from an external climate codebook with the "local spatiotemporal dynamic features" extracted by S3 from the input data. A progressive decoder then generates the final high-resolution strong convection nowcasting result. The core idea of ​​this design is that the global climate prior provided by S2 can compensate for the shortcomings of S3 in cross-regional generalization, while the local dynamic evolution captured by S3 can compensate for the limitations of S2 in temporal detail modeling. The two complement each other and synergistically enhance each other. The dual-path representation enhancement and progressive decoding architecture is the final architecture of the Trans-CodeNet model.

[0117] S4.1: This module receives heterogeneous feature representations from two independent paths as input.

[0118] Path 1 Output: Local spatiotemporal dynamic features extracted by the S3 spatiotemporal decoupling module This feature, learned directly from the input radar echo sequence, contains the fine temporal evolution patterns and spatial distribution structure of the current weather system. Among them, For sample batch size, For the input time dimension, The number of spatial image patches in a single frame of radar image. The hidden channel dimension after decoupling from spatiotemporal features.

[0119] Path 2 Output: Global climate general representation generated by the S2 multi-climate codebook representation extraction module This feature is retrieved and fused from a pre-trained external climate codebook, containing extensive, cross-regional climate background knowledge and typical strong convection patterns. Among them, This represents the number of query vectors (its value is consistent with the number of spatial image patches in a single frame of radar image). This represents a unified feature dimension for the multi-head attention fusion space.

[0120] Due to the characteristics of the two paths in terms of tensor order and channel dimension ( and The feature input exhibits significant spatial and channel heterogeneity, thus it is defined as a heterogeneous feature input. The features of the two paths are complementary in origin and nature: path one focuses on the instantaneous dynamics of current observations, while path two provides long-term accumulated climate priors. Relying solely on either path will lead to information loss.

[0121] S4.2: To achieve the organic integration of the two types of heterogeneous features, they are first aligned in dimensions. The global climate representation output from S2 is then... Broadcast expansion is performed in both the sample and time dimensions to align it with the spatiotemporal dynamic features output by S3. Maintain consistency in shape:

[0122] ,

[0123] Then, the aligned two types of features are concatenated along the feature dimension:

[0124] ,

[0125] To further achieve deep integration, a lightweight cross-attention module is introduced, using features from path one as the query and features from path two as the key and value, to dynamically select and enhance the climate prior information most relevant to the current spatiotemporal dynamics.

[0126] in, For multi-head cross-attention operations:

[0127] The core advantage of this cross-attention mechanism lies in its ability to allow the model to adaptively select and enhance the most relevant parts from the global climate representation based on the actual spatiotemporal characteristics of the current input weather system, rather than simply statically concatenating the two. To dynamically select and enhance the output fusion features based on relevant climate prior information, while also incorporating local dynamics and global climate information. This is the scaling factor for the feature dimension. The final output is the fused feature. It contains information on both local dynamics and climate a priori aspects.

[0128] S4.3 will integrate the enhanced features. The data is fed into the decoder, where a progressive upsampling strategy is used to restore the spatial and temporal resolution layer by layer, ultimately generating a radar echo prediction sequence for future timeframes. First, the fused features need to be reconstructed from a two-dimensional sequence back into a three-dimensional spatial structure for subsequent layer-by-layer restoration:

[0129] ,

[0130] in, This represents the number of patches per frame at the current scale, corresponding to the spatial compression ratio of the deepest layer of the encoder. The decoder consists of L layers, each performing two operations sequentially: feature refinement and spatial upsampling. For the... layer( (e.g., =1,2,…,L) will update the features of the previous layer's output through the Transformer decoder module:

[0131] ,in, For the decoder The output feature tensor of the layer. This is the output of the layer above the decoder. For the first The Transformer module of the layer decoder contains a multi-head self-attention mechanism and a feedforward network to further extract and integrate spatiotemporal information from features at the current scale. For the first Number of patches per layer For the first The feature dimensions of the layer are determined. Secondly, the spatial resolution of the feature map is gradually restored through an upsampling layer. Specifically, for the refined features... Perform an upsampling operation:

[0132] ,

[0133] in, This can be achieved using transposed convolution. After upsampling, As the spatial dimensions increase, the feature dimensions are adjusted accordingly, thereby restoring the spatial details of the radar echo layer by layer.

[0134] S4.4: To improve the model's prediction accuracy for key areas of strong convection, this module designs a weighted loss function based on weights. Strong convection events are low-probability events, occurring in only a small portion of the entire study area. However, traditional loss functions (such as Mean Squared Error (MSE) and Mean Absolute Error (MAE) assign the same weight to all areas during training, causing the model to focus more on widely distributed non-strong convection areas during optimization, while ignoring high-risk areas where strong convection events occur. Therefore, this module assigns differentiated weights to different pixels based on radar echo intensity:

[0135] ,

[0136] in, Here, represents the radar echo value (in dBZ) of the corresponding pixel. Prediction errors in high echo value regions carry a larger weight in the loss calculation, forcing the model to focus heavily on the core regions of strong convection during training. Based on the above weights, the weighted mean square error and weighted average absolute error are defined as follows:

[0137] ,in, To predict the number of sequences, This represents the resolution of a single-frame radar image. This represents the radar echo value of the corresponding pixel. and These are the actual value and the predicted value, respectively. and These represent the weighted mean square error and the weighted average absolute error, respectively. This represents the total number of predicted sequences involved in the error calculation. This represents the horizontal spatial resolution of a single radar image. Differential weights are dynamically assigned based on pixel echo intensity, with higher echo values ​​carrying greater weight, thus forcing the model to focus on the core region of strong convection.

[0138] The final comprehensive loss function is a weighted combination of the two:

[0139] ,

[0140] in, and The weighting coefficients are used to adjust the relative importance of the two types of losses in the overall optimization objective. This loss function combines the sensitivity of MSE to large errors with the robustness of MAE to outliers, while also enhancing the focus on key areas of strong convection through the weighting mechanism, significantly improving the model's prediction accuracy for high echo value regions.

[0141] Experimental verification

[0142] The input and output data of the strong convection prediction experiment are both radar echo images, with 7 frames of input data and 12 frames of output data. The time resolution is 10 minutes. The radar echo images of the next 120 minutes are predicted using radar echo images of a complete 60-minute time interval.

[0143] To systematically evaluate the model's forecasting performance in regions with echoes of varying intensities, this invention starts with different evaluation metrics and thresholds, compares it with other mainstream methods, and combines the evaluation metrics with visualization results to verify the superiority of the proposed model. The experiment selected four representative deep learning models for comparative analysis: SiMVP (SimpleryetBetterVideoPrediction), proposed in 2022 in the field of spatiotemporal series prediction; the currently mainstream severe convection forecasting models RainFormer (RainfallTransformer) and EarthFormer (Earth-SpecificTransformer); and the classic model TrajGRU (TrajectoryGRU) in the field of severe convection nowcasting. The selection of these comparative models covers different architectures and characteristics, providing a multi-faceted reference for evaluating the advantages of the proposed method.

[0144] All methods were evaluated under the same training and test set conditions. Evaluation metrics included: Critical Success Index (CSI), Heidelberg Skill Score (HSS), Point of Detection (POD), and False Alarm Rate (FAR).

[0145] The experimental results are shown in Tables 1 and 2. Analysis of the tables reveals that the model of this invention performs best under different prediction durations and radar echo thresholds, especially demonstrating significant advantages in the two comprehensive indices, CSI and HSS. At all thresholds, Trans-CodeNet maintains high CSI and HSS values, particularly at high thresholds of 35 dBZ and 40 dBZ, where its ability to balance false alarms and missed alarms is more prominent. Trans-CodeNet also performs excellently in the POD and FAR indices. The POD index shows its high efficiency in hit rate, especially in 2-hour predictions, where its hit rate is significantly higher than other models. In the FAR index, Trans-CodeNet exhibits a low false alarm rate, especially under high threshold conditions, indicating that the model can effectively avoid unnecessary predictions, thereby improving prediction accuracy. Compared to other models, SiMVP shows superior hit rate at low thresholds and short time series, but its performance is relatively limited at high thresholds and long time series. RainFormer and EarthFormer performed well at medium thresholds, but their predictive performance was relatively weak at high thresholds. While TrajGRU performed well in FAR (Frequency Assessment Ratio), it lagged behind in key metrics such as CSI (Concentration Indicator) and POD (Problem-Oriented Distance), indicating poor ability to identify severe convective weather and a tendency to miss or underreport cases. Overall, Trans-CodeNet demonstrated excellent performance in complex weather forecasting scenarios due to its strong adaptability to various climatic conditions and thresholds, particularly in predicting severe convective weather across large areas and multiple climate types.

[0146] Table 1 Comparison of model performance under different radar echo thresholds

[0147]

[0148] The preceding text introduced the Köppen climate classification system and labeled eight major climate zones in the dataset. These eight climate zones belong to four climate groups: arid, temperate, polar, and polar. To more comprehensively evaluate the forecasting capabilities of our model across different climate regions, we selected one climate zone with broad coverage from each of these four climate groups for index testing. According to meteorological standards, radar echo intensity exceeding 35 dBZ is generally considered an indicator of severe convection. Given that the above analysis has discussed experimental results at different dBZ thresholds in detail, the experiments below will uniformly use 35 dBZ as the threshold to present relevant experimental results.

[0149] Table 2 shows the performance of each model in severe convective weather forecasting at a threshold of 35 dBZ and in different climate regions. The table shows that Trans-CodeNet performs consistently well across different climate regions, particularly in the temperate desert (BWk) and temperate arid continental (Dwa) regions. In the CSI (Conditional Sound Index), Trans-CodeNet significantly outperforms other models in the BWk and Dwa regions, demonstrating its accurate capture of severe convective weather. In the HSS (Highly Sensitive Situation Index), Trans-CodeNet achieves the highest score in all regions, significantly outperforming RainFormer and EarthFormer in the Dwa region, indicating a higher predictive accuracy for severe convective weather. In the POD (Positive Disturbances Index), Trans-CodeNet leads other models in the Dwa region, significantly improving its ability to identify severe convective events. Meanwhile, in the FAR (Failure Rate) negative index, Trans-CodeNet maintains a low level, especially in the polar climate (ET) region, indicating a significant advantage in reducing false alarms. Overall, Trans-CodeNet is suitable for severe convective weather forecasting tasks in multiple climate regions, and can provide more accurate and stable forecast results.

[0150] Table 2 compares the model's performance in different climate zones at 35 dBZ.

[0151]

[0152] To more intuitively demonstrate the model's forecast performance, a sample was selected for each climate zone for visualization. The visualization results are as follows: Figure 2 , Figure 3 , Figure 4 , Figure 5As shown in the figure, the proposed model demonstrates superior performance in severe convective weather forecasting. Compared to other comparative models, our model more accurately predicts the distribution of high-echo areas, with predictions closer to observations. Furthermore, in long-term forecasts (such as 90-minute and 120-minute forecasts), the model maintains high accuracy, successfully capturing the overall distribution trend and detailed changes in radar echoes. Other comparative models, however, show significant shortcomings in characterizing high-echo areas during these time periods, exhibiting substantial biases in their predictions.

[0153] The forecast model outputs 12 frames. To more intuitively demonstrate the forecast performance of each model at different time frames, this paper uses line graphs to show the four key indicators of each model at different forecast time points: CSI, HSS, POD, and FAR. Figure 6 As shown, these line graphs facilitate a visual comparison of the performance of different models across various forecast periods. The line graphs focus on the results at 35 dBZ. The graphs demonstrate that the proposed model exhibits superior performance in long-term time-series forecasts, particularly maintaining high CSI, HSS, and POD indices for longer forecast periods (e.g., 60 minutes and above), showing a significant advantage over other comparative models. This indicates that the proposed model possesses strong capabilities in capturing temporal features and can effectively predict radar data at multiple future time points.

[0154] In summary, the codebook-based severe convective weather forecasting model proposed in this paper demonstrates excellent performance across multiple evaluation metrics, showcasing its comprehensive advantages in severe convective weather forecasting. It exhibits significant performance improvements, particularly in high-intensity echo areas and long-term forecasts. Compared to various existing classical models, this model not only more accurately captures the spatial distribution and key characteristics of severe convective weather but also demonstrates strong generalization ability and adaptability, providing a new solution for forecasting complex and ever-changing severe convective weather.

[0155] Example 2

[0156] This embodiment provides a large-area severe convective weather forecasting system pre-trained with a multi-climate zone memory codebook.

[0157] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the aforementioned large-area severe convection forecasting method pre-trained with a multi-climate-zone memory codebook.

[0158] A terminal device includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions adapted for loading and execution by the processor of the large-area severe convection forecasting method pre-trained with a multi-climate-zone memory codebook.

[0159] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A large-area severe convection forecasting method using multi-climate zone memory codebook pre-training, characterized in that, include: Acquire radar weather data; Data preprocessing is performed based on the acquired radar weather data; A universal codebook representation for multiple climate regions is constructed based on the Köppen climate classification method; Feature extraction based on a universal codebook representation for multiple climate regions to generate a universal climate representation; Spatiotemporal dynamic features are obtained by decoupling the spatiotemporal characteristics of general climate representation based on the improved Transformer. Strong convection forecast results are generated by dual-path characterization enhancement and progressive decoding based on spatiotemporal dynamic features. The data preprocessing based on the acquired radar weather data includes: performing sliding sampling on the global radar image with a fixed step size to extract local area images of a preset size, thereby reducing computational overhead and ensuring that the sampling results cover different climate zones; constructing a spatiotemporal sequence sample by extracting radar images at 19 consecutive time points along the time dimension based on a sliding window strategy, and removing samples with missing time points to ensure the continuity and integrity of the spatiotemporal sequence data; wherein, outlier screening is performed on the radar images in the spatiotemporal sequence sample, removing radar images that do not contain storm information and contain outliers; for the retained radar images, Pixel values ​​marked as abnormal or missing are replaced with 0dBZ; simultaneously, the effective range of radar reflectivity is set to truncate the data to eliminate interference from low-value weak signals; then, positive and negative samples are divided and data is balanced, and the total number of storm information points contained in the spatiotemporal sequence sample is counted. If the total number is greater than or equal to a preset threshold, the spatiotemporal sequence sample is classified as a positive sample; otherwise, it is classified as a negative sample; to alleviate the data imbalance caused by the long-tail distribution of severe convective weather, negative samples are randomly undersampled; finally, the Min-Max normalization method is used to normalize the processed spatiotemporal sequence samples. The method for constructing a universal codebook representation for multiple climate regions based on the Köppen climate classification method includes an improvement over the traditional single-time-series input-output mechanism for representation learning, which is consistent with downstream forecasts and features input data. It is a five-dimensional tensor, where 𝐵 is the number of samples and T is the time dimension. For the number of channels, and The spatial dimensions are represented by a 3D convolutional network encoder used to compress the features of the input data, as shown below: Where 𝑄 represents the encoded query vector, and 𝑁 represents the number of query vectors. For the feature dimensions of the query vector, The encoder operation is represented; then, between the encoder and decoder, a set of learnable discrete feature vectors is introduced to construct an independent codebook for each climate region. The codebook is initialized randomly and its representational ability is dynamically updated during the training process. The eight climate codebook sets after pre-training are represented as follows: The codebook for each climate region is as follows: ,in For the first A collection of codebooks for each climate region. For the first in the codebook 1 eigenvector The feature dimension of the codebook vector. The number of feature vectors in the codebook for each climate region is determined, and a dynamic codebook matching mechanism is introduced. This mechanism utilizes a Query-Key-Value structure for dynamic feature matching. First, the codebook... Mapping to the key space and value space through two independent linear transformations: ,in, The weight matrix is ​​a learnable matrix. A unified feature dimension for keys and values, then the query vector output by the encoder. As a query, a representative feature is extracted from the current climate codebook through an attention mechanism. This mechanism calculates the similarity between the query and the key vector, selects the value vector most relevant to the input, and extracts features from the codebook. ,in, This is a dynamic codebook operation, where N is the number of query vectors. This is a unified feature dimension for both key and value vectors, and serves as a scaling factor. The query vector is dynamically generated by the encoder. The features that are mapped from the codebook to the key space through a linear transformation. The features that are mapped from the codebook to the value space through a linear transformation. For from the first Feature representations extracted from codebooks of climate regions This represents the number of query vectors; finally, the decoder section uses layer-by-layer upsampling to restore the features retrieved from the codebook to radar data at the original resolution. ,in, This represents the reconstructed radar timing data. This represents the layer-by-layer upsampling operation of the decoder. During the decoding process, the layer-by-layer convolution operation can refine the spatial structure information, preserve and restore the complex spatial details in severe convective weather.

2. The method for large-area severe convection forecasting using a multi-climate zone memory codebook pre-training as described in claim 1, characterized in that, The process of extracting features from a multi-climate region universal codebook to generate a universal climate representation includes using a pre-trained climate region codebook. Deployed as an external memory module, and to achieve effective integration between the codebook and the prediction model, a 3D convolutional encoder with the same structure as the pre-training stage is used to process the radar echo sequence input to the prediction model. Perform feature compression and convert the query vector generated by the encoder. Simultaneously, it is sent to eight pre-trained climate region codebooks, and an independent dynamic matching retrieval operation is performed on each codebook. For the first... Codebook for each climate zone Through a dynamic codebook matching mechanism, the local representations most relevant to the climate characteristics of the region are retrieved and extracted. : ,in, The input sequence to the prediction model is processed by a 3D convolutional encoder to generate query vectors, where N is the number of query vectors. This is a unified feature dimension for both key and value vectors. For deployment as an external memory module A pre-trained climate region codebook, CodebookRetrieve(⋅) represents a dynamic codebook retrieval operation based on an attention mechanism, which obtains a set of 8 local climate feature representations through parallel retrieval: Each local representation It contains information on strong convection patterns specific to the corresponding climate region. Then, an adaptive fusion strategy based on an attention mechanism is constructed to dynamically evaluate and weightedly integrate eight local climate representations. Specifically, the eight local features are first concatenated along the feature dimension: The concatenated features are then mapped to the attention space through a learnable linear transformation. ,in, For learnable parameters, This represents the unified feature dimension of the multi-head attention fusion space. This is a tensor formed by splicing together the local features of eight climate regions along the feature dimension. The concatenated features are linearly mapped and then enter the attention space as a query, key, and value matrix. Finally, an attention mechanism adaptively assigns fusion weights to the representations of different climate regions, generating a unified global climate representation. Where N represents the number of query vectors, This is the output of a global climate general characterization.

3. The method for large-area severe convection forecasting using multi-climate zone memory codebook pre-training as described in claim 2, characterized in that, The method for decoupling spatiotemporal features of general climate representation based on the improved Transformer to obtain spatiotemporal dynamic features includes, to extract temporal features, employing a frame-level patch segmentation strategy to treat the time dimension T as an independent dimension, and performing patch segmentation for each frame of radar image. The input data is initially a radar sequence containing the time dimension. Where 𝐵 is the number of samples and T is the time dimension. For the number of channels, and Each frame represents a spatial dimension and a temporal dimension. It is divided into fixed-size patches and mapped to a high-dimensional feature space: where the spatial size of the patch is set to be... By using a sliding window or an equivalent two-dimensional convolution operation, the spatial dimensions are divided into... For the nth non-overlapping local image patch, The first frame Image blocks It is flattened and mapped to the feature dimension through linear projection. : ,in, It is the first The serialized patch feature of the i-th image patch in a frame after mapping to a high-dimensional feature space. It is the first The first frame Local image patch Indicates the flattening operation. It is a learnable linear projection matrix. As the bias vector, each frame is thus transformed into a serialized patch feature. To preserve temporal and spatial location information, temporal and spatial location codes were subsequently injected into the Patch features. Considering the spatiotemporal evolution characteristics of severe convective weather, spatial location codes were used. Used to label the absolute spatial coordinates and temporal location encoding of each patch in the original 2D image. The injection and fusion process, used to distinguish the chronological order of different time frames, is represented as follows: ,in, It is the first to fuse location information Frame integrity features It is a spatial location encoding, marking the absolute spatial coordinates of the patch. It uses time-position encoding to distinguish the chronological order of different time frames, and then encodes all time dimensions. The feature sequences of the frames are concatenated to obtain the complete input representation: ,in, It is all in the time dimension The complete input representation is obtained after frame concatenation; then, the encoded data is processed through a global attention mechanism to construct a global representation of the input data and apply input features. Perform multi-head self-attention computation, and apply different linear mappings to the input features. Projected into query matrices respectively Key matrix Sum matrix : ,in, For the total number of attention heads, For the first The learnable weight matrix corresponding to each attention head For each attention head, feature dimensions are defined, and scaled dot product attention is computed independently for each attention head to dynamically evaluate the interaction and correlation strength between different spatiotemporal patches: ,in, For the first A query, key, and value matrix for each attention head. For the first The outputs of each attention head are dynamically evaluated for spatiotemporal patch relevance. Finally, the outputs of all attention heads are concatenated along the feature dimension and mapped back to the original dimension through the output projection matrix to generate an intermediate feature representation containing global spatiotemporal dependencies. : , in, This is the learnable weight matrix for the global attention fusion stage. This is an intermediate feature representation that contains global spatiotemporal dependencies, generated after processing by a multi-head self-attention mechanism.

4. The method for large-area severe convection forecasting using multi-climate zone memory codebook pre-training as described in claim 3, characterized in that, The method described above, based on the improved Transformer, decouples spatiotemporal features of general climate representation to obtain spatiotemporal dynamic features. It also includes introducing two decoupling mechanisms—temporal attention and spatial attention—to address the heterogeneity of temporal and spatial patterns in severe convective weather. These mechanisms handle temporal and spatial features separately. To balance computational complexity, a parallel temporal and spatial attention mechanism based on head allocation is constructed, integrating the features of the multi-head attention mechanism. Each attention head is split into two groups: Size allocation is used for spatial attention, focusing on modeling dependencies between different spatial locations within a single frame; the rest... Size allocation is used for temporal attention, focusing on modeling the dynamic evolution of the same spatial location across different time frames. For the spatial attention branch, the input features are reshaped into a sequence of spatial patches, allowing attention computation to be performed only in the spatial dimension. First, a query matrix is ​​generated through a linear transformation. Key matrix Sum matrix : ,in, The learnable weight matrix corresponding to the spatial branch. For the query, key, and value matrix of the spatial attention branch, To reshape the input features into spatial branches in units of spatial patches, A learnable weight matrix specific to the spatial branch is then calculated; subsequently, spatial self-attention weights are calculated and the value matrix is ​​weighted and aggregated. ,in, The spatial branch output is used to capture local structural features at different geographical locations within the same timeframe. In the computation, the dimension of the attention weight matrix is... It is specifically designed to capture the spatial correlation and local structural features between radar echoes from different geographical locations within the same time period, and similarly calculates the output of the temporal attention branch. Finally, the outputs of the two branches are merged through a concatenation operation: in, A learnable linear projection matrix is ​​used to remap the concatenated spatiotemporal features to a unified feature space. After decoupling the temporal and spatial attention calculations, a global attention mechanism is introduced again to refine and integrate the decoupled and fused spatiotemporal features, forming complete spatiotemporal dynamic features that contain both local details and global semantics. The final output This refers to the extracted local spatiotemporal dynamic features. It is the spatiotemporal feature resulting from the splicing and remapping of time and space branches.

5. The method for large-area severe convection forecasting using multi-climate zone memory codebook pre-training as described in claim 4, characterized in that, The process of enhancing strong convection forecasts through dual-path representation and progressive decoding based on spatiotemporal dynamic features includes receiving heterogeneous feature representations from two independent paths as input, namely, local spatiotemporal dynamic features extracted through spatiotemporal decoupling. and global climate general characterization To achieve the organic integration of the two types of heterogeneous features, dimensional alignment is first performed to integrate the global climate characterization. Broadcasting extension is performed in both the sample and time dimensions to align with spatiotemporal dynamic features. Maintain consistency in shape: Then, the aligned two types of features are concatenated along the feature dimension: ,in, This is the result of concatenating temporal dynamic features and prior climate features along the feature dimension. It is a local spatiotemporal dynamic feature. To provide a universal representation of global climate, and to achieve deep integration, a lightweight cross-attention mechanism is introduced. Using features from path one as the query and features from path two as the key and value, the mechanism dynamically selects and enhances the climate prior information most relevant to the current spatiotemporal dynamics. Where Q is set to For multi-head cross-attention operations: ,in, To dynamically select and enhance relevant prior climate information, the output fusion features include both local dynamics and global climate information. The scaling factor is the feature dimension; the final output is the fused feature. It contains information on both local dynamics and climate a priori aspects.

6. The method for large-area severe convection forecasting using multi-climate zone memory codebook pre-training as described in claim 5, characterized in that, The method of generating strong convection forecast results by performing dual-path representation enhancement and progressive decoding based on spatiotemporal dynamic features also includes fusing the enhanced features. The data is fed into the decoder, where a progressive upsampling strategy is used to restore the spatial and temporal resolution layer by layer, ultimately generating a radar echo prediction sequence for future moments. First, the fused features are reconstructed from a two-dimensional sequence back into a three-dimensional spatial structure, represented as: , in, The decoder consists of L layers, where L is the number of patches per frame at the current scale. Each layer performs two operations sequentially: feature refinement and spatial upsampling. For the Lth... layer( =1,2,…,L) The output of the previous layer is updated using the Transformer decoder module: ,in For decoder number The output feature tensor of the layer This is the output of the layer above the decoder. For the first The Transformer module of the layer decoder, which includes a multi-head self-attention mechanism and a feedforward network, is used to further extract and integrate spatiotemporal information from features at the current scale. For the first Number of patches per layer For the first The feature dimension of the layer is first determined, and then the spatial resolution of the feature map is gradually restored through an upsampling layer. Specifically, the refined features are then processed. Perform an upsampling operation: , in, This can be achieved using transposed convolution, after upsampling. As the spatial dimensions increase, the feature dimensions are adjusted accordingly, thereby restoring the spatial details of the radar echo layer by layer.

7. The method for large-area severe convection forecasting using multi-climate zone memory codebook pre-training as described in claim 6, characterized in that, The method for enhancing the dual-path representation and progressively decoding based on spatiotemporal dynamic features to generate severe convection forecast results also includes constructing a weighted loss function based on weights to improve the model's prediction accuracy for key areas of severe convection. First, differentiated weights are assigned to different pixels based on radar echo intensity. , in, Given the radar echo value of the corresponding pixel, the prediction error in the high echo value region accounts for a larger proportion in the loss calculation. Based on the above weights, the weighted mean square error and the weighted average absolute error are defined as follows: , , in, To predict the number of sequences, This represents the resolution of a single-frame radar image. This represents the radar echo value of the corresponding pixel. and These are the actual value and the predicted value, respectively. The final comprehensive loss function is a weighted combination of the two: ,in, and These are weighting coefficients used to adjust the relative importance of the two types of losses in the overall optimization objective. and These represent the weighted mean square error and the weighted average absolute error, respectively. This represents the total number of predicted sequences involved in the error calculation. This represents the horizontal spatial resolution of a single radar frame. Differential weights are dynamically assigned based on pixel echo intensity.

8. A large-area severe convective weather forecasting system pre-trained with a multi-climate codebook, executing the large-area severe convective weather forecasting method pre-trained with a multi-climate codebook as described in claim 1, characterized in that... include: The data acquisition module is configured to acquire radar weather data; The preprocessing module is configured to perform data preprocessing based on the acquired radar weather data; The codebook representation module is configured to construct a universal codebook representation for multiple climate regions based on the Köppen climate classification method. The general characterization module is configured to extract features based on the general codebook characterization of multiple climate regions to generate a general climate characterization. The dynamic feature module is configured to obtain spatiotemporal dynamic features by decoupling the spatiotemporal features of the general climate representation based on the improved Transformer. The decoding module is configured to perform dual-path characterization enhancement and progressive decoding based on spatiotemporal dynamic features to generate strong convection forecast results.

Citation Information

Patent Citations

  • Weather radar reflectivity synthesis method and system based on synchronous stationary satellite

    CN121559504A

  • Expert demonstration-based entropy sensing robot control method and system

    CN122143061A