Deep learning thunder and lightning prediction method and system based on space-time attention mechanism, and storage medium

By introducing a deep learning model of the spatiotemporal attention mechanism in lightning prediction, combined with CNN and GRU, the limitations of spatiotemporal data processing in the prior art are solved, and the accuracy and reliability of lightning prediction are improved.

CN120030329APending Publication Date: 2025-05-23WUHAN NARI LIABILITY OF STATE GRID ELECTRIC POWER RES INST +2
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510170672.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing deep learning-based lightning prediction methods have limitations when processing spatiotemporal data. CNN ignores the time dimension, while RNN is prone to gradient disappearance or explosion when processing long time series, resulting in the inability to effectively capture long-term time dependencies.

Method used

The deep learning model based on the spatiotemporal attention mechanism is adopted, combined with CNN and GRU, and the attention weight of the model for different spatiotemporal attention mechanisms is adaptively adjusted through the spatiotemporal attention mechanism, which not only considers the information in the time dimension, but also captures the features in the space dimension.

Benefits of technology

It improves the accuracy and reliability of lightning prediction, can better process spatiotemporal data, and enhances the ability to identify areas and critical periods of lightning. Compared with traditional numerical forecasting methods, the predicted hit rate is increased by 15% and the false alarm rate is reduced by 10%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030329A_ABST
    Figure CN120030329A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning thunder and lightning prediction method and system based on a space-time attention mechanism, and a storage medium. The method comprises the steps that historical thunder and lightning grid data, satellite cloud picture data and regional numerical forecasting grid data serve as input features, historical approaching thunder and lightning data serve as forecasting labels, and a training data set is constructed; inputting the input features into a deep learning thunder prediction model, extracting the spatial features of the input features by using a CNN convolutional neural network, capturing the time dependence of the input features through a GRU gating circulation unit, adaptively adjusting the attention weights of the model for different space-time regions based on a space-time attention mechanism, and taking a prediction label as a target output. Training the deep learning thunder and lightning prediction model; and inputting the real-time thunder and lightning grid data, the satellite cloud picture data and the regional numerical forecasting grid data into the trained model, and outputting an approaching thunder and lightning forecasting result. The method can better process the spatio-temporal data, and improves the accuracy and reliability of lightning prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lightning prediction, and specifically relates to a deep learning lightning prediction method, system and storage medium based on a spatiotemporal attention mechanism. Background Art

[0002] Lightning is a common natural phenomenon, but its harmfulness cannot be ignored. The strong current and high voltage generated by lightning can cause serious damage and harm to buildings, power equipment and personnel, and even cause dangerous events such as fires and explosions. Therefore, accurate prediction of lightning activity is of great significance to ensuring public safety and reducing disaster losses. In the field of lightning prediction, traditional prediction methods are mainly based on meteorological observation data and statistical models. Although they can predict lightning activity to a certain extent, their prediction accuracy and real-time performance need to be improved. In recent years, with the continuous development of deep learning technology, prediction models based on neural networks have gradually become a research hotspot in the field of lightning prediction.

[0003] At present, some lightning prediction methods based on deep learning have been proposed, such as using convolutional neural networks (CNN) to extract spatial features, or using recurrent neural networks (RNN) to capture temporal dependencies. However, these methods still have certain limitations when processing spatiotemporal data. CNN mainly focuses on the extraction of spatial features, but ignores information in the time dimension; while RNN is prone to gradient vanishing or explosion problems when processing long time series, resulting in the inability to effectively capture long-term temporal dependencies. Because lightning activities in different time and space have different impacts on the prediction of nearby lightning. Some key time and space locations may contain more useful information, while other locations may contain less noise or irrelevant information. Therefore, a deep learning model that can simultaneously consider time and space information is needed to improve the accuracy of lightning prediction, and adaptively adjust the degree of attention to different time and space locations according to the characteristics of the input data to meet the needs of different prediction scenarios. Summary of the invention

[0004] In view of this, the present invention provides a deep learning lightning prediction method, system and storage medium based on spatiotemporal attention mechanism.

[0005] The technical solution adopted by the present invention is: a deep learning lightning prediction method based on spatiotemporal attention mechanism, comprising:

[0006] The historical lightning grid data, satellite cloud image data and regional numerical forecast grid data are used as input features, and the historical nearby lightning data are used as forecast labels to construct a training data set.

[0007] Input features into the deep learning lightning prediction model, use CNN convolutional neural network to extract the spatial features of the input features, capture the temporal dependency of the input features through GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and use the forecast label as the target output to train the deep learning lightning prediction model;

[0008] The real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data are input into the trained model to output the upcoming lightning prediction results.

[0009] It also includes preprocessing of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; rasterizing lightning grid data with different time resolutions to align their time resolution with the time resolution of satellite cloud image data; adjusting the time resolution of regional numerical forecast grid data to make it consistent with the time resolution of satellite cloud image data.

[0010] The preprocessing process also includes: data cleaning of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data to eliminate noise and outliers; standardization and normalization of lightning grid data, satellite cloud image data and regional numerical forecast grid data, mapping the data to the same numerical range to eliminate the influence of dimension and numerical range.

[0011] The deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure:

[0012] Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data;

[0013] Convolutional layer: used to extract spatial features of input data;

[0014] Pooling layer: used to reduce the dimension of the convolutional layer output;

[0015] GRU layer: used to capture the temporal dependency of input data;

[0016] Fully connected layer: used to map the GRU layer output to the prediction space;

[0017] Output layer: used to output the upcoming lightning prediction results.

[0018] The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is:

[0019]

[0020] Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation; b is the bias term, y i,j is the output feature map after convolution.

[0021] The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is:

[0022]

[0023] Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

[0024] The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically:

[0025] z t =σ(w z ·[h t-1 ,x t ]);

[0026] r t =σ(w r ·[h t-1 ,x t ]);

[0027]

[0028] Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h t is the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

[0029] The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as:

[0030] y=σ(w x + b);

[0031] Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

[0032] The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module, which is used to adaptively adjust the model's attention to spatiotemporal data, and its expression is:

[0033] Calculate the attention score:

[0034] socre=X·W q ·X T ;

[0035] Among them, W q is a learnable weight matrix used to map the input data X to the query space;

[0036] Normalized attention score:

[0037] α = softmax(score);

[0038] Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length;

[0039] Compute the weighted representation:

[0040] Y=α·X·W v ;

[0041] Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

[0042] The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

[0043] The present invention also discloses a deep learning lightning prediction system based on a spatiotemporal attention mechanism, wherein the deep learning lightning prediction system based on a spatiotemporal attention mechanism comprises a data acquisition module, a data set construction module, a network training module and a lightning prediction module.

[0044] Data acquisition module, used to collect and integrate historical lightning grid data, satellite cloud image data, regional numerical forecast grid data and historical approaching lightning data;

[0045] The data set construction module is used to construct a training data set by taking historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input features and historical nearby lightning data as forecast labels;

[0046] The network training module is used to input the input features into the deep learning lightning prediction model, extract the spatial features of the input features using the CNN convolutional neural network, capture the temporal dependency of the input features through the GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and train the deep learning lightning prediction model with the forecast label as the target output;

[0047] The lightning prediction module is used to input real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data into the trained model and output the upcoming lightning prediction results.

[0048] It also includes preprocessing of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; rasterizing lightning grid data with different time resolutions to align their time resolution with the time resolution of satellite cloud image data; adjusting the time resolution of regional numerical forecast grid data to make it consistent with the time resolution of satellite cloud image data.

[0049] The preprocessing process also includes: data cleaning of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data to eliminate noise and outliers; standardization and normalization of lightning grid data, satellite cloud image data and regional numerical forecast grid data, mapping the data to the same numerical range to eliminate the influence of dimension and numerical range.

[0050] The deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure:

[0051] Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data;

[0052] Convolutional layer: used to extract spatial features of input data;

[0053] Pooling layer: used to reduce the dimension of the convolutional layer output;

[0054] GRU layer: used to capture the temporal dependency of input data;

[0055] Fully connected layer: used to map the GRU layer output to the prediction space;

[0056] Output layer: used to output the upcoming lightning prediction results.

[0057] The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is:

[0058]

[0059] Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation; b is the bias term, y i,j is the output feature map after convolution.

[0060] The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is:

[0061]

[0062] Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

[0063] The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically:

[0064] z t =σ(w z ·[h t-1 ,x t ]);

[0065] r t =σ(w r ·[h t-1 , x t ]);

[0066]

[0067] Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h t is the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

[0068] The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as:

[0069] y=σ(w x + b);

[0070] Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

[0071] The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module, which is used to adaptively adjust the model's attention to spatiotemporal data, and its expression is:

[0072] Calculate the attention score:

[0073] socre=X·W q ·X T ;

[0074] Among them, W q is a learnable weight matrix used to map the input data X to the query space;

[0075] Normalized attention score:

[0076] α = softmax(score);

[0077] Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length;

[0078] Compute the weighted representation:

[0079] Y=α·X·W v ;

[0080] Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

[0081] The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

[0082] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the deep learning lightning prediction method based on the spatiotemporal attention mechanism are implemented.

[0083] The present invention uses historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input forecast factors and nearby lightning as forecast labels to perform lightning forecasting. Among them, a spatiotemporal attention mechanism is introduced to adaptively adjust the attention weight, which not only considers the information in the time dimension, but also captures the features in the spatial dimension. By combining the advantages of convolutional neural networks and gated recurrent unit neural networks, the model proposed by the present invention can better process spatiotemporal data and improve the accuracy and reliability of lightning forecasting. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0085] Figure 1 It is a flowchart of Embodiment 1 of the deep learning lightning prediction method based on the spatiotemporal attention mechanism of the present invention;

[0086] Figure 2 This is a functional module diagram of Embodiment 2 of a deep learning lightning prediction system based on a spatiotemporal attention mechanism of the present invention;

[0087] Figure 3 Schematic diagram of the deep learning lightning prediction model. DETAILED DESCRIPTION

[0088] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.

[0089] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0090] Embodiment 1:

[0091] like Figure 1 As shown, the present invention provides a deep learning lightning prediction method based on spatiotemporal attention mechanism, comprising:

[0092] The historical lightning grid data, satellite cloud image data and regional numerical forecast grid data are used as input features, and the historical nearby lightning data are used as forecast labels to construct a training data set.

[0093] Input features into the deep learning lightning prediction model, use CNN convolutional neural network to extract the spatial features of the input features, capture the temporal dependency of the input features through GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and use the forecast label as the target output to train the deep learning lightning prediction model;

[0094] The real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data are input into the trained model to output the upcoming lightning prediction results.

[0095] It also includes preprocessing of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; rasterizing lightning grid data with different time resolutions to align their time resolution with the time resolution of satellite cloud image data; adjusting the time resolution of regional numerical forecast grid data to make it consistent with the time resolution of satellite cloud image data.

[0096] In some optional embodiments, discrete lightning grid data with a spatial accuracy of hundreds of meters are rasterized so that their spatial resolution is aligned with the 0.05° resolution of regional numerical forecast grid data; discrete lightning grid data with a time accuracy of microseconds are rasterized so that their time resolution is aligned with the 10-minute resolution of satellite cloud image data; regional numerical forecast grid data with a time resolution of 1 hour are processed using a nonlinear interpolation method to adjust their time resolution to 10 minutes; and low-quality data are eliminated through an anomaly detection algorithm.

[0097] The preprocessing process also includes: data cleaning of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data to eliminate noise and outliers; standardization and normalization of lightning grid data, satellite cloud image data and regional numerical forecast grid data, mapping the data to the same numerical range to eliminate the influence of dimension and numerical range.

[0098] like Figure 3 As shown, the deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure:

[0099] Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data;

[0100] Convolutional layer: used to extract spatial features of input data;

[0101] Pooling layer: used to reduce the dimension of the convolutional layer output;

[0102] GRU layer: used to capture the temporal dependency of input data;

[0103] Fully connected layer: used to map the GRU layer output to the prediction space;

[0104] Output layer: used to output the upcoming lightning prediction results.

[0105] The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is:

[0106]

[0107] Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation, which is used to represent the row index of the local area traversed by the convolution kernel when sliding on the input data; b is the bias term, y i,j is the output feature map after convolution.

[0108] The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is:

[0109]

[0110] Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

[0111] The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically:

[0112] Z:=0(w.[1×D

[0113] r; = 0(w,[1])

[0114]

[0115] Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h tis the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

[0116] The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as:

[0117] y=σ(w x + b);

[0118] Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

[0119] The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module for adaptively adjusting the model's attention to spatiotemporal data. Assuming that the input data is X, its dimension is [batch_size, seq_length, feature_dim], where batch_size represents the batch size, seq_length represents the sequence length, and feature_dim represents the feature dimension. Its expression is:

[0120] Calculate the attention score:

[0121] socre=X·W q ·X T ;

[0122] Among them, W q is a learnable weight matrix used to map the input data X to the query space;

[0123] Normalized attention score:

[0124] α = softmax(score);

[0125] Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length;

[0126] Compute the weighted representation:

[0127] Y=α·X·W v ;

[0128] Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

[0129] In specific implementation, the attention mechanism generates an attention weight map, which is the same size as the input data, and the value at each position represents the attention weight of the corresponding position. The attention weight map is multiplied element by element with the original feature map to obtain the attention-weighted feature map. This feature map is then input into the fully connected layer, and the nearby lightning prediction result is output through the activation function.

[0130] The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

[0131] Through the above implementation, historical lightning grid data, satellite cloud images and regional numerical forecast grid data are used as dual-source input forecasting factors, and the attention weight is adaptively adjusted using the spatiotemporal attention mechanism, combined with a deep learning model based on a convolutional neural network and a gated recurrent unit neural network, to achieve accurate prediction of the approaching (next 2 hours) lightning. The prediction method of the present invention not only improves the prediction accuracy, but also fully considers the correlation between time and space.

[0132] This example combines CNN to extract the spatial distribution characteristics of lightning activity areas (such as the development and movement area of ​​thunderstorm clouds), GRU to capture the temporal evolution of lightning (lightning development trend), and automatically identifies high-incidence areas and key periods of time (such as the development stage of severe convection) through the spatiotemporal attention mechanism, and assigns higher weights. The experimental results show that compared with traditional numerical forecasting methods, the prediction hit rate of key areas under complex meteorological conditions (such as typhoon edge thunderstorms) of this comprehensive model has increased by 15% and reduced the false alarm rate by 10%.

[0133] Embodiment 2:

[0134] like Figure 2 As shown, the present invention provides a deep learning lightning prediction system based on the spatiotemporal attention mechanism, and the deep learning lightning prediction system based on the spatiotemporal attention mechanism includes a data acquisition module, a data set construction module, a network training module and a lightning prediction module.

[0135] Data acquisition module, used to collect and integrate historical lightning grid data, satellite cloud image data, regional numerical forecast grid data and historical approaching lightning data;

[0136] The data set construction module is used to construct a training data set by taking historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input features and historical nearby lightning data as forecast labels;

[0137] The network training module is used to input the input features into the deep learning lightning prediction model, extract the spatial features of the input features using the CNN convolutional neural network, capture the temporal dependency of the input features through the GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and train the deep learning lightning prediction model with the forecast label as the target output;

[0138] The lightning prediction module is used to input real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data into the trained model and output the upcoming lightning prediction results.

[0139] It also includes preprocessing of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; rasterizing lightning grid data with different time resolutions to align their time resolution with the time resolution of satellite cloud image data; adjusting the time resolution of regional numerical forecast grid data to make it consistent with the time resolution of satellite cloud image data.

[0140] The preprocessing process also includes: data cleaning of historical lightning grid data, satellite cloud image data and regional numerical forecast grid data to eliminate noise and outliers; standardization and normalization of lightning grid data, satellite cloud image data and regional numerical forecast grid data, mapping the data to the same numerical range to eliminate the influence of dimension and numerical range.

[0141] The deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure:

[0142] Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data;

[0143] Convolutional layer: used to extract spatial features of input data;

[0144] Pooling layer: used to reduce the dimension of the convolutional layer output;

[0145] GRU layer: used to capture the temporal dependency of input data;

[0146] Fully connected layer: used to map the GRU layer output to the prediction space;

[0147] Output layer: used to output the upcoming lightning prediction results.

[0148] The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is:

[0149]

[0150] Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation, which is used to represent the row index of the local area traversed by the convolution kernel when sliding on the input data; b is the bias term, y i,j is the output feature map after convolution.

[0151] The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is:

[0152]

[0153] Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

[0154] The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically:

[0155] z t =σ(w z ·[h t-1 , x t ]);

[0156] r t =σ(w r ·[h t-1 , x t ]);

[0157]

[0158] Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h t is the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

[0159] The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as:

[0160] y=σ(wx + b);

[0161] Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

[0162] The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module, which is used to adaptively adjust the model's attention to spatiotemporal data, and its expression is:

[0163] Calculate the attention score:

[0164] socre=X·W q ·X T ;

[0165] Among them, W q is a learnable weight matrix used to map the input data X to the query space;

[0166] Normalized attention score:

[0167] α = softmax(score);

[0168] Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length;

[0169] Compute the weighted representation:

[0170] Y=α·X·W v ;

[0171] Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

[0172] The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

[0173] Embodiment three:

[0174] The present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the deep learning lightning prediction method based on the spatiotemporal attention mechanism are implemented.

[0175] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0176] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0177] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit its protection scope. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the pending claims of the invention.

[0180] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.

Claims

1. A deep learning lightning prediction method based on spatiotemporal attention mechanism, characterized by: The historical lightning grid data, satellite cloud image data and regional numerical forecast grid data are used as input features, and the historical nearby lightning data are used as forecast labels to construct a training data set. Input features into the deep learning lightning prediction model, use CNN convolutional neural network to extract the spatial features of the input features, capture the temporal dependency of the input features through GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and use the forecast label as the target output to train the deep learning lightning prediction model; The real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data are input into the trained model to output the upcoming lightning prediction results.

2. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 1 is characterized in that: It also includes preprocessing historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; Rasterize lightning grid data with different time resolutions to align their time resolution with that of satellite cloud image data; The temporal resolution of the regional numerical forecast grid data is adjusted to be consistent with the temporal resolution of the satellite cloud image data.

3. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 1 is characterized in that: The deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure: Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data; Convolutional layer: used to extract spatial features of input data; Pooling layer: used to reduce the dimension of the convolutional layer output; GRU layer: used to capture the temporal dependency of input data; Fully connected layer: used to map the GRU layer output to the prediction space; Output layer: used to output the upcoming lightning prediction results.

4. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 3 is characterized in that: The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is: Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation; b is the bias term, y i,j is the output feature map after convolution.

5. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 3 is characterized in that: The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is: Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

6. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 3 is characterized in that: The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically: z t =σ(w z ·[h t-1 ,x t ]); r t =σ(w r ·[h t-1 ,x t ]); Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h t is the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

7. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 3 is characterized in that: The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as: y=σ(w x +b); Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

8. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 3 is characterized in that: The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module, which is used to adaptively adjust the model's attention to spatiotemporal data, and its expression is: Calculate the attention score: socre=X·W q ·X T ; Among them, W q is a learnable weight matrix used to map the input data X to the query space; Normalized attention score: α = softmax(score); Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length; Compute the weighted representation: Y=α·X·W v ; Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

9. The deep learning lightning prediction method based on spatiotemporal attention mechanism according to claim 1 is characterized in that: The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

10. A deep learning lightning prediction system based on spatiotemporal attention mechanism, characterized by: The deep learning lightning prediction system based on spatiotemporal attention mechanism includes a data acquisition module, a data set construction module, a network training module and a lightning prediction module. Data acquisition module, used to collect and integrate historical lightning grid data, satellite cloud image data, regional numerical forecast grid data and historical approaching lightning data; The data set construction module is used to construct a training data set by taking historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input features and historical nearby lightning data as forecast labels; The network training module is used to input the input features into the deep learning lightning prediction model, extract the spatial features of the input features using the CNN convolutional neural network, capture the temporal dependency of the input features through the GRU gated recurrent unit, adaptively adjust the model's attention weights on different spatiotemporal regions based on the spatiotemporal attention mechanism, and train the deep learning lightning prediction model with the forecast label as the target output; The lightning prediction module is used to input real-time lightning grid data, satellite cloud image data and regional numerical forecast grid data into the trained model and output the upcoming lightning prediction results.

11. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 10, characterized in that: It also includes preprocessing historical lightning grid data, satellite cloud image data and regional numerical forecast grid data. The preprocessing process is: rasterizing lightning grid data with different spatial resolutions to align their spatial resolution with the resolution of regional numerical forecast grid data; Rasterize lightning grid data with different time resolutions to align their time resolution with that of satellite cloud image data; The temporal resolution of the regional numerical forecast grid data is adjusted to be consistent with the temporal resolution of the satellite cloud image data.

12. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 10, characterized in that: The deep learning lightning prediction model is a hybrid model based on CNN convolutional neural network and GRU gated recurrent unit neural network, including the following structure: Input layer: receives historical lightning grid data, satellite cloud image data and regional numerical forecast grid data as input data; Convolutional layer: used to extract spatial features of input data; Pooling layer: used to reduce the dimension of the convolutional layer output; GRU layer: used to capture the temporal dependency of input data; Fully connected layer: used to map the GRU layer output to the prediction space; Output layer: used to output the upcoming lightning prediction results.

13. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 12, characterized in that: The convolution kernel size of the convolution layer is n×n, the step size is s, and the padding is p, and its expression is: Among them, x i,j is the pixel value of the input data; w m,n is the weight of the convolution kernel; m is the intermediate variable used to traverse the local area in the convolution operation; b is the bias term, y i,j is the output feature map after convolution.

14. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 12, characterized in that: The pooling layer adopts the maximum pooling operation, the pooling window size is k×k, the step size is s, and its expression is: Among them, x i,j is the pixel value of the input feature map; i,j is the output feature map after pooling; m′ and n′ represent the offset of rows and columns within the window in the maximum pooling operation.

15. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 12, characterized in that: The expression of the GRU layer includes the calculation process of the update gate, the reset gate and the candidate hidden state, specifically: z t =σ(w z ·[h t-1 ,x t ]); r t =σ(w r ·[h t-1 ,x t ]); Among them, h t-1 is the hidden state of the previous moment; x t is the input at the current moment; z t is the update gate; r t To reset the gate; is the candidate hidden state; h t is the hidden state at the current moment; σ is the activation function: w z 、w r , w are the weight matrices of the update gate, the reset gate, and the candidate hidden state, respectively.

16. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 12, characterized in that: The fully connected layer maps the output of the GRU layer to the prediction space and uses linear transformation and activation function for prediction, which is expressed as: y=σ(w x +b); Among them, x is the output of the GRU layer; W is the weight matrix; b is the bias term; σ is the activation function, and y is the prediction result.

17. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 12, characterized in that: The deep learning lightning prediction model also includes a spatiotemporal attention mechanism module, which is used to adaptively adjust the model's attention to spatiotemporal data, and its expression is: Calculate the attention score: socre=X·W q ·X T ; Among them, W q is a learnable weight matrix used to map the input data X to the query space; Normalized attention score: α = softmax(score); Among them, α is the normalized attention weight, satisfying seq_length indicates the sequence length; Compute the weighted representation: Y=α·X·W v ; Among them, W v is a learnable weight matrix used to map the input data X to the value space; α is the normalized attention weight; Y is the output data after weighted representation, and its dimension is the same as the input data X.

18. The deep learning lightning prediction system based on spatiotemporal attention mechanism according to claim 10, characterized in that: The model training adopts the back propagation algorithm and the stochastic gradient descent SGD algorithm, the loss function is the mean square error MSE loss function, and cross-validation and regularization techniques are used in the training process to prevent overfitting.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the deep learning lightning prediction method based on the spatiotemporal attention mechanism as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Method, device and equipment for predicting electric field in thunderstorm cloud based on deep learning algorithm

    CN121072356A

  • Hail short-term and temporary prediction method and system based on multi-modal fusion and space-time attention

    CN122218845A

  • Hail short-impending prediction method and system based on multi-modal fusion and space-time attention

    CN122218845B