A short-term rainfall prediction method based on an attention mechanism cascaded LSTM
By introducing cascading LSTMs and generative adversarial networks based on attention mechanisms into the short-prone rainfall prediction model, the problem that existing models cannot effectively extract global spatial features is solved, and higher prediction accuracy and visual detail retention are achieved.
Patent Information
- Application Number
- CN202310337088.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-31
AI Technical Summary
When processing radar echo data, the existing short-term rainfall prediction model cannot effectively extract global spatial characteristics, resulting in the inaccurate details of the prediction results in the high-intensity rainfall area.
A cascading LSTM model based on attention mechanism is adopted, combined with a generative adversarial network, local and global spatial features are integrated through attention mechanism, and the model is updated through adversarial loss and perceived learning loss to improve prediction accuracy.
The accuracy of short-term rainfall prediction is significantly improved, especially in high-intensity rainfall areas, which can better preserve the visual details of each frame and implement complex geometric constraints in the global image structure.
Smart Images

Figure CN116381690B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rainfall prediction, and particularly to a short-term rainfall prediction method based on an attention mechanism and cascaded LSTM. Background Art
[0002] Short-term rainfall prediction is one of the important research fields in weather forecasting. It can predict the rainfall intensity prediction value in the next 0-2 hours based on the current weather conditions, and the prediction results can reduce the losses caused by natural disasters such as debris flows and floods. However, short-term rainfall prediction has complexity, non-linearity, and randomness, which is still a very challenging task.
[0003] In recent years, traditional methods have proposed a radar echo tracking algorithm (Cross-correlation Multi-Scale Tracking Radar Echo, MTREC) to solve the problem of affecting the movement of radar echoes at different spatial scales. Most of these methods are linearly determined, while the weather system changes complexly. Non-linear changes need to be considered during prediction, and the prediction results are unstable.
[0004] With the rapid development of deep learning in computer vision, it has attracted the attention of many researchers, and deep learning has begun to be used to solve the problem of short-term rainfall prediction. The Recurrent Neural Network (RNN) and Long Short-Term Memory provide a sequence-to-sequence architecture design for short-term rainfall prediction based on radar echo extrapolation. However, for multi-dimensional graphs, there is complex spatial information and strong correlation between each point, which results in redundancy. Traditional LSTM cannot fully extract this spatial feature. FC-LSTM (Fully Convolutional LSTM) unfolds the data into one dimension for prediction. There is a large amount of redundant information in radar echo data, and FC-LSTM cannot process it. Moreover, FC-LSTM can only extract time series information and cannot extract spatial information. To solve the deficiencies in processing image information, ConvLSTM can not only establish the LSTM time series relationship but also utilize the spatial extraction ability of CNN. However, the convolutional recursive structure in the ConvLSTM model is location invariant, while natural motion and transformation (rotation) are usually location variant. TrajGRU (Trajectory Gated Unit Model) can actively utilize the structure based on location change to cycle and learn, and aggregate states by learning trajectories.
[0005] For spatio-temporal prediction tasks, as a recurrent network PredRNN model with cross-layer interaction, it has the characteristic of zigzag memory flow, which propagates in both bottom-up and top-down directions in all layers, enabling the dynamic visual information learned by different layers of the RNN to communicate with each other. However, this complex structure is still troubled by the problem of gradient vanishing. Through backpropagation through time, the magnitude of the gradient decays exponentially, which is due to the problem of gradient vanishing caused by long-term training and dependence. PredRNN++ uses the cascading operation of spatio-temporal units to increase the non-linear ability, which is beneficial to capturing short-term dynamic changes. At the same time, it uses GHU (Gradient Highway Unit) to connect the past time period and the future time period, and passes the important features through GHU to alleviate the problem of gradient vanishing.
[0006] Based on this, the present invention notices that one point not considered in these models is that due to the limitation of the convolutional kernel size, the convolutional kernel of the current frame can only extract local features and cannot rely on long-distance global spatial features, resulting in the lack of good display of the details of the predicted frame. Summary of the Invention
[0007] The purpose of the present invention is to provide a short-term rainfall prediction method based on an attention mechanism cascaded LSTM, so as to effectively improve the accuracy of short-term rainfall prediction.
[0008] The technical solution of the present invention to solve the above technical problems is as follows:
[0009] The present invention provides a short-term rainfall prediction method based on an attention mechanism cascaded LSTM. The short-term rainfall prediction method based on an attention mechanism cascaded LSTM includes:
[0010] S1: Clean and denoise the historical radar echo map data set at time t to obtain a denoised map data set;
[0011] S2: Input the denoised map data set into the SAC-LSTM model to obtain a predicted radar echo map data set at time t+1;
[0012] S3: Compare the MSE loss between the real radar echo map data set at time t+1 and the predicted radar echo map data set at time t+1 to obtain a first comparison result;
[0013] S4: Input the real radar echo map data set at time t+1 and the predicted radar echo map data set at time t+1 into the adversarial generation network respectively, perform convolutional operations to obtain a first feature map data set corresponding to the denoised map data set and a second feature map data set corresponding to the predicted radar echo map data set at time t+1;
[0014] S5: Compare the perceptual learning loss and adversarial loss of the first feature map dataset and the second feature map dataset to obtain a second comparison result;
[0015] S6: Update the loss function of the SAC-LSTM model using the first comparison result and the second comparison result to obtain an updated SAC-LSTM model;
[0016] S7: Use the updated SAC-LSTM model to predict the current input radar echo map dataset to obtain a radar echo prediction image;
[0017] S8: Perform Z-R transformation on the radar echo prediction image to obtain a transformation result;
[0018] S9: Use the transformation result for rainfall prediction.
[0019] Optionally, in S2, before inputting the denoised map dataset into the SAC-LSTM model, the short-term rainfall prediction method based on the attention mechanism cascaded LSTM further includes:
[0020] Padding and amplifying the length and width dimensions of the denoised map dataset, and then normalizing it to the range of [0,1] to form input data;
[0021] Input the input data into the SAC-LSTM model to obtain a predicted radar echo map dataset at time t+1.
[0022] Optionally, in S2, the SAC-LSTM model includes a cascaded long short-term memory unit based on the attention mechanism and a generative adversarial network;
[0023] The cascaded long short-term memory unit based on the attention mechanism is used to obtain a predicted radar echo map dataset at time t+1 according to the denoised map dataset;
[0024] The generative adversarial network is connected to the output layer of the cascaded long short-term memory unit based on the attention mechanism, and the generative adversarial network includes a discriminator module;
[0025] The discriminator module is used to estimate whether the input frame is a real image or a generated image.
[0026] Optionally, the cascaded long short-term memory unit based on the attention mechanism includes a cascaded LSTM unit and a self-attention mechanism module, and the self-attention mechanism module includes a matrix multiplication layer, a softmax function layer, and a data fusion layer arranged in sequence.
[0027] Optionally, S2 includes:
[0028] Input the denoised image dataset into the cascaded LSTM unit to obtain the output result of the cascaded LSTM unit;
[0029] Split the output result of the cascaded LSTM unit into first weighted data and second weighted data according to the weight matrix;
[0030] Perform matrix multiplication on the first weighted data and the second weighted data to obtain a matrix multiplication result;
[0031] According to the matrix multiplication result, use the softmax function to obtain the position similarities of different channels;
[0032] Fuse the position similarities of different channels and the output result of the cascaded LSTM unit, and output the obtained data fusion result as the predicted radar echo image dataset at time t+1.
[0033] Optionally, the loss function L of the updated SAC-LSTM model P is:
[0034] L P =L MSE +βL LP +ηL GAN (P)
[0035] where L MSE represents the standard mean squared error loss function, β and η are influence factors that determine the respective loss correlations, L LP represents the perceptual learning loss function and D m ( ) represents the feature map of the m-th layer output by the discriminator module, v represents the input frame, L GAN (P) represents the adversarial loss function of the prediction period module and T represents the total number of time steps, represents the prediction frame, t represents the current time step, and D( ) represents the feature map output by the discriminator module.
[0036] The present invention has the following beneficial effects:
[0037] On the one hand, under the action of the cascaded LSTM sub-unit of the attention mechanism, the captured local and global spatial features can be fused together, and then the SAC-LSTM model can significantly retain the visual details of each frame in the high-intensity rainfall area; on the other hand, the updated loss function of the present invention can more accurately implement complex geometric constraints on the global image structure, achieving the effect of improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1Flow chart of the short-term rainfall prediction method based on attention mechanism and cascaded LSTM of the present invention;
[0039] Figure 2 Schematic diagram of the structure of the SAC-LSTM model;
[0040] Figure 3 Schematic diagram of the structure of the cascaded long short-term memory unit with attention mechanism. Detailed implementation manners
[0041] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0042] The present invention provides a short-term rainfall prediction method based on attention mechanism and cascaded LSTM. Referring to Figure 1 as shown, the short-term rainfall prediction method based on attention mechanism and cascaded LSTM includes:
[0043] S1: Clean and denoise the historical radar echo map data set at time t to obtain a denoised map data set;
[0044] First, input the historical radar echo map sequence (X 1 , X 2 ,..., X t ) at the previous t moments. Clean and denoise the input data to obtain a denoised map data set;
[0045] After that, the present invention also pads and amplifies the length and width dimensions of the denoised map data set, and then normalizes the data to the range of [0, 1]. The obtained input data is input into the SAC-LSTM (long short-term memory network based on attention mechanism).
[0046] S2: Input the denoised map data set into the SAC-LSTM model to obtain the predicted radar echo map data set at time t + 1;
[0047] The S2 includes:
[0048] Input the denoised map data set into the cascaded LSTM unit to obtain the output result of the cascaded LSTM unit;
[0049] Split the output result of the cascaded LSTM unit into first weighted data and second weighted data according to the weight matrix;
[0050] Perform matrix multiplication operation on the first weighted data and the second weighted data to obtain the matrix multiplication result;
[0051] According to the matrix multiplication result, using the softmax function, obtain the position similarity α (ranging from 0 to 1) of different channels;
[0052] Fuse the position similarity α of different channels and the output result of the cascaded LSTM unit, and use the obtained data fusion result as the predicted radar echo map dataset at time t+1 for output.
[0053] Reference Figure 2 As shown, the SAC-LSTM model provided by the present invention includes a cascaded long short-term memory unit based on an attention mechanism and a generative adversarial network;
[0054] The cascaded long short-term memory unit based on the attention mechanism is used to obtain the predicted radar echo map dataset at time t+1 according to the denoised map dataset;
[0055] The generative adversarial network is connected to the output layer of the cascaded long short-term memory unit based on the attention mechanism, and the generative adversarial network includes a discriminator module;
[0056] The discriminator module is used to estimate whether the input frame is a real image or a generated image.
[0057] Optionally, as shown in Figure 3 The cascaded long short-term memory unit based on the attention mechanism includes a cascaded LSTM subunit and a self-attention mechanism module, and the self-attention mechanism module includes a matrix multiplication layer, a softmax function layer, and a data fusion layer arranged in sequence.
[0058] As Figure 3 In the first dashed box from the left in t (The radar echo at the current time t), (The hidden state of the k-th layer at time t-1) and (The time memory block of the k-th layer at time t-1) perform a concat operation, and then pass through the sigmoid function to output a value in the range of [0,1]. The closer to 0 means forgetting, and the closer to 1 means retaining. The formula expression is:
[0059]
[0060] σ represents the sigmoid function, represents the time memory block of the k-th layer at time t-1, represents the hidden state of the k-th layer at time t-1, W 1 represents the convolution filter, X t represents the radar echo at the current time t.
[0061] Finally, multiply the output value i of the sigmoid t and the output value g of the tanh t The output value of the sigmoid will determine which important information part of the output value of the tanh is retained, and the cell state of the previous layer is multiplied element-wise with the forget gate vector f t When the product value is close to 0, it means that the new cell state is to be discarded, and then this value is added element-wise to the output value of the input gate to update the information of the neural network into the cell state, obtaining the updated cell state. The formula expression is:
[0062]
[0063] represents the k-th layer memory block at the current time step t, f represents the forget gate of represents the k-th layer memory block at time step t - 1, i t represents the input gate of t represents as the candidate memory unit
[0064] The updated cell state enters the output gate. The previous hidden state and the current input are passed into the sigmoid, and then the updated cell state is passed to the tanh function. Finally, the output of the tanh is multiplied by the output of the sigmoid to determine the information carried by the hidden state. Then, the hidden state is used as the output of the current cell, and the new cell state and the new hidden state are passed to the next time step.
[0065] Such as Figure 3 in the second dashed box from the left in (the k-th layer memory block at the current time step t), (the k - 1-th layer spatial memory block at time step t) and X t (the radar echo at the current time step t) are combined. After passing through the sigmoid function, f t ′ (referred to as the forget gate) and i t ′ (referred to as the input gate) are obtained. After passing through the tanh function, g t ′ is obtained. This step is similar to the function in the first dashed box from the left in the figure, so it is called the cascaded long short-term memory network,
[0066] The formula is expressed as:
[0067]
[0068] Such as Figure 3 in the third dashed box from the left in and Perform a dot product operation through the tanh function, and the final output state at this time depends on (temporal state) and (spatial state). Since C is the C at the previous moment and M is the M of the previous layer, C is related to the time dimension and M is related to the spatial dimension here. The formula expression is as follows:
[0069]
[0070] O t as (the k-th layer spatial memory block at time t) and the output gate of the dual memory block, W 1 ~W 5 represent convolutional filters.
[0071] The long short-term memory structure itself has a certain ability to capture long-distance dependencies. However, since the sequence model allows information to flow through gated units and selectively transmits information. But when there are many visual details in the image, due to information loss in each recursion, the ability to capture dependencies becomes lower and lower.
[0072] Therefore, as shown in Figure 3 the rightmost dashed box in, the present invention introduces the self-attention mechanism into the cascaded LSTM. The formula expression is as follows:
[0073] α = softmax(W 7 tanh(W 6 H'))
[0074] output = H·α
[0075] The weight matrix is labeled as W 6 ∈R B×C×(T×H×W) and W 7 ∈R B×C×C . The Softmax function is used to perform matrix multiplication operations. The value of α is calculated by the sofamax function (ranging from 0 to 1), representing the positional similarity of different channels of the hidden state H. The dimension of the output output is transformed into the same dimension as the hidden state H.
[0076] S3: Compare the MSE losses of the real radar echo map dataset at time t + 1 and the predicted radar echo map dataset at time t + 1 to obtain the first comparison result;
[0077] S4: Input the real radar echo map dataset at the (t + 1) - th moment and the predicted radar echo map dataset at the (t + 1) - th moment into the adversarial generation network respectively, perform convolution operations to obtain the first feature map dataset corresponding to the denoised map dataset and the second feature map dataset corresponding to the predicted radar echo map dataset at the (t + 1) - th moment;
[0078] S5: Compare the perceptual learning loss and the adversarial loss between the first feature map dataset and the second feature map dataset to obtain the second comparison result;
[0079] S6: Update the loss function of the SAC - LSTM model using the first comparison result and the second comparison result to obtain the updated SAC - LSTM model;
[0080] S7: Use the updated SAC - LSTM model to predict the current input radar echo map dataset to obtain a radar echo prediction image;
[0081] S8: Perform Z - R transformation on the radar echo prediction image to obtain a transformation result;
[0082] S9: Use the transformation result for rainfall prediction.
[0083] Optionally, the loss function L of the updated SAC - LSTM model P is:
[0084] L P =L MSE +βL LP +ηL GAN (P)
[0085] where L MSE represents the standard mean square error loss function, β and η are influencing factors that determine the respective loss correlations, L LP represents the perceptual learning loss function and D m () represents the m - th layer of the output of the discriminator module, v represents the input frame, L GAN (P) represents the adversarial loss function of the prediction period module and T represents the total number of time steps, represents the prediction frame, t represents the current time step, and D() represents the feature map output by the discriminator module.
[0086] In the present invention, the long short-term memory structure itself has a certain ability to capture long-distance dependencies. However, since the sequence model enables information to flow through the gating unit and selectively transmits information. However, when there are many visual details in the image vision part, due to the loss of information in each recursion, the ability to capture dependencies becomes lower and lower, and the sequence feature maps cannot be effectively fused to predict the visual details of each frame. Therefore, under the action of the cascaded LSTM sub-unit of the attention mechanism, the captured local and global spatial features can be fused together, and the SAC-LSTM model can significantly retain the visual details of each frame in the high-intensity rainfall area.
[0087] In addition, since the radar echo map contains complex visual details, the encoder encodes and decodes it into high-dimensional features and then the decoder encodes and decodes it back to the radar echo map. During this process, many visual detail features will be lost, resulting in the predicted image becoming blurred. An adversarial generative network is connected to the output layer of the SAC-LSTM. During training, the mean squared error loss (MSE Loss), adversarial loss, and learning perception loss are added. The visual details of each frame are retained in the high-intensity rainfall area (dBZ = 40), and complex geometric constraints can be more accurately implemented in the global image structure, achieving the effect of improving the prediction accuracy.
[0088] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A short-term rainfall prediction method based on an attention mechanism and cascaded LSTM, characterized in that, the short-term rainfall prediction method based on an attention mechanism and cascaded LSTM includes: S1: Clean and denoise the historical radar echo map dataset at time t to obtain a denoised map dataset; S2: Input the denoised map dataset into the SAC-LSTM model to obtain a predicted radar echo map dataset at time t+1; S3: Compare the MSE losses of the real radar echo map dataset at time t+1 and the predicted radar echo map dataset at time t+1 to obtain a first comparison result; S4: Input the real radar echo map dataset at time t+1 and the predicted radar echo map dataset at time t+1 into the adversarial generation network respectively, perform convolution operations, and obtain a first feature map dataset corresponding to the denoised map dataset and a second feature map dataset corresponding to the predicted radar echo map dataset at time t+1; S5: Compare the perceptual learning loss and the adversarial loss between the first feature map dataset and the second feature map dataset to obtain a second comparison result; S6: Use the first comparison result and the second comparison result to update the loss function of the SAC-LSTM model to obtain an updated SAC-LSTM model; S7: Use the updated SAC-LSTM model to predict the current input radar echo map dataset to obtain a radar echo prediction image; S8: Perform Z-R transformation on the radar echo prediction image to obtain a transformation result; S9: Use the transformation result to perform rainfall prediction.
2. The short-term rainfall prediction method based on an attention mechanism and cascaded LSTM according to claim 1, characterized in that, in S2, before inputting the denoised map dataset into the SAC-LSTM model, the short-term rainfall prediction method based on an attention mechanism and cascaded LSTM further includes: Padding and amplifying the length and width dimensions of the denoised map dataset, and then normalizing it to the range of [0,1] to form input data; Input the input data into the SAC-LSTM model to obtain a predicted radar echo map dataset at time t+1.
3. The short-term rainfall prediction method based on an attention mechanism and cascaded LSTM according to claim 1, characterized in that, in S2, the SAC-LSTM model includes a cascaded long short-term memory unit based on an attention mechanism and an adversarial generation network; The cascaded long short-term memory unit based on an attention mechanism is used to obtain a predicted radar echo map dataset at time t+1 according to the denoised map dataset; The adversarial generation network is connected to the output layer of the cascaded long short-term memory unit based on an attention mechanism, and the adversarial generation network includes a discriminator module; The discriminator module is used to estimate whether the input frame is a real image or a generated image.
4. The short-term rainfall prediction method based on an attention mechanism and cascaded LSTM according to claim 3, characterized in that, The cascaded long short-term memory unit based on the attention mechanism includes a cascaded LSTM unit and a self-attention mechanism module. The self-attention mechanism module includes a matrix multiplication layer, a softmax function layer, and a data fusion layer arranged in sequence.
5. The short-term rainfall prediction method based on the cascaded LSTM with attention mechanism according to claim 3, wherein, S2 includes: Inputting the denoised graph data set into the cascaded LSTM unit to obtain the output result of the cascaded LSTM unit; Splitting the output result of the cascaded LSTM unit into first weighted data and second weighted data according to a weight matrix; Performing a matrix multiplication operation on the first weighted data and the second weighted data to obtain a matrix multiplication result; According to the matrix multiplication result, using the softmax function to obtain the position similarities of different channels; Fusing the position similarities of different channels and the output result of the cascaded LSTM unit, and outputting the obtained data fusion result as the predicted radar echo graph data set at time t+1.
6. The short-term rainfall prediction method based on the cascaded LSTM with attention mechanism according to claim 1, wherein, The loss function L of the updated SAC-LSTM model P is as follows: L P = L MSE + βL LP + ηL GAN (P) Among them, L MSE represents the standard mean square error loss function, and β and η are influence factors that determine their respective loss correlations. L LP represents the perceptron learning loss function and D m () represents the feature map of the m-th layer at the output of the discriminator module, v represents the input frame, and L GAN (P) represents the adversarial loss function of the prediction period module and T represents the total number of time steps, represents the predicted frame, t represents the current time step, and D() represents the feature map output by the discriminator module.
Citation Information
Patent Citations
Radar quantitative rainfall estimation method based on deep learning
CN113791415A
KR20230023227A