Optimization of a short-term prediction model for typhoon precipitation and a prediction method
The short-profile prediction model of typhoon precipitation with Cross-Patch multi-layer semantic attention optimization encoder-decoder structure solves the problem of long-distance space-time dependence modeling, improves prediction accuracy and calculation efficiency, and is suitable for short-profile prediction of typhoon precipitation in large-scale weather systems.
Patent Information
- Application Number
- CN202211506522.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-11-28
AI Technical Summary
The existing convolutional neural networks and recurrent neural networks are difficult to effectively model long-distance space-time dependence relationships. The self-attention mechanism is computationally large in high-dimensional space-time and space domains and ignores high-level space-time feature correlations, resulting in insufficient prediction capabilities of short-term prediction models for typhoon precipitation.
The long-distance spatiotemporal modeling method of precipitation with Cross-Patch multi-layer semantic attention is adopted. Through multi-layer semantic full-time information interaction and three-head attention module, the model's perception and modeling ability of long-distance spatiotemporal dependence is improved, and the typhoon precipitation short-profile prediction model with encoder-decoder structure is optimized.
The prediction accuracy and generalization ability of the typhoon precipitation short-term prediction model are improved, and the prediction prediction of precipitation in different types of typhoon scenarios is effectively responded to, and the consumption of computing resources is reduced.
Smart Images

Figure CN115840261B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to an optimization method and prediction method for short-term typhoon precipitation prediction models. Background Art
[0002] There are significant spatio-temporal correlations between large-scale atmospheric circulation, water vapor transport, and other concurrent weather systems and typhoon precipitation. Characterizing the high-order spatio-temporal characteristics of macro weather systems is of great significance for constructing long-distance spatio-temporal dependencies of typhoon precipitation. However, the spatial encoding receptive field based on convolutional neural networks is relatively limited, and the temporal encoding based on recurrent neural networks also has the problem of long-term attenuation. Both are difficult to effectively model long-distance spatio-temporal dependency relationships. Some studies have carried out long-distance spatio-temporal modeling by combining self-attention mechanisms with position encoding. However, in the high-dimensional spatio-temporal domain, the pixel-to-pixel modeling method of the self-attention mechanism brings a huge amount of computation while increasing the fitting difficulty. On the other hand, self-attention only considers the spatio-temporal correlations between pixels during the calculation process and ignores the correlations between higher-level spatio-temporal characteristics. How to improve the model's perception and modeling ability of long-distance spatio-temporal dependencies at the cost of less computing power and further improve the prediction ability of the optimized model is the main goal of this chapter. Summary of the Invention
[0003] The purpose of the present invention is to overcome the existing deficiencies and provide an optimization method and prediction method for short-term typhoon precipitation prediction models.
[0004] To achieve the purpose of the present invention, the following technical solutions are provided:
[0005] In the first aspect, the present invention provides an optimization method for a short-term typhoon precipitation prediction model, which is used to optimize the short-term typhoon precipitation prediction model with an encoder-decoder structure. The steps of the optimization method are as follows:
[0006] S1: Obtain the original feature maps of the encoder and decoder in the short-term typhoon precipitation prediction model to be optimized, divide the original feature maps into regions, and obtain the representations of each region feature in the high-dimensional feature space, so that each original feature map forms a corresponding spatio-temporal semantic feature;
[0007] S2: Further perform multi-layer semantic full spatio-temporal information interaction based on the spatio-temporal semantic features obtained in S1, construct multi-layer semantic spatio-temporal features from spatio-temporal semantic features to a full spatio-temporal view, and provide multi-scale spatio-temporal semantic features with a full spatio-temporal receptive field for subsequent spatio-temporal attention calculation;
[0008] S3: Based on the multi-scale spatio-temporal semantic features with a multi-level semantic full spatio-temporal view, introduce an attention mechanism to construct a spatial, temporal, and channel triple attention module, conduct long-distance spatio-temporal attention calculation for precipitation, calculate the spatio-temporal attention in units of blocks, and finally, the result is decoded by blocks and fused with the original feature map through a residual structure to adaptively enhance the expression of predicted temporal and spatial features and improve the prediction accuracy of the short-term typhoon precipitation prediction model.
[0009] Based on the above technical solutions, each step is preferably implemented in the following specific ways. The preferred implementation methods of each step can be combined accordingly without conflict, which does not constitute a limitation.
[0010] As a preference of the above first aspect, in the step S1, the short-term typhoon precipitation prediction model includes an encoder and a decoder; the encoder includes three cascaded block modules, and each block module is a group of cascaded Convolution layers and ConvLSTM layers; the decoder also includes three cascaded block modules, and each block module is a group of cascaded ConvLSTM operations and Deconvolution operations; in S1, six temporal feature maps of the first block module in the encoder and six predicted temporal feature maps of the last block module in the decoder are extracted to calculate spatio-temporal semantic features, and the temporal feature maps extracted from the encoder are the temporal feature encodings output by the Convolution layer, which are used to construct the keys and values in the self-attention mechanism, and the predicted temporal feature maps extracted from the decoder are the predicted temporal features output by the ConvLSTM layer, which are used to construct the queries in the self-attention mechanism.
[0011] As a preference of the above first aspect, in the step S1, for the feature maps of the encoder and decoder in the original short-term typhoon precipitation prediction model respectively, the following processing is performed according to S11~S13:
[0012] S11: Divide the region for each original feature map with a size of c×h×w, regularly divide and crop the spatial region through a preset p×p patch size, and obtain h / p×w / p patch regions;
[0013] S12: Based on c′ convolutional kernels with a dimension of c×p×p, parallel encoding is performed on each block region, where c′>c, and in order to ensure that the patterns extracted from each block region are consistent, parameter sharing is performed on the encoding convolutional kernels of all block regions. After obtaining h / p×w / p block region feature maps with a dimension of c′×1×1, they are recombined into a spatio-temporal semantic feature with a dimension of c′×h / p×w / p;
[0014] S14: For the feature maps of the six input time series of the encoder and the feature maps of the six predicted time series of the decoder, spatio-temporal semantic features with dimensions of c'×h / p×w / p are respectively obtained through conversion according to S11 and S12.
[0015] As an optimization of the above first aspect, the specific method of step S2 is as follows:
[0016] S21: Concatenate and fuse the six spatio-temporal semantic features corresponding to the encoder and the six spatio-temporal semantic features corresponding to the decoder respectively to obtain two initial spatio-temporal features CPF in and CPF out with dimensions of (s in ×c', h / p, w / p) and (s out ×c', h / p, w / p);
[0017]
[0018]
[0019] In the formula: Concat represents the concatenation and fusion operation; s in and s out respectively represent the feature maps of the input time series in the encoder and the decoder, both of which are 6;
[0020] S22: Use multi-layer stacked depthwise separable convolution (DSC) to perform cross-patch cross-convolution on the initial spatio-temporal features CPF in and CPF out obtained in S21 in the spatial and feature depth directions respectively. The output feature of the i-th layer of stacked depthwise separable convolution is denoted as CPFi;
[0021] S23: After passing through a total of n layers of stacked depthwise separable convolution in S22, n + 1 feature maps CPF i are obtained, i ∈ [0, n], which respectively represent spatio-temporal features of different semantic levels from spatio-temporal semantic features to global receptive fields. Then, the spatio-temporal features of different semantic levels are weighted and fused by pixel-wise summation. Finally, the multi-scale spatio-temporal semantic feature TCPF is:
[0022]
[0023] Among them, λ0, λ1...λ n represent the weights of the n + 1 feature maps CPF i , which can be optimized through model backpropagation.
[0024] S24: Reverse split the multi-scale spatio-temporal semantic features in the concatenation manner adopted in S21 to obtain the multi-scale spatio-temporal semantic features TCPF at each moment with a full spatio-temporal view and semantic information at each level after separation. i , where \(i\in[1,12]\) represents the moment. The first six moments \(i\in[1,6]\) correspond to the input sequence, and the last six moments \(i\in[7,12]\) correspond to the output sequence.
[0025] As a preference of the first aspect above, in S22, each layer of the stacked depthwise separable convolution includes two layers of networks: Depthwise convolution and Pointwise convolution. The calculation process of each layer of the stacked depthwise separable convolution is expressed as:
[0026] CPF i = DSC(CPF i-1 ) = PW(DS(CPF i-1 ))
[0027] where CPFi represents the output feature of the \(i\)-th layer of the stacked depthwise separable convolution. The PW convolution is a weighted combination operation in the depth direction of the feature map, and the number of its convolution kernels is equal to the number of output channels.
[0028] As a preference of the first aspect above, the specific method of step S3 is as follows:
[0029] S31: First, use two 1*1 convolutions Conv k and Conv v to perform convolutions on the multi-scale spatio-temporal semantic features TCPF i at six input moments respectively to obtain the key K and value V features of the attention mechanism. And the convolution operations Conv k used at different moments share parameters with each other, and the convolution operations Conv v used at different moments also share parameters:
[0030] K = Conv k (TCPF1, TCPF2…TCPF6)
[0031] V = Conv v (TCPF1, TCPF2…TCPF6)
[0032] Meanwhile, use the 1*1 convolution operation Conv q to perform convolution on the multi-scale spatio-temporal semantic features TCPF i of six output sequences to obtain the query Q feature of the attention mechanism. The convolution operations Conv q used at different moments share parameters with each other:
[0033] Q = Conv q(TCPF7, TCPF8... TCPF 12 )
[0034] The obtained K, Q, and V have the same dimensions, all being (T, C, H, W).
[0035] S32: To calculate the attention in three different feature spaces of space, time, and channels, three-dimensional conversions are respectively performed on K, Q, and V to obtain three groups of K, Q, and V with dimensions of (T*C, H*W), (C*H*W, T), and (T*H*W, C); the first group of K, Q, and V are Ks, Qs, and Vs, with dimensions of (T*C, H*W), and are used for spatial attention calculation; the second group of K, Q, and V are Kt, Qt, and Vt, with dimensions of (C*H*W, T), and are used for temporal attention calculation; the third group of K, Q, and V are Kc, Qc, and Vc, with dimensions of (T*H*W, C), and are used for channel attention calculation;
[0036] After spatial attention, temporal attention, and channel attention calculations, attention vectors Att s , Att t and Att c ; are respectively obtained;
[0037] S33: The calculated spatial attention vector Att s , temporal attention vector Att t and channel attention vector Att c are uniformly restored to the (T, C, H, W) dimension through dimension transformation, and weighted fusion is performed in a pixel-by-pixel summation manner in combination with weight coefficients to obtain the final spatio-temporal attention calculation result Att fusion :
[0038]
[0039] Among them, Att re represents the Att vector after dimension transformation; the weight coefficients α, β, and γ can be optimized through model backpropagation;
[0040] The final spatio-temporal attention calculation result Att fusion After being decoded by a patch, the decoded result is fused into the feature map output by the decoder in the typhoon precipitation short-term prediction model through a residual structure to adaptively enhance the spatio-temporal feature expression of six prediction time series.
[0041] As a preference of the above first aspect, in S32, the internal calculation logics of the spatial, temporal, and channel attention modules are the same, and the formula is as follows:
[0042] Attention(Q, K, V) = softmax(Q T / K)V
[0043] Where: Attention represents the calculated attention vector.
[0044] As a preference of the first aspect above, the input of the typhoon precipitation short-term prediction model is precipitation grid data and typhoon feature data within the target area at 6 moments.
[0045] As a preference of the first aspect above, the typhoon features are the central pressure and maximum wind speed of the typhoon.
[0046] In the second aspect, the present invention provides a typhoon precipitation short-term prediction method, which is to optimize the typhoon precipitation short-term prediction model with an encoder-decoder structure by using the optimization method described in any one of the above first aspect, and after training, it is used for short-term prediction of typhoon precipitation.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] The present invention provides a method for optimizing a typhoon precipitation short-term prediction model. This method combines a precipitation long-distance spatio-temporal modeling method based on Cross-Patch multi-layer semantic attention, adopts a Cross-Patch multi-layer semantic spatio-temporal encoding strategy and spatio-temporal attention calculation based on multi-layer semantics. Through long-distance spatio-temporal modeling, it improves the model's understanding of large-scale high-order spatio-temporal features of complex weather systems. The method of the present invention addresses the problem of huge resource consumption brought by the pixel-based spatio-temporal attention modeling method in large-scale scenarios, and at the same time considers the influence of high-order spatio-temporal features representing the spatio-temporal evolution pattern of large-scale weather systems, realizing the improvement of the model's perception and modeling ability of long-distance spatio-temporal dependence, comprehensively and effectively improving the accuracy of precipitation intensity prediction. At the same time, the present invention can effectively handle the short-term prediction of precipitation in different types of typhoon scenarios and has a certain generalization ability. Therefore, the present invention has important practical application value for the model optimization and rapid application of typhoon precipitation short-term prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of long-distance spatio-temporal modeling based on cross-patch multi-layer semantic attention.
[0050] Figure 2 It is a flowchart of Patch Embedding.
[0051] Figure 3 It is a flowchart of Cross-Patch multi-layer semantic full spatio-temporal information interaction.
[0052] Figure 4 It is a flowchart of spatio-temporal attention calculation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The present invention will be further described and illustrated below in conjunction with the accompanying drawings and specific embodiments.
[0054] In a preferred embodiment of the present invention, an optimization method for a short-term typhoon precipitation prediction model is provided. This method is used to optimize the short-term typhoon precipitation prediction model with an encoder-decoder structure. The main steps of this optimization method for the short-term typhoon precipitation prediction model include three steps, namely S1 to S3:
[0055] S1: Obtain the original feature maps of the encoder and decoder in the short-term typhoon precipitation prediction model to be optimized, divide the original feature maps into regions, and obtain the representations of each region feature in the high-dimensional feature space, so that each original feature map forms a corresponding spatio-temporal semantic feature respectively. Compared with the original feature map, the feature map obtained by Patch Embedding reduces the spatial dimension and expands the channel dimension, and obtains the high-order expression of the original feature map in the high-dimensional feature space by compressing redundant fine-grained precipitation information;
[0056] S2: Based on the spatio-temporal semantic features PEF obtained by Patch Embedding in S1, further perform multi-layer semantic full spatio-temporal information interaction to construct multi-layer semantic spatio-temporal features from spatio-temporal semantic features to those with a full spatio-temporal view, providing multi-scale spatio-temporal semantic features TCPF with a full spatio-temporal receptive field for the subsequent calculation of spatio-temporal attention.
[0057] S3: Based on the multi-scale spatio-temporal semantic features TCPF with a multi-level semantic full spatio-temporal view, introduce an attention mechanism to construct a spatial, temporal, and channel three-head attention module, carry out long-range spatio-temporal attention calculation for precipitation, and calculate the spatio-temporal attention with blocks (patches) as units. The final result is decoded by patches and fused with the original feature map through a residual structure to adaptively enhance the spatio-temporal feature expressions of six prediction time series.
[0058] In the method shown in S1 to S3 above, aiming at the huge resource consumption problem brought by the spatio-temporal attention modeling method based on global pixel points in large-scale scenarios, and considering the influence of high-order spatio-temporal features representing the spatio-temporal evolution pattern of large-scale weather systems at the same time, a long-range spatio-temporal modeling mechanism based on Cross-Patch multi-layer semantic attention is designed, and its principle is as Figure 1As shown in the figure. On the one hand, in view of the huge computational amount based on global pixel point attention in high-dimensional spatio-temporal scenarios, the present invention proposes to calculate spatio-temporal attention based on patches. By dividing and encoding patches on the feature map, neighborhood information is aggregated to obtain high-order feature expressions based on patches. While compressing redundant fine-grained information in the feature map, the resource consumption of subsequent spatio-temporal attention calculation is significantly reduced. On the other hand, in order to further fully extract and fuse spatio-temporal features at different levels and provide multi-scale spatio-temporal features with a full spatio-temporal view for the calculation of spatio-temporal attention, the present invention proposes a Cross-Patch multi-layer semantic full spatio-temporal information interaction module, enabling subsequent spatio-temporal attention to adaptively select spatio-temporal features at different levels for long-range modeling during the calculation process.
[0059] The specific implementation manners and their effects of S1 to S3 in this embodiment will be described in detail below.
[0060] It should be noted that theoretically, the typhoon precipitation short-term prediction model optimized in the present invention can be any deep neural network model adopting an encoder-decoder structure, as long as the model can predict the short-term precipitation during the occurrence of a typhoon.
[0061] In the embodiment of the present invention, the above typhoon precipitation short-term prediction model includes an encoder (performing the Encoding process) and a decoder (performing the Forecasting process). As Figure 1 shown, the encoder includes three cascaded block modules, and each block module is a group of cascaded Convolution layers and ConvLSTM layers; the decoder also includes three cascaded block modules, and each block module is a group of cascaded ConvLSTM operations and Deconvolution operations. After the decoder, the final output is obtained through double-layer convolution to get the short-term precipitation prediction result.
[0062] In the above short-term typhoon precipitation prediction model, the input of the model needs to be time-series data related to typhoon precipitation. In this embodiment, the precipitation grid data and typhoon feature data in the target area at 6 moments can be selected. Among them, the specific typhoon features in the typhoon feature data can be optimized according to the actual situation. The typhoon features selected in this embodiment are the central pressure and maximum wind speed of the typhoon. Both the precipitation grid data and typhoon feature data input into the model need to be resampled to the same spatial reference and spatial resolution. This multi-source input needs to be fused. The simplest and most intuitive fusion method is to splice (Concatenation) the three grid data and use it as the three-channel feature to input the encoder. In addition, two encoding branches can also be set in the encoder respectively. One encoding branch is used to input the precipitation grid data, and the other encoding branch is used to input the typhoon feature data (splicing the central pressure and maximum wind speed). The outputs of the two encoding branches are fused through a splicing operation, and the fusion result is used as the input of the decoder. In this embodiment, based on the typhoon precipitation short-term prediction model composed of the above 6 block modules, six temporal-spatial semantic features can be calculated by extracting the six temporal feature maps of the first block module in the encoder and the six predicted temporal feature maps of the last block module in the decoder. The temporal feature maps extracted in the encoder are the temporal feature encodings output by the Convolution layer, which are used to construct the key Key (K) and value Value (V) in the self-attention mechanism. The predicted temporal feature maps extracted in the decoder are the predicted temporal features output by the ConvLSTM layer, which are used to construct the query Query (Q) in the self-attention mechanism.
[0063] In the embodiment of the present invention, in the above step S1, for the respective feature maps of the encoder and decoder in the original short-term typhoon precipitation prediction model, through Figure 2 the Patch Embedding method shown, the processing is carried out according to S11 to S13:
[0064] S11: Perform regional division on each feature map with the original size of c×h×w. Regularly divide and crop the spatial area through a preset p×p patch size to obtain h / p×w / p patch areas;
[0065] S12: Based on c′ convolutional kernels with a dimension of c×p×p, perform parallel encoding on each block area, where c′>c. And in order to ensure that the patterns extracted from each block area are consistent, parameter sharing is performed on the encoding convolutional kernels of all block areas. After obtaining h / p×w / p block area feature maps with a dimension of c′×1×1, they are recombined into a temporal-spatial semantic feature with a dimension of c′×h / p×w / p;
[0066] S14: Compared with the original feature map, the feature map obtained by the above Patch Embedding reduces the spatial dimension and expands the channel dimension, and obtains the high-order expression of the original feature map in the high-dimensional feature space by compressing redundant fine-grained precipitation information. Based on this process, for the feature maps of the six input time series of the encoder and the feature maps of the six predicted time series of the decoder, the spatio-temporal semantic features with the dimension of c'×h / p×w / p are respectively obtained through the conversion in accordance with S11 and S12.
[0067] Compared with the original feature map, the feature map obtained by Patch Embedding reduces the spatial dimension and expands the channel dimension, and obtains the high-order expression of the original feature map in the high-dimensional feature space by compressing redundant fine-grained precipitation information. Feature map compression is achieved by performing patch division on the feature map and encoding the regional representations of each patch at each moment, significantly reducing the resource consumption in the subsequent attention calculation process. At the same time, patch encoding aggregates local information, blurs low-order detailed features, and obtains high-order feature expressions based on patches, enabling spatio-temporal attention calculation to be carried out based on higher-level spatio-temporal features, which helps the model capture higher-order global spatio-temporal evolution patterns. To ensure that the encoding patterns of the input sequence and the output sequence are consistent, parameter sharing is maintained during the Patch Embedding process. Compared with the original feature map, the feature map obtained by Patch Embedding reduces the spatial dimension and expands the channel dimension, and obtains the high-order expression of the original feature map in the high-dimensional feature space by compressing redundant fine-grained precipitation information. Based on this process, the feature maps of the six input time series and the feature maps of the six predicted time series can be respectively converted into six patch-based features PEF with the dimension of c'×h / p×w / p, providing refined and efficient spatio-temporal semantic features for subsequent full spatio-temporal information interaction and spatio-temporal attention calculation.
[0068] In the embodiment of the present invention, in the above step S2, a Cross-Patch multi-layer semantic full spatio-temporal information interaction process is designed, as Figure 3 shown, and the specific method is as follows:
[0069] S21: Concatenate and fuse the six spatio-temporal semantic features corresponding to the encoder and the six spatio-temporal semantic features corresponding to the decoder along the time dimension direction respectively, to obtain two spatio-temporal features CPF in with the dimensions of (s out ×c', h / p, w / p) and (s in ×c', h / p, w / p), and the initial values of the two spatio-temporal features CPF out are represented by subscript 0, that is;
[0070]
[0071]
[0072] In the formula: Concat represents the concatenation fusion operation; s in and s out respectively represent the feature maps of the input time series in the encoder and decoder, both of which are 6 in this embodiment.
[0073] S22: Use multi-layer stacked depthwise separable convolution (DSC) to perform cross-patch cross-convolution on the initial spatio-temporal feature CPF obtained in S21 in and CPF out in the spatial and feature depth directions respectively. The output feature of the i-th layer of stacked depthwise separable convolution DSC is denoted as CPFi.
[0074] Among them, each layer of stacked depthwise separable convolution includes two layers of networks: Depthwise (DW) convolution and Pointwise (PW) convolution. The calculation process of each layer of stacked depthwise separable convolution DSC is expressed as:
[0075] CPF i = DSC(CPF i-1 ) = PW(DW(CPF i-1 ))
[0076] Among them, CPFi represents the output feature of the i-th layer of stacked depthwise separable convolution DSC. The PW convolution is a weighted combination operation in the feature map depth direction, and the number of its convolution kernels is equal to the number of output channels.
[0077] S23: After passing through a total of n layers of stacked depthwise separable convolution in S22, n + 1 feature maps CPF i , i ∈ [0, n], respectively represent different semantic-level spatio-temporal features from spatio-temporal semantic features (patch embedding) to those with a global receptive field. Then, the spatio-temporal features of different semantic levels are weighted and fused by the method of pixel-by-pixel summation. Let the weights of the n + 1 features be λ0, λ1... λ n , then the final multi-scale spatio-temporal semantic feature TCPF is:
[0078]
[0079] Among them, λ0, λ1... λ n represent the weights of the n + 1 feature maps CPF i , and can be optimized through model backpropagation.
[0080] S24: Split the multi-scale spatio-temporal semantic features in the reverse concatenation manner adopted in S21 to obtain the multi-scale spatio-temporal semantic features TCPF at each moment with full spatio-temporal vision and semantic information at each level i , where \(i\in[1,12]\) represents the moment. The first six moments \(i\in[1,6]\) correspond to the input sequence, and the last six moments \(i\in[7,12]\) correspond to the output sequence.
[0081] In the embodiment of the present invention, in the above step S3, a spatio-temporal attention calculation method is introduced, as Figure 4 shown. The specific method is as follows:
[0082] S31: First, use two 1×1 convolutions Conv k and Conv v to perform convolutions on the multi-scale spatio-temporal semantic features TCPF i of the six input moments respectively to obtain the key Key (K) and value Value (V) features of the attention mechanism. And the same convolution operation shares parameters between different moments, that is, the convolution operations Conv k used at different moments share parameters, and the convolution operations Conv v used at different moments also share parameters. The calculation formulas for the key Key (K) and value Value (V) features are as follows:
[0083] K = Conv k (TCPF1, TCPF2…TCPF6)
[0084] V = Conv v (TCPF1, TCPF2…TCPF6)
[0085] At the same time, use the 1×1 convolution operation Conv q to perform convolution on the multi-scale spatio-temporal semantic features TCPF i of the six output sequences to obtain the query Query (Q) feature of the attention mechanism. The convolution operations Conv q used at different moments share parameters with each other. The calculation formula for the Query (Q) feature is as follows:
[0086] Q = Conv q (TCPF7 TCPF8…TCPF 12 )
[0087] The obtained K, Q, and V have the same dimension, all being (T, C, H, W).
[0088] S32: To calculate the attention in three different feature spaces of space, time, and channels, three-dimensional transformations are respectively performed on K, Q, and V to obtain three groups of K, Q, and V with dimensions (T*C, H*W), (C*H*W, T), and (T*H*W, C); where the first group of K, Q, and V are Ks, Qs, and Vs, with dimensions (T*C, H*W), and are used for spatial attention calculation; the second group of K, Q, and V are Kt, Qt, and Vt, with dimensions (C*H*W, T), and are used for temporal attention calculation; the third group of K, Q, and V are Kc, Qc, and Vc, with dimensions (T*H*W, C), and are used for channel attention calculation.
[0089] The internal calculation logics of the spatial, temporal, and channel attention modules are the same, and the formula is as follows:
[0090] Attention(Q, K, V) = softmax(Q T K)V
[0091] In the formula: Attention represents the calculated attention vector.
[0092] After spatial attention, temporal attention, and channel attention calculations, the obtained attention vectors Attention are respectively denoted as Att s 、Att t and Att c .
[0093] S33: The calculated spatial attention vector Att s 、temporal attention vector Att t and channel attention vector Att c are uniformly restored to the (T, C, H, W) dimension through dimension transformation, and are weighted and fused in a pixel-by-pixel summation manner by combining weight coefficients to obtain the final spatio-temporal attention calculation result Att fusion :
[0094]
[0095] Among them, Att re represents the Att vector after dimension transformation; the weight coefficients α, β, γ can be optimized through model backpropagation.
[0096] The final spatio-temporal attention calculation result Att fusion After block (patch) decoding, the decoding result is fused into the feature map output by the decoder in the typhoon precipitation short-term prediction model through a residual structure to adaptively enhance the spatio-temporal feature expression of six prediction time series. The fused feature is then input into the double-layer convolution behind the decoder in the model, and the corresponding typhoon precipitation short-term prediction result is output.
[0097] Therefore, the short-term typhoon precipitation prediction model with an encoder-decoder structure is optimized using the optimization methods shown in S1-S3 above, and can be used for short-term prediction of typhoon precipitation after model training.
[0098] Based on the optimization methods shown in the above embodiments S1-S3, the following applies them to specific examples to demonstrate their effects. The specific process is as described above and will not be elaborated here. The following mainly shows the specific parameter settings and implementation effects.
[0099] Embodiment
[0100] The following takes typhoons in the Northwest Pacific region and the precipitation of typhoon "Halong" in 2019 as specific embodiments to specifically describe the present invention. The specific steps are as follows:
[0101] 1) The typhoon data used comes from the Central Meteorological Observatory of the National Meteorological Center, including 85 typhoons in the Northwest Pacific region from 2017 to 2019 ( Figure 1 .2). The observation information of the typhoon data includes cyclone position, central pressure, moving wind speed, maximum wind speed, etc. The spatial resolution of the typhoon data is 0.1°, and the observation intervals are mainly 3h and 6h. As the typhoon intensity increases, the observation interval will be correspondingly shortened, and the minimum time interval is 1h. Based on the spatio-temporally matched data, a sliding window method is used to construct spatio-temporal prediction sequence samples, and on this basis, the typhoon precipitation data set is further screened and divided.
[0102] 2) Since the observation sequence is data for the past 3 hours (6 frames) and the prediction sequence is data for the next 3 hours (6 frames), the sliding window length is set to 6 hours (12 frames). To obtain as many samples as possible while avoiding the problem of data leakage, the sliding window moving step size is set to 2h (4 frames). Sampling is performed on the matched time series data based on the sliding window method, and a series of precipitation sequence samples with a sequence length of 6 hours are finally obtained.
[0103] 3) The precipitation sequence samples in the previous step are screened, and precipitation sequences without typhoon events are deleted. In addition, when the typhoon is in the initial formation or final extinction stage, the number of frames with typhoon marks in the observation sequence (the first 6 frames) in the precipitation sequence sample may be less than 6 frames. To ensure that the samples have enough typhoon marks for the later typhoon information fusion and spatio-temporal inference graph modeling process, this paper further deletes the sequence samples in which the number of frames with typhoon marks in the observation sequence (the first 6 frames) is less than 3, and finally 4581 sequence samples are obtained.
[0104] 4) Randomly divide the sequence samples obtained in the second step into training dataset, validation dataset and test dataset according to the ratio of 10:1:2. Name the divided typhoon precipitation dataset as TCPD (Tropical cyclone Precipitation Dataset), and conduct relevant experiments based on this dataset;
[0105] 5) According to the above steps, use the TCPD dataset for model training and experimental verification.
[0106] Name the reconstruction method of the present invention as ConvLSTM+MCPA. Compared with the original classic ConvLSTM method CSI, it has increased by 1.06% (m = 0.5), 1.83% (m = 2), 2.63% (m = 5), 3.17% (m = 10), 3.52% (m = 20), 3.77% (m = 30) respectively; HSS has increased by 0.79% (m = 0.5), 1.73% (m = 2), 3.02% (m = 5), 4.14% (m = 10), 5.17% (m = 20), 5.87% (m = 30) respectively; F1 has increased by 0.99% (m = 0.5), 1.9% (m = 2), 3.12% (m = 5), 4.19% (m = 10), 5.19% (m = 20), 5.88% (m = 30) respectively. The improvement range of accuracy shows the characteristic of increasing with the increase of precipitation intensity. This is because the MCPA model with multi-layer semantic full spatio-temporal information interaction module further fully extracts and fuses spatio-temporal features at different levels, provides multi-scale spatio-temporal features with a full spatio-temporal view for the calculation of spatio-temporal attention, and ensures that the model can adaptively select spatio-temporal feature information at a reasonable semantic level for long-distance modeling during the calculation of spatio-temporal attention, ultimately improving the prediction ability of the model. The experimental results also verify the effectiveness and superiority of the long-distance modeling method proposed in the present invention.
[0107] The above embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant art can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting the equivalent replacement or equivalent transformation method fall within the protection scope of the present invention.
Claims
1. An optimization method for a short-term prediction model of typhoon precipitation, characterized in that, Optimize the short-term typhoon precipitation prediction model of the encoder-decoder structure. The steps of the optimization method are as follows: S1: Obtain the original feature maps of the encoder and decoder in the short-term typhoon precipitation prediction model to be optimized, divide the original feature maps into regions, and obtain the representations of each region feature in the high-dimensional feature space, so that each original feature map forms corresponding spatio-temporal semantic features respectively; S2: Further perform multi-layer semantic full spatio-temporal information interaction based on the spatio-temporal semantic features obtained in S1, construct multi-layer semantic spatio-temporal features from spatio-temporal semantic features to full spatio-temporal vision, and provide multi-scale spatio-temporal semantic features with full spatio-temporal receptive fields for subsequent spatio-temporal attention calculation; S3: Based on the multi-scale spatio-temporal semantic features with multi-level semantic full spatio-temporal vision, introduce an attention mechanism to construct a spatial, temporal, and channel three-head attention module, carry out long-distance spatio-temporal attention calculation for precipitation, calculate the spatio-temporal attention in units of blocks, and the final result is decoded by blocks and fused with the original feature map through a residual structure to adaptively enhance the spatio-temporal feature expression of the prediction time series and improve the prediction accuracy of the short-term typhoon precipitation prediction model; The specific method of step S2 is as follows: S21: Concatenate and fuse the six spatio-temporal semantic features corresponding to the encoder and the six spatio-temporal semantic features corresponding to the decoder respectively to obtain two initial values of spatio-temporal features CPF in with dimensions of (s out × c′, h / p, w / p) and (s in × c′, h / p, w / p); out Where: Concat represents the concatenation fusion operation; s in and s out respectively represent the feature maps of the input time series in the encoder and the decoder, both of which are 6; S22: Use multi-layer stacked depthwise separable convolutions to perform cross-patch cross convolutions on the initial spatio-temporal features CPF obtained in S21 in and CPF out in the spatial and feature depth directions respectively. The output features of the i-th layer of stacked depthwise separable convolutions are denoted as CPFi; S23: After passing through the n-layer stacked depthwise separable convolutions in S22, n + 1 feature maps CPF are obtained. i , i ∈ [0, n], respectively represent different semantic-level spatio-temporal features from spatio-temporal semantic features to spatio-temporal features with global receptive fields. Then, the spatio-temporal features at different semantic levels are weighted and fused by pixel-wise summation. Finally, the multi-scale spatio-temporal semantic feature TCPF is: Among them, λ0, λ1...λ n represent the weights of n + 1 characteristic maps CPF i and can be optimized through model backpropagation; S24: Reverse split the multi-scale spatio-temporal semantic features in the concatenation manner adopted in S21 to obtain the multi-scale spatio-temporal semantic features TCPF of each moment with full spatio-temporal vision and semantic information at each level after separation. i , where \(i\in[1, 12]\) represents the moment, the first six moments \(i\in[1, 6]\) correspond to the input sequence, and the last six moments \(i\in[7, 12]\) correspond to the output sequence.
2. The method for optimizing the short-term typhoon precipitation prediction model according to claim 1, characterized in that: In step S1, the short-term typhoon precipitation prediction model includes an encoder and a decoder; the encoder includes three cascaded block modules, and each block module is a group of cascaded Convolution layers and ConvLSTM layers; the decoder also includes three cascaded block modules, and each block module is a group of cascaded ConvLSTM operations and Deconvolution operations; in S1, six temporal feature maps of the first block module in the encoder and six predicted temporal feature maps of the last block module in the decoder are extracted to calculate spatio-temporal semantic features, and the temporal feature maps extracted from the encoder are the temporal feature encodings output by the Convolution layer, which are used to construct the keys and values in the self-attention mechanism, and the predicted temporal feature maps extracted from the decoder are the predicted temporal features output by the ConvLSTM layer, which are used to construct the queries in the self-attention mechanism.
3. The method for optimizing the short-term typhoon precipitation prediction model according to claim 1, wherein: In step S1, for the feature maps of the encoder and decoder in the original short-term typhoon precipitation prediction model respectively, they are processed according to S11~S13: S11: Divide the region of each original feature map with size c×h×w, regularly divide and crop the spatial region with a preset p×p patch size, and obtain h / p×w / p patch regions; S12: Parallelly encode each block region based on c' convolutional kernels with dimension c×p×p, where c'>c, and in order to ensure that the patterns extracted from each block region are consistent, the parameters of the encoding convolutional kernels for all block regions are shared. After obtaining h / p×w / p block region feature maps with dimension c'×1×1, they are recombined into spatio-temporal semantic features with dimension c'×h / p×w / p; S14: For the feature maps of the six input time series of the encoder and the feature maps of the six predicted time series of the decoder, spatiotemporal semantic features with a dimension of c'×h / p×w / p are obtained through conversion according to S11 and S12 respectively.
4. The method for optimizing the short-term typhoon precipitation prediction model according to claim 1, wherein: In S22, each layer of the stacked depthwise separable convolution includes two layers of networks, namely the Depthwise convolution and the Pointwise convolution. The calculation process of each layer of the stacked depthwise separable convolution is expressed as: CPF i = DSC(CPF i-1 ) = PW(DW(CPF i-1 )) Among them, CPFi represents the output feature of the i-th layer of the stacked depthwise separable convolution. The PW convolution is a weighted combination operation in the depth direction of the feature map, and the number of its convolution kernels is equal to the number of output channels.
5. The method for optimizing the short-term typhoon precipitation prediction model according to claim 1, characterized in that: The specific method of step S3 is as follows: S31: First, two 1*1 convolutions Conv k and Conv v are used to perform convolutions on the multi-scale spatio-temporal semantic features TCPFi at six input time instants respectively to obtain the key K and value V features of the attention mechanism, and the convolution operations Conv k used at different time instants share parameters, and the convolution operations Conv v used at different time instants also share parameters: K = Conv k (TCPF1, TCPF2...TCPF6) V = Conv v (TCPF1, TCPF2...TCPF6) Meanwhile, 1*1 convolution operation Conv is adopted q to perform convolution on the multi-scale spatio-temporal semantic features TCPF of six output sequences i to obtain the query Q features of the attention mechanism. The convolution operations Conv used at different times q share the same parameters with each other: Q = Conv q (TCPF7, TCPF8... TCPF 12 ) The obtained K, Q, and V have the same dimension, all being (T, C, H, W); S32: In order to calculate the attention in three different feature spaces of space, time, and channel, K, Q, and V are respectively subjected to three-dimensional conversions to obtain three groups of K, Q, and V with dimensions of (T*C, H*W), (C*H*W, T), and (T*H*W, C); the first group of K, Q, and V are Ks, Qs, and Vs, with a dimension of (T*C, H*W), and are used for spatial attention calculation; the second group of K, Q, and V are Kt, Qt, and Vt, with a dimension of (C*H*W, T), and are used for temporal attention calculation; the third group of K, Q, and V are Kc, Qc, and Vc, with a dimension of (T*H*W, C), and are used for channel attention calculation; After spatial attention, temporal attention, and channel attention calculations, the attention vectors Att s , Att t and Att c ; S33: The calculated spatial attention vector Att s , the temporal attention vector Att t and the channel attention vector Att c are uniformly restored to the (T, C, H, W) dimension through dimensionality transformation, and weighted fusion is performed in a pixel-by-pixel summation manner by combining the weight coefficients to obtain the final spatio-temporal attention calculation result Att fusion : Among them, Att re represents the Att vector after dimensionality transformation; the weight coefficients α, β, and γ can be optimized through backpropagation of the model; Final spatio-temporal attention calculation result Att fusion After block decoding, the decoding result is fused into the feature map output by the decoder in the short-term typhoon precipitation prediction model through a residual structure to adaptively enhance the spatio-temporal feature expression of six prediction time series.
6. The method for optimizing the short-term prediction model of typhoon precipitation according to claim 1, wherein: In S32, the internal calculation logics of the spatial, temporal, and channel attention modules are the same, and the formula is as follows: Attention(Q,K,V) = softmax(Q T K)V In the formula: Attention represents the calculated attention vector.
7. The method for optimizing the short-term typhoon precipitation prediction model according to claim 2, characterized in that, The input of the short-term typhoon precipitation prediction model is the precipitation grid data and typhoon feature data within the target area at 6 moments.
8. The method for optimizing the short-term typhoon precipitation prediction model according to claim 2, characterized in that, The typhoon features are the central pressure and maximum wind speed of the typhoon.
9. A short-term prediction method for typhoon precipitation, characterized in that, The short-term typhoon precipitation prediction model with an encoder-decoder structure is optimized by using the optimization method described in any one of claims 1 to 8, and is used for short-term prediction of typhoon precipitation after training.