A short-term precipitation prediction method based on multi-modal feature fusion
By adopting a multimodal feature fusion method with a weighted attention mechanism in short-term precipitation prediction, the problem of scale mismatch and semantic confusion when data fusion of different modalities is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510067445.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-01-16
AI Technical Summary
When using multimodal data for short-term precipitation prediction in the prior art, there are problems of scale mismatch and semantic confusion when data fusion is integrated in different modes, and it is impossible to effectively play the advantages of multimodal data.
A multimodal feature fusion method based on the weighted attention mechanism is adopted, data features are extracted through the feature extraction network module, and weighted fusion module is used for weighted splicing. Combined with the weighted attention module, the weighted attention module is used to calculate the weights of different modal features, and the final multimodal fusion feature is output.
It effectively avoids the problem of inability to fusion of features at different scales and confusing results of fusion of different semantic features, reduces information loss and information redundancy, and improves the accuracy of the multimodal feature fusion process.
Smart Images

Figure CN119862534B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of precipitation prediction and deep learning in the meteorological field, and specifically refers to a short-term precipitation prediction method based on multi-modal feature fusion. Background Art
[0002] Precipitation prediction plays a key role in meteorological research and is of crucial significance to multiple fields such as agriculture, water resource management, disaster prevention and mitigation. Accurate prediction of precipitation on a short time scale can significantly improve the ability of disaster prevention and mitigation and reduce the losses caused by flood disasters.
[0003] In recent years, with the rapid development of artificial intelligence technology, short-term precipitation forecasting based on artificial neural networks has become the mainstream method. The artificial neural network needs to analyze and learn relevant historical meteorological data (wind speed, precipitation, cloud cover, etc.) to complete end-to-end rainfall forecasting.
[0004] However, the current mainstream neural network models face a difficult problem. Precipitation data is diverse, including but not limited to radar precipitation maps, radar cloud maps, wind speed maps, station data, etc. Different modal data extracts data features with different scales and semantics through the neural network. According to the existing situation, the current models often use methods such as short links, skip connections, and splicing when using multi-modal data. However, the above methods often cause problems such as mismatched scales and semantic confusion in the fusion of different modal data, and cannot give full play to the advantages of multi-modal data. Summary of the Invention
[0005] Aiming at the problems of different scales and semantic confusion in the fusion of different modal data in the prior art, the present invention proposes a short-term precipitation prediction method based on multi-modal feature fusion.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A short-term precipitation prediction method based on multi-modal feature fusion, comprising the following steps:
[0008] Step 1: Use the feature extraction network module to extract data features. Taking the radar precipitation map and station precipitation data as examples, the former often uses the CNN feature extraction network, and the latter often uses the GCN feature extraction network. The output feature of the former is M∈R CxHxW , and the latter is N∈R CxWxH , where C is the number of image channels; H is the image length; W is the image width.
[0009] Step 2: The features of both are weighted and spliced through the weighted fusion module designed in this application to obtain the initial fused feature Z.
[0010] Step 3: The initial fused feature Z passes through the weight attention module to obtain the weight coefficient K of the original feature X 1 , the weight coefficient K of the original feature Y 2 , where K 1 +K 2 = 1. The output weighted feature is M K = K 1 M and N k = K 2 N.
[0011] Step 4: The original feature branch and the weight branch output the final multi-modal fused feature Z through the concat splicing method K = M k + N k .
[0012] Preferably, the calculation process of the weighted fusion module mentioned in step 2 is as follows:
[0013] Z = αM + βN
[0014] where α represents the proportion weight of feature X, and β represents the proportion weight of feature Y. The change of the weight is determined by the loss function. The calculation method of its loss function is as follows:
[0015]
[0016] where x i represents the radar precipitation map data, and x α represents the corresponding prediction result. y i represents the input station data, and y α represents the corresponding prediction result.
[0017] Preferably, the weight attention module described in step 3 uses a global average pooling layer, an X-axis max pooling layer, a Y-axis max pooling layer, a Relu activation layer, and a Sigmoid activation layer respectively. The global average pooling layer obtains global information, and the X-axis pooling layer and the Y-axis pooling layer obtain local feature information. After passing through the Relu activation layer, they are concatenated and then passed through the Sigmoid activation layer to obtain the corresponding weight values K 1 and K 2 , and finally the proportion weights of different modal features are calculated using the weight values. The definition of K is as follows:
[0018]
[0019] where Sig represents the Sigmoid activation layer, Avg represents the global average pooling layer, MaxX represents the X-axis max pooling layer, MaxY represents the Y-axis max pooling layer, represents the Relu activation layer.
[0020] Preferably, the MSE and RMSE are used as evaluation indicators in the testing and evaluation phase, and the calculation methods are as follows:
[0021]
[0022] The present invention has the following characteristics and beneficial effects:
[0023] Adopting the above technical solution, this method uses a multi-modal feature fusion mechanism based on a weighted attention mechanism. Using the fusion mechanism proposed in this application, it is possible to avoid the problems of inability to fuse features of different scales and chaotic fusion results of different semantic features during the fusion of multi-modal features. At the same time, during the initial feature fusion process, a weighted fusion method is used to reduce the information loss and information redundancy problems existing in the initial splicing, effectively improving the accuracy of the multi-modal feature fusion process. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 It is a basic model framework diagram of the multi-modal network in the embodiment of the present invention.
[0026] Figure 2 It is an overall flowchart of the multi-modal feature fusion mechanism based on the weighted attention mechanism of the present invention.
[0027] Figure 3 It is a flowchart of the weight attention module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0029] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0030] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0031] The present invention provides a short-term precipitation prediction method based on multi-modal feature fusion. When applying deep learning technology to the short-term precipitation prediction task, the weighted attention multi-modal fusion method can be used to process the features from different modal data (such as radar precipitation maps, station data), thereby improving the accuracy of short-term precipitation prediction.
[0032] The specific steps are as follows:
[0033] Step 1: Obtain historical multi-source precipitation data, where the multi-source precipitation data includes radar precipitation maps and station precipitation data. In this embodiment, the data used is the radar precipitation map of Zhejiang Province and the station data of Zhejiang Province.
[0034] Step 2: Construct a basic multi-modal network model framework composed of a CNN feature extraction network and a GCN feature extraction network. In this embodiment, the radar precipitation map outputs the original feature M ∈ R through the CNN feature extraction network CxHxW , and the station precipitation data outputs the original feature N ∈ R through the GCN feature extraction network CxWxH , where C is the number of image channels; H is the length of the image; W is the width of the image.
[0035] It should be noted that both the CNN feature extraction network and the GCN feature extraction network are conventional feature extraction networks. Therefore, in this embodiment, the methods for extracting features by the CNN feature extraction network and the GCN feature extraction network will not be specifically described.
[0036] Step 3: Construct a multi-modal fusion network mechanism composed of a weighted fusion module and a weight attention module.
[0037] Specifically, as Figure 2 shown, in the weighted fusion module, the feature M ∈ R CxHxW and the feature N ∈ R CxWxH are weighted and fused according to the loss function to obtain the initial fusion feature Z.
[0038] Z = αM + βN
[0039] where α represents the proportion weight of the feature M, β represents the proportion weight of the feature N, and the change of the weight is determined by the loss function.
[0040] Furthermore, the initial fusion feature Z passes through the weight attention module to obtain the weight coefficient K 1 of the original feature M, and the weight coefficient K 2 of the original feature N, where K 1 + K 2 = 1, and the output weight feature is M K = K 1 M and N k = K 2 N, and finally the final multi-modal fusion feature Z K = M k + N k .
[0041] In this embodiment, the weighted fusion module uses the dual-unit error loss function designed in this application, and the expression is as follows:
[0042]
[0043] where x i represents the radar precipitation map data, x α represents the corresponding prediction result, y i represents the input station data, and y α represents the corresponding prediction result.
[0044] Furthermore, as Figure 3 shown, the weight attention module includes a global average pooling layer, an X-axis maximum pooling layer, a Y-axis maximum pooling layer, a Relu activation layer, a Sigmoid activation layer, a concat splicing, and a multiplication splicing.
[0045] The input features first pass through the global average pooling layer to obtain global information, the X-axis pooling layer and the Y-axis pooling layer to obtain local feature information. After passing through the Relu activation layer, they are concatenated and then pass through the Sigmoid activation layer to obtain the corresponding weight value K 1 and K 2 , and finally, the proportion of different modality features is calculated using the weight values respectively.
[0046] Among them, the calculation method of the weight value K is as follows:
[0047]
[0048] Among them, Sig represents the Sigmoid activation layer, Avg represents the global average pooling layer, MaxX represents the X-axis max pooling, MaxY represents the Y-axis max pooling layer, represents the Relu activation layer.
[0049] It can be understood that since K 1 +K 2 =1, therefore, through the calculation formula of the weight value, one of the weight values K 1 or K 2 is calculated, and the other weight value can be obtained.
[0050] Step 4: Input the radar precipitation map into the network model architecture composed of Step 2 and Step 3 to obtain the corresponding result.
[0051] For the network model architecture composed of Step 2 and Step 3, in this embodiment, the experiments are all carried out under the Pytorch deep learning framework. The Pytorch version is 2.0.1, and an NVIDIA 4090 GPU is used for training. The num-workers parameter is set to 32. The weight decay is set to 0.0004. The learning rate is set to 0.1, and the learning rate is adjusted every 10 epochs according to the performance change of the model. If the model performance is not significantly optimized, it is adjusted; if it is normally optimized, it is not adjusted.
[0052] The following are the specific experiments added and the descriptions:
[0053] This example is compared with multiple advanced methods on the same dataset, and the results are shown in Table 1.
[0054] Table 1: Comparison with advanced methods on the same dataset
[0055]
[0056] In this embodiment, MSE and RMSE are used as evaluation metrics in the test and evaluation phases, and the calculation methods are as follows:
[0057]
[0058] Among them, MSE and RMSE are error metrics, and the lower they are, the smaller the model error. It can be seen that the results predicted by the multi-modal feature fusion mechanism based on the weighted attention mechanism of the present invention are more accurate and have smaller errors.
[0059] The above has described the embodiments of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principles and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.
Claims
1. A short-term precipitation prediction method based on multimodal feature fusion, characterized in that: The steps include: Step 1: Obtain historical multi-source precipitation data, and extract features through feature extraction networks respectively, so as to obtain multi-dimensional original features, wherein the multi-source precipitation data includes radar precipitation map and site precipitation data; the feature extraction network includes CNN feature extraction network and GCN feature extraction network, and the radar precipitation map outputs original features M∈R through CNN feature extraction network CxHxW , the site precipitation data is extracted through the GCN feature network to output the original features N∈R CxWxH , where C is the image channel; H is the image length; W is the image width; the weighted fusion method of the multi-dimensional original features is as follows: Z=αM+βN Among them, α represents the weight of feature M, β represents the weight of feature N, and the change of weight is determined by the loss function; Step 2: Perform weighted splicing through the weighted fusion module to obtain the initial fusion features; The weighted fusion module uses a dual-unit error loss function, which is expressed as follows: where x i Represents radar precipitation map data, x α Represents the corresponding prediction result, y i Indicates the input site data, y α Indicates the corresponding prediction results; Step 3: The initial fusion features are passed through the weighted attention module to obtain the weight coefficients of the multi-dimensional original features, and then the weight features of the multi-dimensional original features are obtained, and the multi-modal fusion features are obtained by splicing; The weighted attention module includes a global average pooling layer, an X-axis maximum pooling layer, a Y-axis maximum pooling layer, a Relu activation layer, a Sigmoid activation layer, a concat splicing, and a multiplication splicing; Step 4: Analyze the error between the input data and the predicted data, calculate the error index, analyze the model performance based on the input data and the predicted data, and calculate the performance index.
2. The short-term precipitation prediction method based on multimodal feature fusion according to claim 1 is characterized in that: The method for extracting weight features by the weight attention module is: The global average pooling layer obtains global information, the X-axis pooling layer and the Y-axis pooling layer obtain local feature information, and the concatenation passes through the Relu activation layer and then passes through the Sigmoid activation layer again to obtain the corresponding weight value K 1 With K 2 , and finally use the weight values to calculate the proportion of different modal features.
3. The short-term precipitation prediction method based on multimodal feature fusion according to claim 2 is characterized in that: The weight value is calculated as follows: Among them, Sig represents the Sigmoid activation layer, Avg represents the global average pooling layer, MaxX represents the maximum pooling on the X axis, and MaxY represents the maximum pooling layer on the Y axis. represents the Relu activation layer, and Z represents the initial fusion feature.
4. The short-term precipitation prediction method based on multimodal feature fusion according to claim 3 is characterized in that: In step 3, the initial fusion feature Z is passed through the weight attention module to obtain the weight coefficient K of the original feature M 1 , the weight coefficient K of the original feature N 2 , where K 1 +K 2 =1, the output weight feature is M K =K 1 M and N k =K 2 N, and finally output the final multimodal fusion feature Z through splicing K =M k +N k .
Citation Information
Patent Citations
Short-term rainfall prediction method of convolutional network based on small attention mechanism and application thereof
CN118606628A
Precipitation prediction system, precipitation prediction method, program, base station selection system, and base station selection method
WO2023162482A1