Meteorological feature prediction method based on Transform and CNN parallel structure

By adopting the Transformer and CNN parallel structure method in meteorological prediction, combining self-attention mechanism and convolutional neural network, the problem that traditional methods are difficult to capture the local and global characteristics of meteorological data is solved, and higher meteorological prediction accuracy and timeliness are achieved.

CN120123686APending Publication Date: 2025-06-10GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510195050.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional meteorological prediction methods are difficult to capture the local characteristics and global trends of meteorological data at the same time, resulting in insufficient prediction accuracy, especially when facing variable meteorological conditions and long-term prediction tasks.

Method used

The meteorological feature prediction method based on the parallel structure of Transformer and CNN is adopted to capture the global timing dependence through the self-attention mechanism, and local features are extracted in combination with convolutional neural networks to form a joint modeling of global and local features.

Benefits of technology

It improves the accuracy and timeliness of meteorological prediction, can better adapt to rapidly changing meteorological conditions, and improves the accuracy and stability of short-term meteorological prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123686A_ABST
    Figure CN120123686A_ABST
Patent Text Reader

Abstract

The invention provides a meteorological feature prediction method based on a Transform and CNN (Convolutional Neural Network) parallel structure. The method comprises the following steps: firstly, normalizing meteorological sequence data; then, the normalized matrix is subjected to filling and blocking; then, linear projection is carried out to map the image to a high-dimensional space, and position codes are added to the image; a multi-head self-attention mechanism of a Transform encoder is used to extract global time sequence features, a convolutional neural network is used to extract local time sequence features, and a ReLU activation function is used to enhance non-linear features after convolution extraction. Afterwards, features extracted by the self-attention mechanism and the convolutional network are fused, and a global time sequence relation is captured while local features are reserved; and finally, generating a future weather prediction result through the prediction head, and optimizing the model by using a mean square error (MSE) loss function. Experimental results show that the method significantly improves the meteorological prediction precision, and has high efficiency and low error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is a meteorological feature prediction method based on a parallel structure of Transformer and CNN. This method is mainly applied to the high-precision prediction of meteorological data, especially in aspects such as meteorological change trend analysis, short-term weather forecasting, and meteorological anomaly detection. The present invention combines the self-attention mechanism in deep learning with the convolutional neural network, aiming to improve the accuracy and timeliness of meteorological prediction by jointly modeling the global temporal relationship and local temporal features of meteorological data. This method is widely applied in fields such as agricultural production, environmental monitoring, meteorological warning, and disaster prediction, and has important practical application value and economic significance. Background Art

[0002] Meteorological prediction relies on accurately extracting and modeling key features from complex time series data. Traditional meteorological prediction methods usually face the problem that they cannot capture both local features and global trends of meteorological data simultaneously. This results in insufficient prediction accuracy. Especially when facing variable meteorological conditions and long-term prediction tasks, the performance of traditional methods often fails to meet the accuracy requirements. With the development of deep learning technology, especially the wide application of the self-attention mechanism and convolutional neural network, new ideas and solutions have been provided for meteorological prediction.

[0003] This paper proposes a meteorological feature prediction method based on a parallel structure of Transformer and CNN. The self-attention mechanism can capture global temporal dependencies when processing time series data and is suitable for modeling long-term dependencies in meteorological data; while the convolutional neural network is good at extracting local features and effectively capturing short-term temporal patterns. Combining these two technologies can simultaneously improve the model's learning ability for local and global features, thereby improving prediction accuracy. In recent years, models combining the self-attention mechanism and convolutional neural network have become an important research direction in the field of meteorological prediction, providing strong support for efficient and accurate meteorological feature prediction. Summary of the Invention

[0004] The present invention provides a meteorological feature prediction method based on a parallel structure of Transformer and CNN, which has excellent performance in dealing with complex time dependencies and improving prediction accuracy. The technical solution is as follows:

[0005] 1. A meteorological feature prediction method based on a parallel structure of Transformer and CNN, characterized by comprising the following steps:

[0006] Step 1: Assume that the input meteorological data sequence contains L time steps, and each time step has M variables, represented as an L×M matrix Among them, M represents that there are M meteorological-related features at each time step, including various meteorological features such as temperature, humidity, and air pressure;

[0007] Step 2: Normalize the input data to obtain the normalized feature matrix x norm , and the specific formula is as follows, where μ is the mean of the input data x, and σ is the standard deviation of the input data x:

[0008]

[0009] Step 3: Expand the matrix x norm through the padding layer, and then decompose the matrix x with a stride of S norm to obtain N block-wise feature matrices x patch ;

[0010] Step 4: Perform linear projection on the block-wise feature matrix x patch and add positional encoding to map it to a specified high-dimensional space and add positional information. The specific formula is as follows:

[0011] x pos = W P ·x patch + W pos

[0012] where W P is the linear projection matrix, W pos is the positional encoding matrix, and x pos is the output after linear mapping and adding positional information;

[0013] Step 5: Pass the output matrix x of Step 4 pos to the parallel structure processing module CNN-Transformer composed of a Transformer encoder and a CNN. The processing process of this module is as follows:

[0014] Step 5.1: The Transformer encoder part in the CNN-Transformer module uses the multi-head attention mechanism to divide the input matrix x pos into multiple heads, perform linear transformation on each attention head to obtain the query matrix Q, key matrix K, and value matrix V, and then calculate the scaled dot-product attention. The specific formula is as follows:

[0016]

[0017] where, is the scaling factor, and softmax is used to normalize the attention scores;

[0018] Step 5.2: The CNN part in the CNN-Transformer module processes the input matrix xpos Perform a convolution operation with a convolution kernel of size k×k and a stride of s. The local features of the time series are extracted during the convolution process

[0019] The specific formula is as follows:

[0020] x conv = Conv1(x pos , W conv )

[0021] where W conv is the convolution kernel and x conv is the output matrix after convolution;

[0022] Step 5.3: Concatenate the outputs of multiple attention heads in the Transformer encoder part of the CNN-Transformer module, and then obtain the final attention output matrix x attn ;

[0023] Step 5.4: Use the ReLU activation function for the output matrix x conv in the CNN part of the CNN-Transformer module to increase non-linear features and obtain the activated convolution feature matrix x relu ;

[0024] Step 5.5: Add the output matrix x attn after passing through the Transformer encoder and the activated convolution output matrix x relu to obtain the complete output matrix, which not only retains the local spatial features extracted by the convolutional layer but also introduces the global temporal relationship captured by the self-attention mechanism;

[0025] Step 6: Map the obtained complete output matrix to the target output space through the prediction head to generate the prediction results for the next T time steps:

[0026]

[0027] Step 7: Use the mean squared error (MSE) loss function to calculate the loss through the following formula:

[0028]

[0029] where represents the predicted value of the i-th time series, represents the true value of the i-th time series, M represents the number of features in the meteorological data series, represents averaging the losses for each feature channel, summarizing the losses for M channels, averaging the summarized losses over the entire time series to obtain the overall objective loss, and evaluating the prediction results. Brief Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 It is a flowchart of a meteorological feature prediction method based on a parallel structure of Transformer and CNN provided by an embodiment of this patent.

[0032] Figure 2 It is a schematic diagram of the model of a meteorological feature prediction method based on a parallel structure of Transformer and CNN provided by an embodiment of this patent.

[0033] Figure 3 It is the experimental comparison result of the model C_Transformer of this patent and the current mainstream architecture of meteorological prediction. Detailed Embodiments

[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further elaborate on the embodiments of the present invention in conjunction with the drawings.

[0035] As Figure 2 shown, an embodiment of the present invention provides a meteorological feature prediction method based on a parallel structure of Transformer and CNN, including:

[0036] Step 1: Select the public dataset weather, which is meteorological data recorded every ten minutes in 2020. Each group of meteorological data contains 21 features, including temperature, humidity, wind speed, etc. Set the look-back window length of the time series to 336, then the input sequence can be represented as a 336×21 matrix Set the output stride to 336. In the subsequent steps, the meteorological data of the input 336 time steps is used to predict the meteorological conditions of the adjacent next 336 time steps. The 336 input data and the 336 predicted data are adjacent but non-overlapping, and each feature contains multiple blocks of data with 336 as the unit;

[0037] Step 2: Normalize the input data to obtain a normalized feature matrix

[0038] Step 3: Expand the matrix through a padding layer, and then decompose the matrix with a stride of 8 and a block length of 16 Obtain 41 feature matrices after partitioning

[0039] Step 4: For the feature matrix after partitioning Perform linear projection and add positional encoding, map it to a specified high-dimensional space and add position information. The specific formula is as follows:

[0040] x pos = W P ·x patch + W pos

[0041] where is the linear projection matrix, is the positional encoding matrix, is the output after linear mapping and adding position information;

[0042] Step 5: Pass the output matrix of Step 4 to the parallel structure processing module CNN-Transformer composed of a Transformer encoder and a CNN. The processing of this module is as follows:

[0043] Step 5.1: The Transformer encoder part in the CNN-Transformer module uses the multi-head attention mechanism to divide it into 32 attention heads, and the dimension of each head is Perform a linear transformation on each attention head to obtain the query matrix key matrix value matrix Then calculate the scaled dot-product attention. The specific formula is as follows:

[0045]

[0046] where, is the scaling factor, and softmax is used to normalize the attention scores;

[0047] Step 5.2: The CNN part in the CNN-Transformer module performs a convolution operation on the input matrix The convolution kernel size is 3×3, and the stride is 1. The convolution process extracts the local features of the time

[0049] series features. The specific formula is as follows:

[0050] x conv = Conv1(x pos ,W conv )

[0051] where, Wconv is the convolutional kernel, is the output matrix after convolution;

[0052] Step 5.3: Concatenate the outputs of multiple attention heads in the Transformer encoder part of the CNN-Transformer module, and then obtain the final attention output matrix through a linear transformation

[0054] Step 5.4: The output matrix of the CNN part in the CNN-Transformer module Use the ReLU activation function to increase non-linear features and obtain the activated convolutional features

[0056] matrix

[0057] Step 5.5: Add the output matrix after passing through the Transformer encoder and the activated convolutional output matrix to obtain the complete output matrix, which not only retains the local spatial features extracted by the convolutional layer but also introduces the global temporal relationship captured by the self-attention mechanism;

[0058] Step 6: Map the obtained complete output matrix to the target output space through the prediction head to generate the prediction results for the next 336 time steps:

[0059]

[0060] Step 7: Use the mean squared error (MSE) loss function to calculate the loss through the following formula:

[0061]

[0062] where, represents the predicted value of the i-th time series, represents the true value of the i-th time series, 21 represents that the meteorological data contains 21 features, represents taking the average of the losses in each feature channel, summarizing the losses in each channel, averaging the summarized losses over the entire time series to obtain the overall objective loss, and evaluating the prediction results. Figure 3Shows the experimental comparison results between the patent model C_Transformer and the current mainstream architecture of meteorological prediction. Among them, DataName represents the dataset, InputLength represents the input length, OutputLength represents the output length, Model_name represents the model name, mse is the mean square error, and mae is the mean absolute error. The results of mse and mae show that the patent model has improved the average performance of the prediction results of 21 meteorological features in the weather dataset, and the model has good prediction effects in aspects such as temperature, humidity, and air pressure. The model not only improves the prediction accuracy but also can maintain good stability and response speed in the environment of real-time meteorological data update, adapting to rapidly changing meteorological conditions, which is particularly important for short-term meteorological prediction in practical applications.

Claims

1. A meteorological feature prediction method based on the parallel structure of Transformer and CNN, characterized in that The following steps are involved: Step 1: Assume that the input meteorological data sequence contains L time steps, each time step has M variables, represented as an L×M matrix Where M means that there are M meteorological related features in each time step, including temperature, humidity, air pressure and other meteorological features; Step 2: Normalize the input data to obtain the normalized feature matrix x norm , the specific formula is as follows, where μ is the mean of the input data x and σ is the standard deviation of the input data x: Step 3: Transform the matrix x norm Expand through the padding layer and then decompose the matrix x with stride S norm Get the feature matrix x after N blocks patch ; Step 4: After the block feature matrix x patch Perform linear projection and add positional encoding, map it to the specified high-dimensional space and add position information. The specific formula is as follows: x pos =W P ·x patch +W pos Where W P is the linear projection matrix, W pos is the position encoding matrix, x pos This is the output after linear mapping and adding position information; Step 5: Convert the output matrix x of step 4 to pos The data is passed to the parallel structure processing module CNN-Transformer, which consists of the Transformer encoder and CNN. The processing process of this module is as follows: Step 5.1: The Transformer encoder part of the CNN-Transformer module uses a multi-head attention mechanism to convert the input matrix x pos Divide into multiple heads, perform linear transformation on each attention head to obtain the query matrix Q, key matrix K, value matrix V, and then calculate the scaled dot product attention. The specific formula is as follows: in, is a scaling factor, softmax is used to normalize the attention score; Step 5.2: The CNN part of the CNN-Transformer module processes the input matrix x pos Perform convolution operation, the size of the convolution kernel is k×k, the stride is s, and the convolution process extracts the local features of the time series features. The specific formula is as follows: x conv =Conv1(x pos ,W conv ) Among them, W conv is the convolution kernel, x conv is the output matrix after convolution; Step 5.3: Concatenate the outputs of multiple attention heads of the Transformer encoder part of the CNN-Transformer module, and then perform a linear transformation to obtain the final attention output matrix x attn ; Step 5.4: Transform the output matrix x of the CNN part of the CNN-Transformer module conv Use the ReLU activation function to increase the nonlinear features and obtain the activated convolution feature matrix x relu ; Step 5.5: Pass the output matrix x of the Transformer encoder attn And the convolution output matrix x after activation relu Add them together to get the complete output matrix, which not only retains the local spatial features extracted by the convolutional layer, but also introduces the global temporal relationship captured by the self-attention mechanism; Step 6: Map the obtained complete output matrix to the target output space through the prediction head to generate the prediction results for the next T time steps: Step 7: Calculate the loss using the mean squared error (MSE) loss function using the following formula: in, represents the predicted value of the i-th time series, represents the true value of the i-th time series, M represents the number of features of the meteorological data series, It means to average the loss of each feature channel, summarize the losses of M channels, average the summarized losses over the entire time series, obtain the overall target loss, and evaluate the prediction results.

Citation Information

Cited By

  • Distributed photovoltaic access area-oriented net load prediction method

    CN121546548A