Transformer oil temperature prediction method and system
Through the improved Transformer model, a multi-scale feature tree is constructed using the CGEM module and the Skip-PAM module, which solves the integration of medium- and long-term dependence and short-term dynamic changes in transformer oil temperature prediction, and achieves higher prediction accuracy and stability.
Patent Information
- Application Number
- CN202510165659.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to effectively capture long-term dependence and short-term dynamic changes in transformer oil temperature prediction, and it is difficult to integrate short-term and long-term characteristics without sacrificing computing efficiency, resulting in insufficient prediction accuracy.
Using the improved Transformer model, multi-scale time features are extracted through the CGEM module, and a Skip-PAM module with a step-length pyramid-type attention mechanism is introduced for feature fusion, a multi-scale feature tree is constructed, and a full-connection layer is used for prediction.
It significantly improves the accuracy of transformer oil temperature prediction, can keenly capture short-term fluctuations and long-term trends on different time scales, and enhances the generalization ability and prediction accuracy of the model.
Smart Images

Figure CN120277601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transformer detection, and particularly to a transformer oil temperature prediction method and system. Background Art
[0002] The dynamic load capacity and insulation aging rate of power transformers mainly depend on their thermal characteristics. The top oil temperature is an important indicator for measuring the thermal characteristics of transformers and is also one of the important monitoring quantities during the operation of transformers. Accurately and reliably predicting the top oil temperature is of great significance for reasonably guiding the arrangement of the dynamic load of transformers and preventing transformer thermal faults.
[0003] Time series prediction has wide applications in production and life. In recent years, time series prediction methods based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs) can well capture non-linear relationships and handle long-term dependencies. The data self-adaptability of these models and their processing capabilities on large-scale data also make up for the deficiencies of traditional methods. However, when dealing with time series prediction problems, the following problems exist:
[0004] 1. Challenges in capturing long-term dependencies: A key feature of time series data is its inherent temporal dependence, and capturing long-term dependence relationships is particularly important for accurate prediction. Although RNNs and their variants such as long short-term memory networks (LSTMs) are designed to handle sequence dependence problems, they are often limited by gradient vanishing or explosion in practical applications and are difficult to effectively learn dependence relationships over long time spans. Although Transformers improve the ability to capture long-term dependence relationships through self-attention mechanisms, they have high computational complexity and large resource consumption when dealing with very long sequences.
[0005] 2. Insufficient attention to short-term dynamics: In time series prediction, short-term dynamic changes are also very crucial. For example, in financial market analysis, short-term price fluctuations may have an important impact on future trends. CNNs can capture local features, but their fixed window size limits their adaptability to features at different time scales. Although RNNs and Transformers can handle sequence data, they may not be as flexible and effective as specially designed mechanisms when emphasizing features that are closely related in the short term.
[0006] 3. Challenges in integrating short-term and long-term features: In long time series analysis, there is an urgent need to capture short-term features that affect immediate decision-making, as well as to understand the factors driving long-term trends. Existing technologies face significant challenges in simultaneously processing these two types of features. For example, although RNNs and their variants can capture the dynamic characteristics of time series through recursive processing of sequences, they often focus more on recent information in practice and have difficulty balancing and integrating the influence of long-term historical data. Similarly, Transformers can capture dependencies within the entire sequence through their self-attention mechanism, but their equal attention to all time points often makes the model insufficient in distinguishing which are the key short-term signals and long-term trends. In addition, existing models are difficult to simultaneously capture and utilize these two types of features without sacrificing computational efficiency. In practical applications, this leads to for complex time series prediction tasks, the model either tends to ignore the subtle patterns crucial for predicting short-term behavior or fails to fully utilize the trends and periodic information contained in long-term historical data.
[0007] Therefore, an improved time series prediction algorithm is needed to accurately predict the transformer oil temperature. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a transformer oil temperature prediction method and system that can accurately capture and utilize key information in time series data and improve the accuracy of transformer oil temperature prediction.
[0009] The technical solution adopted by the present invention to solve its technical problem is: to provide a transformer oil temperature prediction method, including the following steps:
[0010] Collect the transformer oil temperature to obtain target time series data;
[0011] Use a time series prediction model to analyze the target time series data to obtain a prediction result; the time series prediction model is constructed based on an improved Transformer and includes:
[0012] An encoding part, embedding a CGEM module at the encoder input end to extract time features with different granularities at multiple different scales of the target time series data, and introducing a stride pyramid attention mechanism through a Skip-PAM module to perform feature fusion on the time features;
[0013] A decoding part, analyzing and obtaining a prediction result according to the fused time features.
[0014] Further, the CGEM module includes:
[0015] A first linear layer, converting the initial feature dimension of the input time series data into a set feature dimension;
[0016] A number of feature convolutional layers extract feature vectors of different scales from the input time series data in sequence and splice the feature vectors;
[0017] A second linear layer outputs after restoring the feature dimension of the spliced feature vectors to the initial feature dimension.
[0018] Further, after the feature convolutional layer extracts the feature vectors of the input time series data through a convolutional kernel with a stride of step and splices them, it then performs a convolution operation on the spliced feature vectors through a convolutional kernel with a stride of where 1 is the length of the input time series data.
[0019] Further, the strides of the number of feature convolutional layers are not all the same.
[0020] Further, the stride pyramid attention mechanism includes:
[0021] Obtain a time feature tree based on the multi-scale time features;
[0022] For any node in the time feature tree, calculate the attention value of the current node according to the relationship between the current node and its adjacent nodes of the same scale, the relationship between the current node and its parent node, and the relationship between the current node and its child node.
[0023] Further, the attention value is calculated by the following formula:
[0024]
[0025] where, is the total number of nodes for which the attention mechanism should be calculated. For node represents the s-th layer scale from the bottom to the top, l ∈ [1, l s represents the first node of this layer, q i is the query matrix corresponding to node x i k l and v l are the key matrix and value matrix respectively, d x is the sequence length, is the attention value.
[0026] Further, the total number of nodes for which the attention mechanism should be calculated is calculated by the following formula:
[0027]
[0028] where, for node represents its adjacent sibling nodes, and A represents the number of its adjacent sibling nodes. represents its child nodes. represents its parent node.
[0029] Furthermore, the prediction result obtained by analyzing the fused time features is achieved through a fully connected layer.
[0030] The present invention also provides a transformer oil temperature prediction system, including:
[0031] An input module that collects the transformer oil temperature to obtain target time series data;
[0032] An analysis module that analyzes the target time series data using a time series prediction model to obtain a prediction result;
[0033] The time series prediction model is constructed based on an improved Transformer and includes:
[0034] An encoding part that embeds a CGEM module at the input end of the encoder to extract different granularity time features of the target time series data at multiple different scales, and introduces a stride pyramid attention mechanism through a Skip-PAM module to perform feature fusion on the multi-scale time features;
[0035] A decoding part that analyzes the fused time features to obtain a prediction result.
[0036] Furthermore, the CGEM module includes:
[0037] A first linear layer that converts the initial feature dimension of the input time series data into a set feature dimension;
[0038] A number of feature convolutional layers that sequentially extract feature vectors of different scales of the input time series data and splice the feature vectors;
[0039] A second linear layer that restores the feature dimension of the spliced feature vectors to the initial feature dimension and then outputs.
[0040] Furthermore, the attention value output by the Skip-PAM module is calculated by the following formula:
[0041]
[0042] where is the sum of all nodes for which the attention mechanism should be calculated. For node represents the s-th layer scale from the bottom to the top, l ∈ [1, l s represents the first node of this layer, q i is the query matrix corresponding to node x i kl and v l are the key matrix and the value matrix respectively, and d x is the sequence length, is the attention value.
[0043] Furthermore, the sum of nodes for which the attention mechanism should be calculated is calculated by the following formula:
[0044]
[0045] where, for node represents its sibling adjacent node, A represents the number of its sibling adjacent nodes, represents its child node, represents its parent node.
[0046] Beneficial effects
[0047] Due to the above technical solutions, compared with the prior art, the present invention has the following advantages and positive effects: By improving the Transformer model, the present invention can more accurately capture and utilize the key information in time series data. Whether it is short-term fluctuations or long-term trends, the method of the present invention can comprehensively integrate various data characteristics, significantly improve the prediction accuracy, and enhance the generalization ability for different time series characteristics and change patterns. This means that the model can achieve stable and accurate predictions on a wide range of data sets without the need for a large amount of manual adjustment for specific data. Especially when encountering new or unseen time series data, the present invention can effectively adapt and maintain its prediction accuracy, thus providing reliable support and decision-making basis in various application scenarios. The present invention particularly emphasizes the comprehensive consideration and processing of short-term and long-term dependencies in time series. By innovatively integrating short-term dynamics and long-term trends through the Skip-PAM module, it can keenly capture the subtle changes affecting short-term decisions without sacrificing long-term prediction ability, thereby achieving excellent prediction performance on multiple time scales. Brief description of the drawings
[0048] Figure 1 is a flowchart of the first embodiment of the present invention;
[0049] Figure 2 is the overall framework diagram of the MSFformer of the first and second embodiments of the present invention;
[0050] Figure 3 is a schematic diagram of the feature convolutional layer FCNN of the first and second embodiments of the present invention;
[0051] Figure 4 is a schematic diagram of the CGEM module of the first and second embodiments of the present invention;
[0052] Figure 5 It is a schematic diagram of the Skip-PAM module in the first and second embodiments of the present invention;
[0053] Figure 6 It is a comparison chart of experimental results of the ETTh1 oil temperature characteristics in the first embodiment of the present invention;
[0054] Figure 7 It is a comparison chart of the true value and predicted value of the synthetic dataset in the first embodiment of the present invention. Specific Embodiments
[0055] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0056] The first embodiment of the present invention relates to a transformer oil temperature prediction method based on an improved Transformer, as Figure 1 shown, including the following steps:
[0057] Collect the transformer oil temperature to obtain target time series data;
[0058] Use a time series prediction model to analyze the target time series data to obtain a prediction result; wherein, the time series prediction model is constructed based on an improved Transformer and includes:
[0059] An encoding part, embedding a CGEM module at the encoder input end to extract the time features of different granularities of the target time series data at multiple different scales, and introducing a stride pyramid attention mechanism through a Skip-PAM module to perform feature fusion on the multi-scale and multi-granularity time features;
[0060] A decoding part, analyzing the fused multi-scale time features to obtain a prediction result.
[0061] The following further illustrates this embodiment in conjunction with a specific time series prediction model MSFformer (Multi-Scale Feature Transformer).
[0062] The overall framework of the MSFformer proposed in this embodiment is as Figure 2As shown in the figure, this model improves the attention mechanism in the encoder on the basis of Transformer and introduces a pyramid-shaped attention mechanism across time steps (Skip-PAM, Skip-Pyramidal Attention Module). To construct multi-scale time information, a feature convolution is designed to build the CGEM (Coarse-Grained Extraction Module) module to obtain multi-granularity time information.
[0063] To adapt to the feature extraction with a stride length, a feature convolution layer (FCNN, Feature CNN) is constructed as shown in Figure 3 the figure. Specifically, for time series data with a length of l, the FCNN first extracts feature vectors through a convolution kernel with a stride length of step and then concatenates them together. Then, a convolution operation is performed through a convolution kernel with a size of and a stride length of . The formula is as follows:
[0064] Output = FCNN(Input)
[0065] To better apply the Skip-PAM structure, a pyramid-shaped feature tree needs to be constructed using the CGEM module. The CGEM module is as shown in Figure 4 the figure. The input time series is first passed through a linear layer to expand the feature dimension to a fixed dimension, and then gradually passed through the feature convolution layer to obtain feature information at different scales. Then, they are concatenated to form a pyramid-shaped feature structure. Finally, the feature dimension is restored through a linear layer:
[0066] M0 = Linear1(X)
[0067] M1 = FCNN1(M0)
[0068] M2 = FCNN2(M1)
[0069] M3 = FCNN3(M2)
[0070] M CGEM = Linear2([M0; M1; M2; M3])
[0071] where X is the input of the CGEM module, Linear is the linear layer, FCNN is the feature convolution module mentioned in the above formula, M1, M2, and M3 are the results of successive feature convolutions, M CGEM is the output of the CGEM module, and the ';' operation concatenates M0, M1, M2, and M3 in the time dimension.
[0072] By separately setting the stride lengths of FCNN1, FCNN2, and FCNN3, time feature information with different granularities is obtained at different scales. By stacking feature convolutional layers to construct a pyramid-shaped feature tree, not only can the model understand data from multiple levels, but it also provides a good foundation for the implementation of the Skip-PAM attention mechanism.
[0073] As Figure 5 shown is the time feature tree constructed by Skip-PAM based on the CGEM module. The time series extracts information through the attention mechanism from multiple scales, which can be divided into intra-scale connections and inter-scale connections. Intra-scale connections mean that a node performs attention calculations with adjacent nodes in its own scale layer, and inter-scale connections mean that a node performs attention calculations with its parent node (each parent node has P children) and C child nodes. Specifically, for a node where s ∈ [1, S] represents the s-th scale from the bottom layer to the top layer, and l ∈ [1, l s represents the l-th node in this layer:
[0074]
[0075]
[0076] Among them represents adjacent nodes within the scale, A represents the number of adjacent nodes within the scale, represents child nodes, represents the parent node, is the sum of nodes for which the attention mechanism should be calculated. Then, according to the principle of the self-attention mechanism, for the node the attention can be expressed as:
[0077]
[0078] where q i is the query matrix corresponding to x i k l and v l are the key matrix and value matrix respectively, d x is the sequence length, is the attention value of the node x i .
[0079] Through such a pyramid-shaped attention mechanism, combined with the multi-scale feature convolution to be mentioned later, a powerful feature extraction network is formed, which can adapt to the dynamic changes of various time scales, whether it is short-term fluctuations or long-term evolution.
[0080] Other modules follow the Transformer model structure, where the decoder uses a fully connected layer to avoid the problem of error accumulation in the Transformer autoregressive decoder.
[0081] The following is a further illustration through the practical application and data of the above method.
[0082] (1) Dataset Preparation
[0083] In this embodiment, three datasets of time series with different stationarity are used to verify the model, specifically as follows:
[0084] ETTh1: The ETT dataset contains the transformer oil temperature records collected from July 2016 to July 2018 and six time series of indicators related to voltage load. The hourly sampling frequency of ETTh1 data is very suitable for evaluating the model's ability to capture the periodic daily variations and long-term seasonal trends in energy use related to ambient temperature fluctuations.
[0085] ETTm1: As a finer-grained subset of the ETT dataset, ETTm1 provides data at 15-minute intervals. This high-resolution dataset requires the model to identify more subtle short-term variations and sudden changes in power load, which is crucial for operational decisions in energy distribution and responses to rapid demand changes.
[0086] Electricity: This dataset records the hourly electricity consumption of 321 customers from 2012 to 2014. It provides a rich information source for studying consumer behavior over the years, including consumption changes caused by personal lifestyles, social events, and different business operation models.
[0087] All three datasets are divided into training set, validation set, and test set in a ratio of 6:2:2. The data is shown in Table 1 specifically.
[0088] Table 1 Datasets
[0089]
[0090]
[0091] (2) Evaluation Metrics
[0092] In this embodiment, two widely used performance evaluation metrics are adopted to measure the prediction accuracy of the model, namely the Mean-Square Error (MSE) and the Mean Absolute Error (MAE). The lower the values of these two metrics, the better the prediction effect of the model.
[0093] MSE: Mean Squared Error measures the predicted value y' i and the true value y i The average of the squared differences. This metric assigns a higher penalty to larger errors, thus paying more attention to larger prediction errors when evaluating the model. The formula can be expressed as:
[0094]
[0095] MAE: Similar to MSE, this metric calculates the average absolute difference between the predicted value y' i and the actual observed value y i It reflects the average degree to which the predicted value deviates from the true value. The formula can be expressed as:
[0096]
[0097] (3) Parameter settings
[0098] In this embodiment, four different prediction time point lengths are set for the dataset, namely {96, 192, 336, 720}. For the datasets ETTh1 and Electricity, it corresponds to predicting data points for the next 4 days, 8 days, 2 weeks, and 1 month. For the dataset ETTm1, it corresponds to predicting data points for the next 1 day, 2 days, 3.5 days, and 7.5 days. The number of encoder layers is set to 4, and the number of attention heads is set to 6. A time series of length 96 is used as the input, and the stride window for constructing the number of features is set to {4, 4, 6}, resulting in a tree layer structure of {96, 24, 6, 1}. During the construction of Skip - PAM, the scale - interconnection nodes are set to 5.
[0099] In addition, Adam is uniformly used as the optimization algorithm during neural network training, and ten rounds of iterative training are performed. The initial learning rate is set to 1e -4 , and it is decayed by a factor of five every two rounds of iteration.
[0100] (4) Comparative experiments
[0101] To verify the effectiveness of the model, the MSFformer model is compared with the following three baseline models in terms of training and results.
[0102] Transformer: A neural network model based on the multi - head attention mechanism, which has applications in various fields.
[0103] Informer: A variant neural network model based on Transformer that reduces the time complexity of Transformer and is more suitable for long - time series prediction.
[0104] Pyraformer: It is also a variant neural network model based on Transformer, which extracts temporal features by establishing a pyramid-shaped attention mechanism.
[0105] Table 2 Comparative experiments with baseline methods
[0106]
[0107] Table 2 describes the experimental data comparison of the MSFformer model proposed in this chapter and the three baseline methods on the ETTh1, ETTm1 and Electricity datasets, and the best results are in bold. During the training process, the input time series length is 96, and the prediction lengths are {96, 192, 336, 720} respectively. From Table 2, we can conclude that:
[0108] Among all datasets, especially ETTh1 and ETTm1 datasets, the MSFformer model has the best overall performance, followed by Pyraformer. This shows that MSFformer can effectively obtain temporal feature information from multiple scales in long time series prediction tasks, and can combine long-term and short-term features at the same time through the Skip-PAM module to improve the model's prediction ability.
[0109] For the Electricity dataset, MSFformer performs poorly in the evaluation index MSE, while MAE performs best among all model methods. This may be because MSFformer has large prediction errors at some time points, and the MSE evaluation index amplifies this error. However, the MAE data shows that the MSFformer model still performs well in time series prediction tasks.
[0110] Overall, the MSFformer model outperforms the three baseline methods. On the three datasets, the MAE indicators are improved by up to 26.95%, 31.03% and 35.87% respectively compared with Transformer, and the MSE indicators of datasets ETTh1 and ETTm1 are improved by up to 34.84% and 42.60% respectively.
[0111] To further illustrate the advanced nature of MSFformer, the prediction results of oil temperature features in the dataset ETTh1 are visualized. Figure 6 As shown in the figure, 16 sub-graphs with four rows and four columns are shown. The input lengths represented by each column are {96, 192, 336, 720}, and the prediction lengths represented by each row are {96, 192, 336, 720}. Each sub-graph shows the visualization results of the true value, the MSFformer prediction value, and the Pyraformer prediction value. Figure 6It can be seen that:
[0112] The predicted results shown in the figure clearly demonstrate the significant superiority of the MSFformer model in terms of accuracy and dynamically capturing time series data. By adopting a multi-scale feature extraction strategy, the model exhibits outstanding performance in different prediction periods, which becomes particularly evident in the comparative analysis with the Pyraformer model. Especially in identifying and predicting peak phenomena in time series, the MSFformer shows higher sensitivity and excellent prediction ability. This is attributed to its complex mechanism, which can identify long-term and short-term dependencies.
[0113] Even when the prediction window is expanded, bringing more prediction challenges, the MSFformer still maintains commendable accuracy. In addition, the increase in the input length enhances the performance of the model, highlighting the effectiveness of multi-scale feature extraction in collecting extensive historical insights and precisely pointing out key time dependencies. These findings not only emphasize the robustness of the MSFformer in dealing with the complexity of time series prediction but also highlight its adaptability at different data scales, enhancing its potential for wide application in complex prediction scenarios.
[0114] To verify that the MSFformer model can effectively combine multi-scale features for time series prediction, an hourly dataset called the "synthetic dataset" was cited. This dataset has multi-range time dependencies and experiments were conducted based on it. This dataset is formed by linearly combining three sine functions with different periods (24 hours, 168 hours, 720 hours), representing daily (short-term), weekly (medium-term), and monthly (long-term) time scale dependencies respectively. In the experimental setup, both the input length and the prediction length are 720, and other settings remain the same as in the previous comparative experiment except for the window size and the intra-scale span parameter. Since both the deterministic part and the random part of the synthetic time series exhibit long-range correlations, the model must effectively capture this dependency to accurately predict the next 720 data points.
[0115] Table 3 Comparative Experiments on the Synthetic Dataset
[0116]
[0117]
[0118] The experimental results are shown in Table 3, which presents the MSE and MAE values of various methods on the synthetic dataset, where the best results are marked in bold and the second-best results are underlined. Among them, the MSFformer 24,6,5The representative window size is {24, 6, 5}, the corresponding constructed tree structure is {720, 30, 5, 1}, and the intra-scale span parameter is 7; MSFformer 20,6,6 The representative window size is {20, 6, 6}, the constructed tree structure is {720, 36, 6, 1}, and the intra-scale span parameter is 13. The results show that both variants of MSFformer with the two window sizes adopted achieve the best performance among all methods, especially MSFformer 20,6,6 improves by 16.22% compared to the best Pyraformer method in terms of the MSE metric and by 8.55% in terms of the MAE metric. Figure 7 Shows the visualization of the prediction results of MSFformer 20,6,6 The results indicate that by introducing the Skip-PAM module, MSFformer can effectively capture information at different time scales.
[0119] The second embodiment of the present invention relates to a transformer oil temperature prediction system based on an improved Transformer for implementing the method as described above, including:
[0120] An input module that collects the transformer oil temperature to obtain target time-series data;
[0121] An analysis module that uses a time-series prediction model to analyze the target time-series data to obtain a prediction result;
[0122] Among them, the time-series prediction model is constructed based on an improved Transformer and includes:
[0123] An encoding part that embeds a CGEM module at the encoder input end to extract different granularity time features of the target time-series data at multiple different scales, and introduces a stride pyramid attention mechanism through the Skip-PAM module to perform feature fusion on the multi-scale and multi-granularity time features;
[0124] A decoding part that analyzes and obtains a prediction result based on the fused multi-scale time features.
[0125] The following further illustrates this embodiment in combination with the specific time-series prediction model MSFformer (Multi-Scale Feature Transformer).
[0126] The overall framework of MSFformer proposed in this embodiment is as Figure 2As shown, based on the Transformer, this model improves the attention mechanism in the encoder and introduces a pyramid-shaped attention mechanism across time steps (Skip-PAM, Skip-Pyramidal Attention Module). To construct multi-scale time information, a feature convolution is designed to build the CGEM (Coarse-Grained Extraction Module) module to obtain time information at a coarse-grained level.
[0127] To adapt to feature extraction with a stride, a feature convolutional layer (FCNN, Feature CNN) is constructed as shown in Figure 3 . Specifically, for time series data of length l, it first passes through a convolutional kernel with the required stride step to extract and concatenate the feature vectors, and then passes through a convolutional kernel with a size of and a stride of for convolution operation. The formula is as follows:
[0128] Output = FCNN(Input)
[0129] To better apply the Skip-PAM structure, a pyramid-shaped feature tree needs to be constructed using the CGEM module. The CGEM module, as shown in Figure 4 , first passes the input time series through a linear layer to expand the feature dimension to a fixed dimension, then gradually passes through the feature convolutional layer to obtain feature information at different scales, and then concatenates them to form a pyramid-shaped feature structure. Finally, the feature dimension is restored through a linear layer:
[0130] M0 = Linear1(X)
[0131] M1 = FCNN1(M0)
[0132] M2 = FCNN2(M1)
[0133] M3 = FCNN3(M2)
[0134] M CGEM = Linear2([M0; M1; M2; M3])
[0135] where X is the input of the CGEM module, Linear is the linear layer, FCNN is the feature convolutional module mentioned in the above formula, M1, M2, and M3 are the results of successive feature convolutions, M CGEM is the output of the CGEM module, and the ';' operation concatenates M0, M1, M2, and M3 in the time dimension.
[0136] By separately setting the stride lengths of FCNN1, FCNN2, and FCNN3, time feature information with different granularities is obtained at different scales. A pyramidal feature tree is constructed by stacking feature convolutional layers, which not only enables the model to understand data from multiple levels but also provides a good foundation for the implementation of the Skip-PAM attention mechanism.
[0137] As Figure 5 shown is the time feature tree constructed by Skip-PAM based on the CGEM module. The attention mechanism extracts information from the time series at multiple scales, which can be divided into intra-scale connections and inter-scale connections. Intra-scale connections mean that a node performs attention calculations with adjacent nodes in its own scale layer, and inter-scale connections mean that a node performs attention calculations with its parent node (each parent node has P children) and C child nodes. Specifically, for a node where s ∈ [1, S] represents the s-th scale from the bottom layer to the top layer, and l ∈ [1, l s represents the l-th node in this layer:
[0138]
[0139] where represents adjacent nodes within the scale, A represents the number of adjacent nodes within the scale, represents child nodes, represents the parent node, is the sum of nodes for which the attention mechanism should be calculated. Then, according to the principle of the self-attention mechanism, for the node the attention can be expressed as:
[0140]
[0141] where q i is the query matrix corresponding to x i , k l and v l are the key matrix and value matrix respectively, d x is the sequence length, is the output result through the attention.
[0142] Through such a pyramidal attention mechanism, combined with the multi-scale feature convolution to be mentioned later, a powerful feature extraction network is formed, which can adapt to the dynamic changes of various time scales, whether it is short-term fluctuations or long-term evolution.
[0143] Other modules follow the Transformer model structure, where the decoder uses a fully connected layer to avoid the problem of error accumulation in the Transformer autoregressive decoder.
Claims
1. A transformer oil temperature prediction method, characterized in that, It includes the following steps: Collect the oil temperature of the transformer to obtain target time-series data; Use the time-series prediction model to analyze the target time-series data to obtain a prediction result; The time-series prediction model is constructed based on an improved Transformer, including: An encoding part, embedding a CGEM module at the input end of the encoder to extract different-granularity time features of the target time-series data at multiple different scales, and introducing a stride pyramid attention mechanism through a Skip-PAM module to perform feature fusion on the time features; A decoding part, analyzing and obtaining a prediction result according to the fused time features.
2. The method according to claim 1, characterized in that The CGEM module includes: A first linear layer, converting the initial feature dimension of the input time-series data into a set feature dimension; Several feature convolutional layers, sequentially extracting feature vectors of the input time-series data at different scales and splicing the feature vectors; A second linear layer, restoring the feature dimension of the spliced feature vectors to the initial feature dimension and then outputting.
3. The method according to claim 2, characterized in that, After the feature convolution layer extracts the feature vectors of the input time series data through a convolution kernel with a stride of step and concatenates them, it then performs a convolution operation on the concatenated feature vectors through a convolution kernel with a size of and a stride of , where l is the length of the input time series data.
4. The method according to claim 3, wherein The strides of the several feature convolutional layers are not all the same.
5. The method according to claim 1, characterized in that, The stride pyramid attention mechanism includes: Obtaining a time feature tree based on the multi-scale time features; For any node in the time feature tree, calculating the attention value of the current node according to the relationship between the current node and its adjacent nodes at the same scale, the relationship between the current node and its parent node, and the relationship between the current node and its child node.
6. The method according to claim 5, characterized in that The attention value is calculated by the following formula: Among them, is the sum of nodes for which the attention mechanism should be calculated. For node s ∈ [1, S] represents the s-th layer scale from the bottom layer to the top layer, and l ∈ [1, l s represents the l-th node in this layer. q i is the query matrix corresponding to node x i . k l and v l are the key matrix and value matrix respectively. d x is the sequence length, is the attention value.
7. The method according to claim 6, characterized in that The sum of nodes for which the attention mechanism should be calculated is calculated by the following formula: Among them, for a node represents its adjacent sibling nodes, and A represents the number of its adjacent sibling nodes. represents its child nodes. represents its parent node.
8. The method according to claim 1, wherein The analyzing and obtaining a prediction result according to the fused time features is implemented through a fully connected layer.
9. A transformer oil temperature prediction system, characterized in that, It includes: An input module, collecting the oil temperature of the transformer to obtain target time-series data; An analysis module, using the time-series prediction model to analyze the target time-series data to obtain a prediction result; The time-series prediction model is constructed based on an improved Transformer, including: An encoding part, embedding a CGEM module at the input end of the encoder to extract different-granularity time features of the target time-series data at multiple different scales, and introducing a stride pyramid attention mechanism through a Skip-PAM module to perform feature fusion on the time features; A decoding part, analyzing and obtaining a prediction result according to the fused time features.
10. The system according to claim 9, wherein The CGEM module includes: A first linear layer, converting the initial feature dimension of the input time-series data into a set feature dimension; Several feature convolutional layers, sequentially extracting feature vectors of the input time-series data at different scales and splicing the feature vectors; A second linear layer, restoring the feature dimension of the spliced feature vectors to the initial feature dimension and then outputting.