Power transformer temperature prediction method based on fragmented frequency domain enhancement
Through the chipped frequency domain enhancement method, discrete cosine transformation and cross attention mechanism are used to transfer the time series from the time domain to the frequency domain, and a multi-layer cross attention decoder module is built, which solves the problem of insufficient frequency domain feature extraction in the existing technology and improves the accuracy and robustness of the temperature prediction of power transformers.
Patent Information
- Application Number
- CN202510402267.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
AI Technical Summary
The existing time series prediction methods lack the frequency domain feature extraction capability in the temperature prediction of power transformers, resulting in insufficient prediction accuracy and insufficient adaptability of existing models to noise and small sample scenarios.
Using the sharded frequency domain enhancement method, the timing features are transferred from the time domain to the frequency domain through discrete cosine transformation and cross-attention mechanism, and a multi-layer cross-attention decoder module is built for feature extraction, enhancing the frequency domain feature capture and local feature extraction capabilities.
The accuracy and robustness of the temperature prediction of the power transformer are improved, and the adaptability and generalization ability of the model to noise scenes are enhanced.
Smart Images

Figure CN120354069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series prediction, and more specifically to a method for predicting the temperature of a power transformer based on slice-type frequency domain enhancement. Background Art
[0002] The power distribution problem is that the power grid manages the distribution of power to different user areas according to the sequentially changing demand. However, it is difficult to predict the future demand of a specific user area because it varies with different factors such as weekdays, holidays, seasons, weather, temperature, etc. Existing prediction methods cannot be applied to high-precision long-term predictions of long-term real-world data, and any wrong predictions may have serious consequences. Therefore, there is currently no effective way to predict future power consumption, and managers have to make decisions based on empirical values, and the threshold of empirical values is usually much higher than the actual demand. Conservative strategies lead to unnecessary waste of electricity and equipment depreciation. It is worth noting that the oil temperature of the transformer can effectively reflect the working condition of the power transformer. Therefore, predicting the oil temperature of the transformer can also try to avoid unnecessary waste.
[0003] The temperature data of power transformers is essentially time series data, which is mainly composed of the temperature and load values of each transformer at different time points. Time series prediction is a technology that uses a large amount of historical data to extract time-related features to predict future development trends. For a long time, a lot of research has focused on time series prediction, which can provide insights for decision support; however, most time series prediction models mainly focus on feature extraction in the time domain and ignore the features in the frequency domain, which will lead to insufficient extraction of details, cycles and other features during the model training process.
[0004] Benefiting from the progress of deep learning, time series prediction methods have made great progress, such as recurrent neural network (RNN), convolutional neural network (CNN) and Transformer and its related variants. However, since RNN is prone to the problem of gradient vanishing, and CNN is limited by its local receptive field in capturing long-term dependencies, these two types of methods still have many shortcomings and challenges. Transformer has become the focus of research in the field of time series prediction due to its inherent ability to perceive long-distance temporal dependencies.
[0005] In recent years, many Transformer-based time series prediction methods have been proposed. Usually, time series are marked using multi-resolution, such as time points, and self-attention mechanisms are used to model their dependencies. These methods have shown impressive performance. Despite the success of many current methods, they still have problems in feature extraction in the frequency domain. Most models focus on feature extraction in the time domain and lack feature extraction in the frequency domain, which may lead to insufficient feature extraction. As a result, the final prediction effect does not reach the best and there is still room for improvement.
[0006] Therefore, improving the ability to extract features in the frequency domain to enhance the final prediction accuracy is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a power transformer temperature prediction method based on piecewise frequency domain enhancement. By using discrete cosine transform and cross-attention mechanism, the time series feature extraction is transferred from the time domain to the frequency domain. At the same time, considering the locality of the time series, the corresponding features are extracted by segmenting the time series and applying discrete cosine transform.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] The present invention provides a power transformer temperature prediction method based on piecewise frequency domain enhancement, including the following steps:
[0010] S1. Obtain the historical temperature data of the power transformer;
[0011] S2. Normalize the historical temperature data and perform segmentation processing;
[0012] S3. Perform discrete cosine transform on the segmented data to obtain frequency domain components, and perform data dimensionality increase to obtain the first data dimensionality increase result;
[0013] S4. Construct a virtual output according to the space of the frequency domain components and the time series length T of the temperature data to be predicted, and perform the data dimensionality increase to obtain the second data dimensionality increase result;
[0014] S5. Construct a decoder module, input the first data dimensionality increase result and the second data dimensionality increase result into the decoder module for feature extraction to obtain the frequency domain features corresponding to the time series;
[0015] S6. Perform time series prediction on the frequency domain features, and the obtained prediction result is the temperature data result within the future time series length T.
[0016] Furthermore, it further includes:
[0017] S7. Use error evaluation indicators to compare the predicted results with the actual results and optimize the model.
[0018] Furthermore, the historical temperature data in step S1 is multivariate time series data, which is composed of time series, temperature and load of different transformers.
[0019] Further, step S2 specifically includes:
[0020] S21, the temperature data x from time point t-L+1 to t t-L+1:t After normalization, the formula is expressed as:
[0021]
[0022] Among them, mean(x t-L+1:t ) is the temperature data x t-L+1:t The mean, std(x t-L+1:t ) is the temperature data x t-L+1:t The standard deviation of M is the dimension of the power transformer temperature data, L is the length of the input time series; R represents the real number set, R M*L Represents the real number space of M*L dimensions;
[0023] S22, after normalization, the data x norm Add padding values at the end of , and perform sharding according to the configuration of shard size as patch and step length as stride; the formula is expressed as:
[0024]
[0025] Where P is the size of each patch, N is the number of patches, and padding is the number of values added at the end of the data sequence for a complete patch.
[0026] Further, step S3 specifically includes:
[0027] S31, shard data x patch Perform discrete cosine transform to obtain the frequency domain component x dct , the formula is:
[0028]
[0029] Wherein, DCT stands for discrete cosine transform;
[0030] S32, the frequency domain component x is processed by a linear layer of size P*D dct Perform dimension upgrade to obtain the first data dimension upgrade result The formula is:
[0031]
[0032] Among them, Proj represents a linear layer of size P*D, and D is the dimension size after mapping.
[0033] Furthermore, step S4 specifically includes:
[0034] Construct a virtual output x of size and perform dimensionality increase on the virtual output x through the linear layer of size P*D dummy to obtain the second dimensionality increase result of the data. dummy It is expressed by the formula: It is expressed by the formula:
[0035]
[0036] where N pred is the number of shards calculated according to the time series length T of the temperature data to be predicted.
[0037] Furthermore, step S5 specifically includes:
[0038] Construct a decoder module, which is composed of multiple layers of cross-attention;
[0039] Use the first dimensionality increase result of the data as the key and value inputs of the decoder module, and use the second dimensionality increase result of the data as the query input of the decoder module, and perform feature extraction through multiple layers of cross-attention to obtain the frequency domain features corresponding to the time series; It is expressed by the formula:
[0040]
[0041] where are the frequency domain features corresponding to the time series.
[0042] Furthermore, in step S5, the feature extraction through multiple layers of cross-attention specifically includes:
[0043] Split query, key, and value into n heads sub-heads, and the dimension of each head is d head ; Each sub-head calculates the attention score attention using scaled dot product, and the formula is expressed as:
[0044]
[0045] Apply softmax to obtain the attention weights, and use the results of combining multiple heads to obtain the final output of the attention layer. The formula is expressed as:
[0046]
[0047] Among them, and as the input sequence of each attention head, representing the learned frequency-domain features; Q represents the query vector after the input sequence is transformed by the weight matrix W Q ; K represents the key vector after the input sequence is transformed by the weight matrix W K ; V represents the value vector after the input sequence is transformed by the weight matrix W V ; d head represents the output dimension of each attention head, n heads represents the number of heads in the multi-head attention mechanism; h represents the h-th sub-attention head, Q h , K h , V h represent the vectors of Q, K, and V input to the h-th sub-attention head; W O represents the output weight matrix, which is used to map the concatenated vector to the final output dimension.
[0048] Furthermore, step S6 specifically includes:
[0049] S61. Perform data dimensionality reduction on the frequency-domain features through a linear layer of size D*P to obtain a dimensionality reduction result; and perform inverse discrete cosine transform on the dimensionality reduction result through an IDCT layer to obtain an inverse discrete cosine transform result The formula is expressed as:
[0050]
[0051] S62. Process the inverse discrete cosine transform result through an expansion layer and an inverse normalization layer to obtain a prediction result, and the formula is expressed as:
[0052]
[0053] where T is the length of the future time series, is the temperature data result within the predicted future time series length T.
[0054] Through the above technical solutions, it can be seen that compared with the prior art, the present invention discloses a power transformer temperature prediction method based on piecewise frequency-domain enhancement, which has the following beneficial effects:
[0055] 1) Enhanced frequency-domain feature extraction ability: The present invention transforms the time series from the time domain to the frequency domain through discrete cosine transform, and uses cross-attention to learn the weights of the frequency-domain components, learning the features in the frequency domain, enabling the model to better capture and utilize the features in the time series, enhancing the model's feature extraction ability for time series, and thus improving the prediction accuracy.
[0056] 2) Enhanced local feature extraction ability: The present invention divides the time series into continuous local segments through slicing operations, and captures fine-grained time series patterns through frequency-domain attention within each segment, enhancing the model's local feature extraction ability.
[0057] 3) Improved generalization ability: The decoder module constructed by the present invention can effectively suppress the influence of noise in the high frequency on the data through the cross-attention mechanism, avoid being affected by noise or small sample problems, and strengthen the adaptability and robustness of the model in the face of noise scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0059] Figure 1 It is a schematic diagram of a method for predicting the temperature of a power transformer based on slice-based frequency-domain enhancement provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0061] Embodiment 1
[0062] The embodiment of the present invention discloses a method for predicting the temperature of a power transformer based on slice-based frequency-domain enhancement. Referring to Figure 1 as shown, it includes the following steps:
[0063] S1. Obtain the historical temperature data of the power transformer;
[0064] S2. Perform normalization processing on the historical temperature data and perform slicing processing;
[0065] S3. Perform discrete cosine transform on the sharded data to obtain frequency domain components, and perform data dimensionality increase to obtain the first data dimensionality increase result;
[0066] S4. Construct a virtual output based on the space of the frequency domain components and the time series length of the temperature data to be predicted, and perform data dimensionality increase to obtain the second data dimensionality increase result;
[0067] S5. Construct a decoder module, input the first data dimensionality increase result and the second data dimensionality increase result into the decoder module for feature extraction to obtain the frequency domain features corresponding to the time series;
[0068] S6. Perform time series prediction on the frequency domain features, and the obtained prediction result is the temperature data result within the future time series length.
[0069] In the embodiment of the present invention, the above steps first extract the historical operation data of the power transformer from the database or log file of the power substation; through sharding, the locality of the time series is taken into account, and then for each shard, discrete cosine transform is used to convert the time series from the time domain to the frequency domain to obtain the frequency domain components of the time series, and the cross-attention mechanism is introduced to autonomously learn the weight ratio of different components in the final result, and the ratio of effective components is increased as much as possible. Finally, the frequency components are converted back to the time series through inverse discrete cosine transform to obtain the final prediction result.
[0070] In this embodiment, through sharding, the local characteristics of the time series can be better reflected, and through discrete cosine transform and cross-attention mechanism, the feature extraction of the time series can be improved to enhance the performance of the model in time series prediction.
[0071] The above steps are described in detail below:
[0072] In step S1, obtain the historical temperature data of the power transformer.
[0073] The temperature data of the power transformer is not just a time series data of a single variable. It is composed of a time series, the temperatures of different transformers, and the load, and is multi-variable time series data.
[0074] Among them, the time series means that each measurement point (such as temperature, load) has a timestamp, recording the state of this point at a specific moment; multiple variables, including but not limited to oil temperature, winding temperature, ambient temperature, humidity, load current, power, etc.; cross-device data, for multiple transformers, the mutual influence between different transformers needs to be considered.
[0075] In step S2, first perform normalization processing on the historical temperature data.
[0076] Considering the multivariate characteristics of historical temperature data, there may be significant differences in the data ranges corresponding to different attributes. Therefore, in order to reduce the problems caused by these data range differences, it is necessary to perform reversible normalization on the input of the model and denormalization on the output of the model before inputting the original data into the model.
[0077] Normalize the temperature data x from time point t - L + 1 to t. t-L+1:t The formula is as follows:
[0078]
[0079] where mean(x t-L+1:t ) is the mean of the temperature data x t-L+1:t , std(x t-L+1:t ) is the standard deviation of the temperature data x t-L+1:t , M is the dimension of the power transformer temperature data, L is the length of the input time series; R represents the set of real numbers, and R M+L represents the real space of dimension M + L.
[0080] Secondly, slice the normalized data.
[0081] First, add padding values at the end of the normalization result x norm . The added values are the last value of the sequence to extend the sequence. Then, slice the time series x norm to obtain the sliced result x patch . The purpose of slicing is to consider the locality of the time series, which is beneficial for capturing local features.
[0082] The following details the process of slicing:
[0083] First, slice the time series x norm with a patch size of patch and a stride of stride; x patch = Patch(x norm ), which belongs to
[0084] The formula is as follows:
[0085]
[0086] where P is the size patch of each slice, N is the number of slices, and padding is the number of values added at the end of the data sequence for complete slicing.
[0087] In step S3, use the discrete cosine transform to obtain the frequency components and map them to a higher dimension through a linear layer to obtain the first data dimensionality increase result.
[0088] Perform discrete cosine transform on the sharding result x of step S2, that is, DCT transform; transfer feature extraction from the time domain to the frequency domain to obtain the frequency domain components x of the time series patch ; Then, in order to learn more features, a linear layer of size P*D is used to dct dimensionality increase x dct .
[0089] The DCT transform formula is:
[0090]
[0091] The DCT function is:
[0092]
[0093] The normalization factor is:
[0094]
[0095] The data dimensionality increase formula is:
[0096]
[0097] Among them, DCT represents discrete cosine transform, and the number of frequency components obtained after DCT transform of the time series x with length N is N. X[k] is the k-th component in the frequency components, a k is the normalization factor, k represents the position of the frequency domain component, k = 0, 1, 2,..., N - 1; j represents the position index in the time series, j = 0, 1, 2…, N - 1, x j represents the j-th value in the time series; Proj represents a linear layer of size P*D, D is the size of the mapped dimension, and N is the number of shards
[0098] In step S4, construct a virtual output as the input and perform dimensionality increase on it to obtain the second data dimensionality increase result
[0099] First, construct a virtual output x of size as one of the inputs to the model decoding layer, where N dummy is the number of shards calculated according to the time series length T of the temperature data to be predicted. Then, use the same linear layer in step S3 to perform dimensionality increase on the virtual output to obtain pred This process of data dimensionality increase formula is expressed as:
[0100]
[0101]
[0102] In this embodiment, steps S1 - S4 serve as the input module of the prediction model of the present invention, and are used to construct features for the input multi - variable time series, including normalizing each column, segmenting and slicing, and performing DCT transformation, so as to enhance the features in the frequency domain provided by the sequence.
[0103] In step S5, a decoder module composed of multiple layers of cross - attention is used for feature extraction to learn the features of the time series in the frequency domain.
[0104] The decoder module is composed of multiple layers of cross - attention layers. By using the up - dimension result of the first data output in step S3 as the key and value input of the decoder module, and using the up - dimension result of the second data output in step S4 as the query input of the decoder module, after multiple layers of cross - attention, all the features extracted by the decoder module are obtained.
[0105] The formula representation of this process is:
[0106]
[0107] Among them, is the frequency - domain feature corresponding to the time series.
[0108] The process of the decoder extracting features through multiple layers of cross - attention is introduced in detail below:
[0109] The decoder module is a TransformerDecoder architecture composed of multiple layers of cross - attention layers. For each input sequence, it is mapped to query Q (query), key K (key), and value V (value) through three groups of learnable weight matrices. Split Q, K, and V into n heads sub - heads, and the dimension of each head is d head ; each sub - head uses scaled dot - product to calculate the attention score attention, and this process is represented by the formula:
[0110]
[0111] Among them, the product of Q and K represents the correlation between subsequences in different context contexts; then apply softmax to obtain the attention weights, indicating the importance in different contexts, and these weights are used to calculate V by weighting. The results of combining multiple heads are used to obtain the final output of the attention layer.
[0112] This process is represented by the formula:
[0113]
[0114] Among them, and is the input sequence for each attention head, representing the learned frequency domain features; Q represents the query vector after the input sequence is transformed by the weight matrix W Q ; K represents the key vector after the input sequence is transformed by the weight matrix W K ; V represents the value vector after the input sequence is transformed by the weight matrix W V ; d head represents the output dimension of each attention head, n heads represents the number of heads in the multi-head attention mechanism; h represents the h-th sub-attention head, Q h , K h , V h represent the vectors of Q, K, and V input to the h-th sub-attention head; W O represents the output weight matrix, which is used to map the concatenated vector to the final output dimension.
[0115] In this embodiment, step S5 serves as the backbone module of the prediction model of the present invention and is used to extract features from the constructed time series frequency components. The backbone module consists of a decoder module composed of a multi-layer cross-attention layer. Each layer of the cross-attention layer consists of a multi-head attention and a feed-forward neural network layer, which is used to learn the proportion of different components in the frequency components to achieve feature extraction in the frequency domain.
[0116] In step S6, a prediction module is used to predict the time series.
[0117] The prediction module consists of a linear layer, an IDCT layer, an unfolding layer, and an inverse normalization layer. Among them, the linear layer first reduces the dimension of the output of step S5 , and the size of this linear layer is D*P; then the IDCT layer performs an inverse discrete cosine transform on the output of the linear layer to obtain Finally, it is restored through the unfolding layer and the inverse normalization layer to obtain the final prediction result
[0118] The formulas for dimension reduction and IDCT are:
[0119]
[0120] The formulas for unfolding and inverse normalization are:
[0121]
[0122] where T is the length of the future time series, is the temperature data result within the length T of the predicted future time series.
[0123] This embodiment further includes step S7, comparing the predicted result with the true result using an error evaluation index, and optimizing the model.
[0124] Among them, the error evaluation index is the combination of the mean squared error MSE and the trend loss Trend, and the formula is expressed as:
[0125]
[0126] Among them, represents the predicted value corresponding to the training result of the i-th sample of the prediction model; y i represents the true value corresponding to the training result of the i-th sample of the prediction model, and n is the total number of samples in the prediction model.
[0127] This embodiment further verifies the performance of the prediction method of the present invention.
[0128] The dataset for the performance evaluation of the present invention comes from the load and temperature characteristics of seven power transformers from July 2016 to July 2018. And the dataset is divided into a training set, a validation set, and a test set according to 7:1:2.
[0129] In this example, two commonly used indicators in time series prediction are adopted: Mean Absolute Error (MAE) and Mean Square Error (MSE). And the prediction performance of the method of the present invention is evaluated using these two error evaluation models.
[0130] The comparison of the prediction effects of the model of the present invention and other models on the above dataset is shown in Table 1 below. In the table, 96, 192, 336, and 720 represent the length of the prediction sequence. The bold indicates the optimal index under the same standard, and the underline indicates the index that is second best under the same standard:
[0131]
[0132] Table 1 Comparison of model prediction results
[0133] In Table 1, the FreqCATS method is the model of the present invention, while CATS, TimeMixer, PatchTST, Times, Crossformer, MICN, FiLM, DLinear, Autoformer, and Informer are respectively more advanced methods in the field of time series prediction in recent years. Among them, there are methods improved based on Transformer and methods that perform predictions through simple linear layers.
[0134] It can be seen from the comparison with the prediction results of other models that the prediction method proposed in the present invention has higher accuracy than other models, indicating that the prediction method proposed in the present invention can obtain better results in the temperature prediction of power transformers.
[0135] In the present invention, the extraction of time series features is transferred from the time domain to the frequency domain through the discrete cosine transform (DCT) and the cross-attention mechanism. At the same time, considering the locality of the time series, the time series is sliced and the DCT is applied to extract the corresponding features. The method of the present invention can effectively extract the frequency domain features of the time series, improving the accuracy and robustness of the model prediction.
[0136] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.
[0137] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A temperature prediction method for power transformers based on sharded frequency domain enhancement, characterized in that It includes the following steps: S1. Obtain the historical temperature data of the power transformer; S2. Normalize the historical temperature data and perform slicing processing; S3. Perform discrete cosine transform on the sliced data to obtain frequency domain components, and perform data dimensionality increase to obtain the first data dimensionality increase result; S4. Construct a virtual output according to the space of the frequency domain components and the time series length of the temperature data to be predicted, and perform data dimensionality increase to obtain the second data dimensionality increase result; S5. Construct a decoder module, input the first data dimensionality increase result and the second data dimensionality increase result into the decoder module for feature extraction to obtain the frequency domain features corresponding to the time series; S6. Perform time series prediction on the frequency domain features, and the obtained prediction result is the temperature data result within the future time series length.
2. The method for predicting the temperature of a power transformer based on piecewise frequency domain enhancement according to claim 1, wherein, It also includes: S7. Compare the prediction result with the real result using an error evaluation index and optimize the model.
3. A method for predicting the temperature of a power transformer based on piecewise frequency-domain enhancement as claimed in claim 1, wherein The historical temperature data in step S1 is multivariate time series data, which consists of a time series, the temperatures and loads of different transformers.
4. The method for predicting the temperature of a power transformer based on piecewise frequency domain enhancement according to claim 1, wherein Step S2 specifically includes: S21. Normalize the temperature data x from time point t - L + 1 to t, which is expressed by the formula as follows: t-L+1:t to be where, mean(x t-L+1:t ) is the mean value of the temperature data x t-L+1:t , std(x t-L+1:t ) is the standard deviation of the temperature data x t-L+1:t , M is the dimension of the power transformer temperature data, and L is the length of the input time series; R represents the set of real numbers, and R M*L represents the real number space of M * L dimensions; S22. Add padding values at the end of the normalized data x norm and slice it according to the configuration with the slice size of patch and the stride of stride; The formula is expressed as: Where P is the size of each slice patch, N is the number of slices, and padding is the number of values added at the end of the data sequence for complete slicing.
5. The temperature prediction method for power transformers based on piecewise frequency domain enhancement according to claim 4, characterized in that Step S3 specifically includes: S31. Perform discrete cosine transform on the sharded data x patch to obtain the frequency domain component x dct , which is expressed by the formula as follows: Where DCT represents discrete cosine transform; S32. Perform dimensionality increase on the frequency domain component x through a linear layer of size P*D dct to obtain a first data dimensionality increase result The formula is expressed as: Where Proj represents a linear layer of size P*D, and D is the size of the mapped dimension.
6. The method for predicting the temperature of a power transformer based on piecewise frequency-domain enhancement according to claim 5, wherein, Step S4 specifically includes: Construct a virtual output \(x\) with a size of , and dimensionally up-sample the virtual output \(x\) dummy through the linear layer of size \(P\times D\) to obtain a second data dimensional up-sampling result dummy . It is expressed by the formula as: Formula representation: where N pred is the number of shards calculated according to the time series length T of the temperature data to be predicted.
7. The method for predicting the temperature of a power transformer based on piecewise frequency domain enhancement according to claim 6, characterized in that, Step S5 specifically includes: Construct a decoder module, and the decoder module is composed of multiple layers of cross attention; Regarding the upsampled result of the first data as the key and value inputs of the decoder module, and regarding the upsampled result of the second data as the query input of the decoder module, perform feature extraction through multiple layers of cross-attention to obtain the frequency-domain features corresponding to the time series; the formula is expressed as: Among them, is the frequency domain feature corresponding to the time series.
8. The method for predicting the temperature of a power transformer based on piecewise frequency domain enhancement according to claim 7, characterized in that, In step S5, the feature extraction through multiple layers of cross attention specifically includes: Split query, key, and value into n heads sub - heads, each with dimension d head ; each sub - head calculates the attention score attention using scaled dot - product, which is expressed by the formula: Apply softmax to obtain the attention weights, and use the result of merging multiple heads to obtain the final output of the attention layer. The formula is expressed as: Among them, and serve as the input sequences for each attention head, representing the learned frequency-domain features; Q represents the query vector after the input sequence is transformed by the weight matrix W Q ; K represents the key vector after the input sequence is transformed by the weight matrix W K ; V represents the value vector after the input sequence is transformed by the weight matrix W V ; d head represents the output dimension of each attention head, and n heads represents the number of heads in the multi-head attention mechanism; h represents the h-th sub-attention head, and Q h , K h , V h represent the vectors of Q, K, and V input to the h-th sub-attention head; W O represents the output weight matrix, which is used to map the concatenated vector to the final output dimension.
9. The method for predicting the temperature of a power transformer based on piecewise frequency domain enhancement according to claim 7, characterized in that, Step S6 specifically includes: S61. Perform data dimensionality reduction on the frequency domain features through a linear layer of size D*P to obtain a dimensionality reduction result; and perform inverse discrete cosine transform on the dimensionality reduction result through an IDCT layer to obtain an inverse discrete cosine transform result The formula is expressed as: The formula is expressed as: S62. Process the inverse discrete cosine transform result through an expansion layer and an inverse normalization layer to obtain a prediction result, which is expressed by the formula: where T is the length of the future time series, is the temperature data result within the predicted future time series length T.
Citation Information
Cited By
Oil temperature prediction method for oil-immersed power transformer
CN122470941A
Oil temperature prediction method for oil-immersed power transformer
CN122470941B