Method for predicting top oil temperature of oil-containing equipment, electronic equipment and storage medium

Through dual-channel feature extraction and multi-scale time convolution network model, the problem of insufficient spatial and frequency domain feature extraction in the top-level oil temperature prediction of power equipment is solved, and efficient and accurate top-level oil temperature prediction is achieved, which improves the safety, stability and energy efficiency of power equipment.

CN120408080APending Publication Date: 2025-08-01STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510482698.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing deep learning algorithms are difficult to effectively extract the spatio-temporal and frequency domain characteristics of the top oil temperature of power equipment, and single-scale features are difficult to accurately predict the top oil temperature, and they are unable to fully capture the correlation and important feature information of the oil temperature time series.

Method used

A dual-channel feature extraction module is adopted. One channel uses a spatiotemporal attention mechanism to extract spatiotemporal features, and the other channel uses fast Fourier changes to convert to the frequency domain and uses a convolutional neural network to extract frequency domain features. Feature fusion is performed through the cross attention mechanism, and a multi-scale temporal convolution network model of multi-scale expansion causal convolution and deep learning algorithms is designed to capture multi-scale features from global and local angles, and finally output the prediction results through a fully connected layer and non-normalized output.

Benefits of technology

It improves the accuracy and robustness of the top-level oil temperature prediction of power equipment, reduces the calculation complexity, and can quickly and accurately predict the top-level oil temperature to meet the needs of actual engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408080A_ABST
    Figure CN120408080A_ABST
Patent Text Reader

Abstract

The invention discloses a top oil temperature prediction method for oil-containing equipment, and solves the problems that spatial-temporal features and frequency-domain features representing an oil temperature sequence are difficult to effectively extract by using a single deep learning algorithm, and top oil temperature prediction is difficult to accurately carry out by using single-scale features. Wherein one channel uses a space-time attention mechanism to extract space-time features of normalized data, and the other channel uses fast Fourier transform to convert original top oil temperature data from a time domain to a frequency domain, and uses a convolutional neural network to extract frequency domain features; performing spatial-temporal feature and frequency domain feature fusion by using a cross attention mechanism; designing a multi-scale time convolution network model based on multi-scale expansion causal convolution and a deep learning algorithm, and capturing multi-scale features from global and local angles; outputting a prediction result by using a full connection layer and non-standardization; according to the method, the calculation complexity is relatively low, and the top oil temperature of the power equipment can be quickly predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil temperature prediction for power equipment, and specifically relates to a method for predicting the top oil temperature of oil-containing equipment, an electronic device, and a storage medium. Background Art [[ID=X]]

[0002] With the rapid advancement of power development in China, the capacity of newly commissioned power generation units has increased significantly, and the achievements in the construction of the national power grid are remarkable. Predicting the top oil temperature is crucial for oil-containing power equipment in the power system, as it can detect equipment abnormalities in advance and prevent failures and safety accidents. Excessive temperature will accelerate the aging of insulating materials and shorten the service life of the equipment. Predicting the oil temperature helps optimize equipment maintenance and extend the operation cycle. By monitoring the change of oil temperature, the load dispatching of the equipment can be optimized, and the stability and reliability of the system can be improved. The oil temperature prediction can also be used as a fault warning indicator to diagnose equipment problems in advance and reduce the impact of sudden failures. Finally, reasonable temperature management can improve the energy efficiency of the equipment and reduce the operation cost. Therefore, ensuring the reliability of power equipment itself is a key prerequisite for ensuring the reliability of the power grid. During the use of power equipment, improving the maintenance and repair level and eliminating or reducing the probability of various latent and operating failures are of great significance for the safe operation of the power system.

[0003] Currently, the prediction of the top oil temperature of power equipment mainly includes numerical models, semi-physical models, and data-driven models. Numerical models usually rely on physical laws and thermodynamics principles to construct mathematical equations for heat transfer inside power equipment to predict the top oil temperature. This method requires obtaining the structural parameters of power equipment and involves a large amount of calculation, which is not suitable for monitoring the top oil temperature of large-scale power equipment. The semi-physical model based on heat transfer has the disadvantages of large errors and difficult-to-obtain heat transfer parameters, and is not widely used. Machine learning methods based on data-driven provide a new solution for the intelligent prediction of the oil temperature of power equipment, such as support vector machines, artificial neural networks, etc. These methods show good prediction accuracy and generalization ability. However, machine learning methods do not fully consider the influence between different operating parameters and have not explored the correlation between data. It can be seen that traditional oil temperature prediction methods often have problems such as low prediction accuracy, high model complexity, and difficulty in adapting to the changes of large-scale power grids. Physical model prediction depends on a deep understanding of the internal heat conduction process of the equipment, but the structure of power equipment is complex, and it is difficult to establish an accurate physical model. Although the prediction method based on statistics is simple and easy to implement, its prediction effect is often unsatisfactory when facing the complexity and uncertainty of large-scale power grids. Therefore, developing an efficient and accurate oil temperature prediction method is of great significance for ensuring the safe and stable operation of the power system.

[0004] In recent years, with the continuous development of deep learning methods, algorithms such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have emerged. With their outstanding nonlinear fitting capabilities, they have achieved great success in the field of top-layer oil temperature prediction. However, some problems still exist: 1) Current prediction models extract a single top-layer oil temperature feature, which cannot fully characterize the spatiotemporal and frequency domain characteristics of the top-layer oil temperature series. 2) Current prediction models extract single-scale features that make it difficult to accurately predict the top-layer oil temperature. 3) Current prediction models do not consider the correlations and important feature information in the oil temperature time series, and fail to capture both global and local features. Therefore, a single deep learning algorithm is difficult to effectively predict the top-layer oil temperature of power equipment. Summary of the Invention

[0005] The technical solution of the present invention is used to solve the problem that it is difficult to effectively extract the spatiotemporal features and frequency domain features that characterize the oil temperature series using a single deep learning algorithm, and that single-scale features are difficult to accurately predict the top oil temperature.

[0006] The present invention solves the above technical problems through the following technical solutions:

[0007] The present invention provides a method for predicting the top oil temperature of oil-containing equipment, comprising:

[0008] Normalize the original top oil temperature dataset;

[0009] Design a dual-channel feature extraction module. One channel uses a spatiotemporal attention mechanism to extract spatiotemporal features of the normalized data. The other channel uses a fast Fourier transform to transform the raw top oil temperature data from the time domain to the frequency domain and uses a convolutional neural network to extract frequency domain features.

[0010] Use cross-attention mechanism to fuse spatiotemporal features and frequency domain features;

[0011] A multi-scale temporal convolutional network model is designed based on multi-scale dilated causal convolution and deep learning algorithms to capture multi-scale features from both global and local perspectives;

[0012] Use fully connected layers and unnormalize the output predictions.

[0013] Furthermore, the spatiotemporal attention mechanism is composed of a temporal attention unit and a spatial attention mechanism connected in series. The temporal attention unit is used to adaptively assign attention weights to all time points and extract temporal features; the spatial attention mechanism is used to adaptively assign attention weights to input variables and extract spatial features.

[0014] Furthermore, the temporal attention unit is designed based on a bidirectional gated recurrent unit.

[0015] Further, the formula for extracting frequency-domain features using a convolutional neural network is as follows:

[0016]

[0017] where X(t) represents the normalized data, is the frequency-domain signal after the fast Fourier transform, and w is the angular frequency.

[0018] Further, the calculation formula for fusing spatio-temporal features and frequency-domain features using the cross-attention mechanism is as follows:

[0019] F S = Softmax((W q F x )(W k h x ))W T )W v h x

[0020] where W k and W v are the linear transformation matrices of the key and value respectively, W q is the linear transformation matrix of the query, F x is the extracted frequency-domain feature, h x is the extracted spatio-temporal feature, F S is the fused feature, and softmax() is the normalization activation function.

[0021] Further, the multi-scale temporal convolutional network model consists of multi-scale dilated causal convolutions, self-attention, and fully connected layers. After obtaining features of different scales through multi-scale dilated causal convolutions, deep learning algorithms are used to capture global and local connections.

[0022] Further, the calculation formula of the multi-scale temporal convolutional network model is as follows:

[0023]

[0024] where, is the DCC result of the i-th channel, f is the convolution kernel size, F S is the input fused feature, T is the time length of the input data, X c is the output of the multi-scale dilated causal convolution, W Q , W K , W V are the linear transformation matrices of the triple, is the scaling factor, H is the output of the deep learning algorithm, FC() is the calculation of the fully connected layer, is the prediction result.

[0025] Furthermore, the non - normalized formula is as follows:

[0026]

[0027] where x max is the maximum value in the oil temperature data, and x min is the minimum value in the oil temperature data. and are the predicted values of the oil temperature before and after inverse normalization respectively.

[0028] The present invention also provides an electronic device, including a memory and a processor. The memory is used to store a program that supports the processor to execute the above - mentioned method for predicting the top - layer oil temperature of the oil - containing device, and the processor is configured to execute the program stored in the memory.

[0029] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above - mentioned method for predicting the top - layer oil temperature of the oil - containing device.

[0030] The beneficial effects of the present invention are as follows:

[0031] First, the present invention performs a normalization process on the original top - layer oil temperature data set, reduces the influence of outliers caused by the external environment, reduces the computational amount of the model, and meets the requirements of rapidity and accuracy in practical engineering applications. Then, a dual - channel feature extraction module is designed. One channel uses a newly designed spatio - temporal attention module (STAM) to extract spatio - temporal features of the normalized data. Specifically, a time attention unit (TAU) based on BiGRU is used to extract temporal features, and then an improved spatial attention (SAM) is used to extract spatial features; the other channel uses the fast Fourier transform to convert the original top - layer oil temperature data from the time domain to the frequency domain, and uses a CNN to extract frequency - domain features; further considering the sensitivity between features, cross - attention is used for feature fusion; then, based on multi - scale dilated causal convolution (MSDCC) and deep learning algorithms, a multi - scale temporal convolutional network (MSTCN) is designed to enable the model to focus on the correlation and important feature information in the oil temperature time series, and capture multi - scale features from both global and local perspectives; finally, a fully - connected layer is used to output the predicted result of the top - layer oil temperature. The technical solution of the present invention has a small computational complexity, can quickly predict the top - layer oil temperature of power equipment, and can improve the accuracy and robustness of the prediction of the top - layer oil temperature of power equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the overall flowchart of the method for predicting the top - layer oil temperature of the oil - containing device based on multi - scale dilated causal convolution in the first embodiment of the present invention;

[0033] Figure 2It is the CNN structure diagram of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0034] Figure 3 It is the GRU structure diagram of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0035] Figure 4 It is the BiGRU structure diagram of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0036] Figure 5 It is the STAM structure diagram of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0037] Figure 6 It is the feature fusion flow chart of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0038] Figure 7 It is the structure diagram of MSTCN of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention;

[0039] Figure 8 It is the top oil temperature prediction curve graph of the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the first embodiment of the present invention. Detailed implementation manners

[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0041] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments:

[0042] Embodiment 1

[0043] As Figure 1 shown, the top oil temperature prediction method for oil-containing equipment based on multi-scale dilated causal convolution in the embodiment of the present invention is specifically as follows:

[0044] 1) Data normalization

[0045] To reduce the influence of outliers caused by interference factors such as the external environment and reduce the computational complexity of the model to meet the requirements of fast and accurate in practical engineering applications, the present invention normalizes the original top-layer oil temperature data. Linear function normalization is used to constrain the oil temperature data set between [0, 1]. The calculation formula is as follows:

[0046]

[0047] where x represents the oil temperature data before normalization, x' represents the oil temperature data after normalization, x max is the maximum value in the oil temperature data, x min is the minimum value in the oil temperature data.

[0048] 2) Fast Fourier Transform (FFT)

[0049] The frequency domain of the oil temperature data contains important features not included in the time domain. Therefore, it is necessary to obtain the frequency domain signal. The Discrete Fourier Transform (DFT) is a classic method for obtaining the frequency domain signal of the original signal, but its computational complexity is huge. To reduce the number of DFT calculations on the computer, the present invention uses the Fast Fourier Transform (FFT) to obtain the frequency domain signal. The FFT formula is as follows:

[0050]

[0051] where X(t) represents the normalized data, is the frequency domain signal after the Fast Fourier Transform, and w is the angular frequency.

[0052] 3) Convolutional Neural Network (CNN)

[0053] The network structure of the CNN includes convolutional layers, pooling layers, and fully connected layers, etc., which can automatically learn features and perform effective feature representation, as Figure 2 shown. The CNN gradually extracts low-level and high-level features in the data through convolutional operations and pooling operations. After these features are combined through multiple convolutional layers, they are finally output by the fully connected layer. The present invention uses a one-dimensional convolutional neural network to extract features from the obtained frequency domain signal. The calculation formula of the one-dimensional convolutional neural network is as follows:

[0054]

[0055] where w i is the convolutional kernel, x t-i+1 is the time series, y t is the output feature after convolution, and k is the size of the convolutional kernel.

[0056] 4) Bidirectional Gated Recurrent Unit (BiGRU)

[0057] Due to the problems of vanishing gradients and exploding gradients in traditional Recurrent Neural Networks (RNNs), the introduction of forget gates, reset gates, and input gates in LSTM solves the above problems. The Gated Recurrent Unit (GRU) network is a variant of LSTM. It inherits the advantages of RNN and LSTM in time series prediction, integrates the forget gate and input gate in LSTM into the update gate, can effectively discard invalid information, capture long-range dependencies, greatly reduce the number of parameters, and avoid overfitting. The structure of the Gated Recurrent Unit network model is as Figure 3 shown.

[0058] The formula of the Gated Recurrent Unit network model is as follows:

[0059] z t = σ(W z x t + U z h t-1 ) (4)

[0061] r t = σ(W r x t + U r h t-1 ) (5)

[0063]

[0064] Among them, z t and r t represent the states of the update gate and the reset gate at time t respectively; h t-1 represents the input at the previous moment; h t represents the output; x t represents the input data; W z , W r , W h , U z , U r and U h all represent weight matrices; σ() represents the sigmoid function, represents the Hadamard product, represents the candidate hidden state.

[0065] The unidirectional GRU transmits unidirectionally from front to back, ignoring the influence of subsequent time series on the previous time series. Therefore, in the present invention, a bidirectional gated recurrent network with GRU is applied in both the forward and reverse directions, and a time attention unit (TAU) is designed based on this to improve the model prediction performance. The structure of BiGRU is as Figure 4 shown.

[0066] 5) Spatio-Temporal Attention Mechanism (STAM)

[0067] To achieve better prediction results, it is necessary to not only focus on the trend changes of temporal features but also emphasize the positional relationships of spatial features. Although CNN and GRU networks can extract spatio-temporal features from data, they cannot display spatio-temporal correlations in an interpretable way, resulting in unreliable prediction and monitoring results and difficulties in on-site applications. In addition, the GRU network has an inherent drawback of early information loss, which limits the accuracy of modeling. Therefore, the present invention designs a spatio-temporal attention mechanism, and its unique structure enables it to be interpreted by expressing the correlation between input data and attention weights. Specifically, the present invention designs a temporal attention unit that can adaptively assign attention weights to all time points. By expressing temporal correlations and attention weights, the extraction of temporal features becomes interpretable, and at the same time, the problem of early information loss in the GRU network is solved. In addition, a spatial attention mechanism is designed to effectively explore the spatial correlations of data and adaptively assign attention weights to input variables. By using attention weights that make the extraction of spatial features interpretable, the role of key input variables is strengthened, and redundant variables are weakened. The spatio-temporal attention mechanism (STAM) designed by the present invention is composed of a temporal attention unit (TAU) and a spatial attention mechanism (SAM), as Figure 5 shown.

[0068] The oil temperature data obtains the hidden temporal feature h in the TAU by BiGRU g , and further obtains the high-level temporal feature H through depthwise separable convolution. The dot product between h g and each token of is calculated by the scoring function F, and then the result score is obtained through linear mapping. Then, the Softmax function of score is multiplied by H and concatenated with h g , and then multiplied by the sigmoid function of score to obtain the output hg. The calculation formula of TAU is as follows:

[0069] h g = BiGRU(x) (8)

[0070] H = DWC(h g ) (9)

[0071] score = F(H, h g ) (10)

[0072] α = Softmax(score) (11)

[0073] h′ g = contact(αH, h g ) (12)

[0074]

[0075] Among them, h g is the hidden time feature, H is the mined high-level feature, α is the obtained weight, h′ g is the weighted sum of the hidden states at each moment in H combined with the hidden state h g h t is the output of TAU, softmax() is the normalization activation function, BiGRU() represents passing through a bidirectional gated recurrent unit, DWC() represents performing a depthwise separable convolution operation, F() represents the scoring function, contact() represents concatenating features, sigmoid() represents the activation function, and x represents the input.

[0076] In addition, the extracted temporal features are fed into SAM to further obtain spatial features. First, the input features are subjected to global max pooling and global average pooling in the channel dimension to compress the channel size, facilitating the subsequent learning of spatial features. Then, the results of global max pooling and global average pooling are concatenated along the channels. Different from the traditional CAM, in the present invention, the concatenated result is passed through 2 convolutions of different sizes in parallel, and the results of the convolutions are added together. Through parallel convolutions, two convolutional kernels of different scales can respectively focus on regions of different sizes, thereby capturing more detailed and extensive spatial information, learning more diverse spatial patterns, enhancing the model's ability to capture complex patterns, and improving the model's performance. Finally, the output is obtained by multiplying with the input through the Sigmoid activation function. The calculation formula of SAM is as follows:

[0077] F a = Avgpool(h t ) (14)

[0078] F m = Maxpool(h t ) (15)

[0079] S1 = Conv1(Concat(F a , F m )) (16)

[0080] S2 = Conv2(Concat(F a , F m )) (17)

[0081]

[0082] Among them, h t is the output of TAU, F a and F mThey are global average pooling and max pooling spatial features respectively. S1 and S2 are the concatenated two spatial features, which are the feature results after convolution operations of different sizes, F x is the output of SAM. Avgpool() represents average pooling, Maxpool() represents max pooling, and Conv1() and Conv2() both represent convolution operations.

[0083] 6) Cross-attention mechanism

[0084] During the fusion process, the frequency-domain features extracted by CNN through FFT transformation are used as the query sequence, and the spatio-temporal features extracted by STAM are used as the key-value pair sequence. The cross-attention mechanism is used to fuse the time-domain and frequency-domain features. In this way, the model can focus more on important features by calculating the attention weights. The cross-attention fusion is as Figure 6 shown.

[0085] The calculation formula for fusing spatio-temporal features and frequency-domain features using the cross-attention mechanism is as follows:

[0086] F S = Softmax((W q F x )(W k h x ))W T )W v h x (19)

[0087] Among them, W k and W v are the linear transformation matrices of the key and value respectively, W q is the linear transformation matrix of the query, F x is the extracted frequency-domain feature, h x is the extracted spatio-temporal feature, and F S is the fused feature.

[0088] 7) Multi-scale Temporal Convolutional Network (MSTCN) model

[0089] In order to focus on multi-scale features from both global and local perspectives, the present invention designs an MSTCN model. This model consists of multi-scale dilated causal convolution, self-attention, and fully connected layers. After obtaining features of different scales through multi-scale dilated causal convolution (MSDCC), deep learning algorithms are used to capture global and local connections, overcoming the problem of insufficient global information acquisition ability of convolution, exploring the oil temperature changes in different periods more deeply, and paying attention to the internal connections. Finally, the prediction result is output through the fully connected layer (FC) and non-normalization. The structure of the multi-scale temporal convolutional network model is as Figure 7 shown, and the calculation formula of the multi-scale temporal convolutional network model is as follows:

[0090]

[0091] Among them, is the DCC result of the i-th channel, f is the convolutional kernel size, and F S is the input fusion feature, T is the time length of the input data, and X c is the output of the multi-scale dilated causal convolution, and W Q , W K , W V is the linear transformation matrix of the triple, is the scaling factor, H is the output of the deep learning algorithm, and FC() is the calculation of the fully connected layer, is the prediction result.

[0092] At the end of the prediction, it is necessary to denormalize the result to evaluate the deviation from the true value. The denormalization formula is as follows:

[0093]

[0094] Among them, x max is the maximum value in the oil temperature data, and x min is the minimum value in the oil temperature data, and are the predicted values of the oil temperature before and after inverse normalization, respectively.

[0095] Through experiments, the predicted curve of the top oil temperature of the power equipment is obtained, and it can be seen that the method proposed by the present invention can well fit Figure 8 the true oil temperature in

[0096] The present invention provides a method for predicting the top oil temperature of power equipment based on multi-scale dilated causal convolution. First, the top oil temperature data is normalized. Then, the normalized data is input into the time channel and the frequency domain channel respectively. In the time channel, the designed STAM is used to extract spatio-temporal features. In the frequency domain channel, the data is first subjected to FFT transformation to convert it from the time domain to the frequency domain, and CNN is used to extract frequency domain features. Next, the obtained spatio-temporal features and frequency domain features are fused using cross-attention. Finally, the designed MSTCN is used to focus on the multi-scale features of the oil temperature from both the global and local perspectives, and the prediction result is output through the fully connected layer and denormalization. The method proposed in this invention patent can quickly and accurately predict the top oil temperature of power equipment. A new time attention unit (TAU) and a new spatial attention mechanism (SAM) are designed and connected in series, and further a spatio-temporal attention mechanism (STAM) is designed. STAM first mines the temporal features of the data, and then captures the spatial features on this basis, so as to extract the desired spatio-temporal features. This can dynamically adjust the attention degree to different time steps and spatial positions, and endow the model with unique interpretability advantages. It quantitatively describes the model training process with attention weights, thereby improving the reliability of the prediction result. Aiming at the problem that the single-scale features extracted by the current prediction model are difficult to accurately represent the top oil temperature, a dual-channel is designed to extract multiple features and fuse them. Specifically, the designed STAM is used to extract spatio-temporal features, CNN is used to extract the frequency domain features after FFT transformation, and cross-attention is used for feature fusion. Aiming at the problem that the single-scale features extracted by the prediction model are difficult to accurately represent the top oil temperature and lack feature extraction from both the global and local perspectives, a multi-scale temporal convolutional network (MSTCN) is designed. MSTCN first uses MSDCC to extract multi-scale features, then uses deep learning algorithms to make up for the deficiency of the convolution's ability to obtain global information, and finally uses the fully connected layer and denormalization to output the prediction result.

[0097] Embodiment 2

[0098] An electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the method for time synchronization of distributed nodes within a local area network in Embodiment 1, and the processor is configured to execute the program stored in the memory.

[0099] Embodiment 3

[0100] A storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the method for time synchronization of distributed nodes within a local area network in Embodiment 1.

[0101] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the top oil temperature of an oil-containing device, characterized in that, Including: Normalize the original top-layer oil temperature dataset; Design a dual-channel feature extraction module. In one channel, a spatio-temporal attention mechanism is used to extract spatio-temporal features of the normalized data. In the other channel, the fast Fourier transform is used to transform the original top-layer oil temperature data from the time domain to the frequency domain, and a convolutional neural network is used to extract frequency domain features; Use a cross-attention mechanism to fuse spatio-temporal features and frequency domain features; Design a multi-scale temporal convolutional network model based on multi-scale dilated causal convolution and deep learning algorithms to capture multi-scale features from both global and local perspectives; Use a fully connected layer and non-normalized output to predict the result.

2. The method for predicting the top oil temperature of the oil-containing equipment according to claim 1, wherein, The spatio-temporal attention mechanism is composed of a time attention unit and a spatial attention mechanism in series. The time attention unit is used to adaptively assign attention weights to all time points to extract temporal features; the spatial attention mechanism is used to adaptively assign attention weights to input variables to extract spatial features.

3. The oil-containing equipment top layer oil temperature prediction method according to claim 2, characterized in that The time attention unit is designed based on a bidirectional gated recurrent unit.

4. The method for predicting the top oil temperature of the oil-containing equipment according to claim 1, wherein The formula for using a convolutional neural network to extract frequency domain features is as follows: Among them, X(t) represents the normalized data, is the frequency-domain signal after fast Fourier transform, and w is the angular frequency.

5. The method for predicting the top oil temperature of the oil-containing equipment according to claim 1, wherein The calculation formula for using a cross-attention mechanism to fuse spatio-temporal features and frequency domain features is as follows: F S = Softmax((W q F x )(W k h x )) T )W v h x Among them, W k and W v are the linear transformation matrices of the key and the value respectively, W q is the linear transformation matrix of the query, F x is the extracted frequency-domain feature, h x is the extracted spatio-temporal feature, F S is the fused feature, and softmax() is the normalization activation function.

6. The method for predicting the top oil temperature of the oil-containing equipment according to claim 1, wherein The multi-scale temporal convolutional network model is composed of multi-scale dilated causal convolution, self-attention, and a fully connected layer. After obtaining features of different scales through multi-scale dilated causal convolution, deep learning algorithms are used to capture global and local connections.

7. The method for predicting the top oil temperature of the oil-containing equipment according to claim 6, characterized in that, The calculation formula of the multi-scale temporal convolutional network model is as follows: Among them, is the DCC result of the i-th channel, f is the convolution kernel size, F S is the input fusion feature, T is the time length of the input data, X c is the output of the multi-scale dilated causal convolution, W Q , W K , W V is the linear transformation matrix of the triple, is the scaling factor, H is the output of the deep learning algorithm, FC() is the calculation of the fully connected layer, is the prediction result.

8. The method for predicting the top oil temperature of the oil-containing equipment according to claim 1, wherein The formula for non-normalization is as follows: Among them, x max is the maximum value in the oil temperature data, and x min is the minimum value in the oil temperature data. and are the predicted values of the oil temperature before and after inverse normalization, respectively.

9. An electronic device, comprising a memory and a processor, characterized in that The memory is used to store a program that supports the processor to execute the top-layer oil temperature prediction method for the oil-containing equipment described in any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A storage medium has a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the top-layer oil temperature prediction method for the oil-containing equipment described in any one of claims 1 to 8.