Gearbox oil temperature prediction method based on hybrid deep learning model

By adopting a hybrid deep learning model in gearbox oil temperature prediction, combining TTAO algorithm, hierarchical clustering method and phase space reconstruction technology, the chaotic characteristics of gearbox are identified and processed, and the problem of insufficient prediction accuracy in the existing technology is solved, achieving more efficient and accurate oil temperature prediction.

CN120030476APending Publication Date: 2025-05-23BEIFANG UNIV OF NATITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114054.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture its complex dynamic characteristics and chaotic behavior in the prediction of gearbox oil temperature, resulting in insufficient prediction accuracy.

Method used

Using a method based on a hybrid deep learning model, the parameters of VMD are optimized through the TTAO algorithm, combined with hierarchical clustering method and phase space reconstruction technology, the gearbox oil temperature data is decomposed and reconstructed, chaotic characteristics are identified, and short-term prediction is used using the GRU-Informer hybrid model.

Benefits of technology

It improves the accuracy and stability of gearbox oil temperature prediction, reduces the calculation complexity and storage occupancy of the model, reflects the health status inside the wind turbine, and reduces the cost of equipment maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030476A_ABST
    Figure CN120030476A_ABST
Patent Text Reader

Abstract

The invention discloses a gearbox oil temperature prediction method based on a hybrid deep learning model, and belongs to the technical field of fault prediction, and the method comprises the steps: optimizing two regulatory factors of variational mode decomposition (VMD) through employing a TTAO algorithm, and carrying out the decomposition of oil temperature data through employing a TTAO-VMD model; performing sequence reconstruction according to the fuzzy entropy values of different frequency components after decomposition; according to the Lyapunov exponent of the reconstructed sequence, the chaos characteristic is identified, and lagging order selection is carried out on the chaos sequence and the non-chaos sequence; and carrying out short-term prediction by using a deep learning model GRU-Informer. According to the method, short-term prediction is carried out on the oil temperature of the gearbox through sequence reconstruction and chaos recognition, and the prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault prediction, and in particular to a gearbox oil temperature prediction method based on a hybrid deep learning model. Background Art

[0002] At present, the research on time series prediction inside wind turbines mainly focuses on physical models, statistical models, artificial intelligence models and hybrid models. Physical models include heat network models and stochastic fluid dynamics models. However, such models need to establish and solve relevant partial differential equations, and the data requirements are large, which brings challenges to the prediction accuracy. Statistical models mainly include autoregressive integrated moving average model (ARIMA), autoregressive moving average model (ARMA) and generalized autoregressive conditional heteroskedasticity model (GARCH). Statistical models overcome the shortcomings of physical modeling to a certain extent due to their relatively simple modeling, but ignore the nonlinear characteristics of gearbox internal data. In recent years, the excellent performance of artificial intelligence technology has been discovered by more and more scholars and applied to the field of gearbox oil temperature prediction. Long short-term memory network (LSTM) has been used by many scholars for oil temperature prediction. With the increasing application of Transformer in various fields, and because of its own self-attention mechanism, it performs better than LSTM in time series prediction tasks, Transformer is increasingly used, but Transformer needs to process each time step, which makes the model computationally complex and storage occupancy high. Informer introduces a probabilistic sparse attention mechanism based on Transformer to reduce the temporal and spatial complexity and improve prediction efficiency. In addition, the decoder can complete the output of long sequences in a single forward calculation, thereby effectively avoiding the expansion of cumulative errors in the reasoning stage. Although the above methods perform well in prediction tasks, the prediction performance of a single model is always limited.

[0003] The emergence of hybrid models makes up for the shortcomings of single models. The hybrid model is a hybrid model that combines signal decomposition technology, intelligent optimization algorithms and multiple prediction models. It contains a decomposition algorithm. It clearly presents the relationship between the frequency domain and the time domain through time-frequency analysis, so that the model has obvious advantages in mining the intrinsic characteristics of complex data, thereby achieving more accurate predictions. In the prior art, variational mode decomposition VMD is often used for decomposition, but the problem of hyperparameter selection is not considered when processing by the VMD method. This reduces the decomposition quality due to the subjective setting of parameters when decomposing the original data, thereby affecting the accuracy of the prediction model. Combining intelligent optimization algorithms with signal decomposition technologies such as VMD can effectively solve this problem and improve the performance of the model.

[0004] By integrating the advantages of different algorithms, hybrid models can demonstrate strong adaptability when processing diverse data. However, faced with data with complex dynamic characteristics, traditional data processing methods often find it difficult to fully capture its inherent laws. In view of these characteristics, Beyaoui et al. analyzed the nonlinear dynamics of the wind turbine gearbox system and demonstrated the chaotic behavior inside the gearbox. However, no scholar has yet taken the chaotic behavior of the gearbox into account when modeling. Summary of the invention

[0005] The purpose of the present invention is to provide a gearbox oil temperature prediction method based on a hybrid deep learning model, which performs short-term prediction of the gearbox oil temperature through sequence reconstruction and chaos identification to improve the prediction accuracy.

[0006] To achieve the above object, the present invention provides a gearbox oil temperature prediction method based on a hybrid deep learning model, the steps comprising:

[0007] S1. Use TTAO algorithm to optimize the two adjustment factors of variable mode decomposition VMD, and use TTAO-VMD model to decompose the gearbox oil temperature data;

[0008] S2. Reconstruct the gearbox oil temperature time series using a hierarchical clustering method based on the fuzzy entropy values ​​of different frequency components after decomposition;

[0009] S3, identify the chaotic characteristics of the reconstructed sequence according to the Lyapunov index, and use phase space reconstruction and partial autocorrelation function to select the lag order for chaotic and non-chaotic sequences respectively;

[0010] S4. Use the GRU-Informer hybrid deep learning model for short-term prediction of gearbox oil temperature.

[0011] Preferably, in step S1, VMD uses a non-recursive method to decompose the original signal of the gearbox oil temperature, including:

[0012] Establish the constraint equation, the formula is:

[0013]

[0014] In the formula, * represents the convolution calculation symbol, x(t) represents the original data, ω k and u K They are respectively represented as the center frequency and frequency band of the Kth intrinsic mode function, represents the partial derivative in time, δ(t) represents the unit impulse function, j is the imaginary unit, and t represents time;

[0015] The penalty parameter α and the Lagrange multiplier λ(t) are introduced to solve the constraint equation. The formula is:

[0016]

[0017] In the formula, L represents the objective function, f(t) represents the input signal, is a complex exponential signal, indicating the spectrum deviation of the modulation, u k (t) represents the kth mode function, ω k represents the center frequency of the kth mode;

[0018] The optimal solution is obtained using the alternating direction multiplier method, the formula is:

[0019]

[0020] In the formula, represents the spectrum of K modes in n+1 iterations, represents the spectrum of the signal x(t), represents the spectrum of the jth mode, represents the frequency-dependent Lagrange multiplier.

[0021] Preferably, the TTAO algorithm is used to optimize the penalty parameters and the number of modes in the VMD model, and the envelope entropy is used as the fitness function. The optimization function is:

[0022]

[0023] In the formula, α and K represent the penalty factor and the number of modes respectively, λ 1 , 2 represents the trade-off parameter in the optimization process, E k (t) represents the energy distribution of the kth mode, represents the positive part of the variable, represents the objective function to be optimized.

[0024] Preferably, the fuzzy entropy data of different frequency components in step S2 are reconstructed into the gearbox oil temperature time series using a hierarchical clustering method. The fuzzy entropy is used to measure the time complexity of the sequence. For a time series U={u 1 ,u 2 ,…,u N}, starting from the first element, select m consecutive u values ​​to form a sequence segment U with a total of N-m+1 k = {u k ,u k+1 ,...,u k+m-1}, and then construct a new time series, expressed as:

[0025] U_new (k) =U k -u 0 (k);

[0026]

[0027] Where: U_new (k) is the new local sequence segment, u 0 (k) is a sequence of m consecutive u k The mean of , l represents the index variable of the summation process, and its value range is from 0 to 1;

[0028] Calculate any two m-dimensional vectors U k and U l The maximum absolute distance between The formula is:

[0029]

[0030] In the formula, u 0 (l) represents the smoothed signal value, u l+p-1 Represents the value of signal U at position l+p-1;

[0031] Introducing fuzzy membership function The formula is:

[0032]

[0033] In the formula, r represents tolerance;

[0034] Calculate the average value for each k and define the relationship dimension function φ under the function m dimension m (r), the formula is:

[0035]

[0036] In the formula, represents a calculated statistic, Nm represents the signal length N minus a window size m;

[0037] Increase the dimension, increase the dimension to m+1, repeat the above calculation steps, and get the relationship dimension function φ under m+1 dimension m+1 (r), the formula is:

[0038]

[0039] According to the above steps, the fuzzy entropy of the original sequence is obtained, and the formula is:

[0040] F E (m,r,N)=lnφ m (r)-lnφ m+1 (r).

[0041] Preferably, in step S3, selecting the lag order of the chaotic sequence using phase space reconstruction includes:

[0042] The phase space reconstruction condition is set as: d>2D+1, where d is the embedding dimension and D is the system correlation dimension;

[0043] The time delay and embedding dimension are estimated using the correlation integral. Considering the time delay τ and the embedding dimension d, the correlation of the time series is obtained by using the correlation integral analysis of the embedded time series, and the statistic is obtained. S cor (τ) and according to S cor (τ), and τ to obtain the optimal delay time τ d and the embedding window τ w , and finally find the embedding dimension d.

[0044] Preferably, in step S3, the lag order of the non-chaotic sequence is selected according to the partial autocorrelation function and the Akaike information criterion.

[0045] Preferably, step S4 specifically includes:

[0046] Input the data obtained in step S3 into the GRU-Informer hybrid deep learning model;

[0047] The GRU model extracts short-term dependencies of time series and encodes dynamic information within the time window as a hidden state sequence;

[0048] The encoding layer of Informer receives the hidden state sequence of GRU and generates intermediate results using the probabilistic sparse attention and distillation modules. The decoding layer uses these intermediate results and the hidden states extracted by GRU to perform multi-head attention operations.

[0049] Finally, the output dimension is adjusted through the fully connected layer to obtain the prediction result, and the reconstructed sequence components are predicted and superimposed to obtain the final oil temperature prediction result.

[0050] Therefore, the present invention adopts the above-mentioned gearbox oil temperature prediction method based on the hybrid deep learning model, which has the following beneficial effects:

[0051] (1) The TTAO algorithm is used to solve the optimal parameters of VMD penalty factor and mode number. The introduction of the optimization algorithm minimizes the subjectivity of VMD parameter selection and improves the quality of data decomposition.

[0052] (2) The oil temperature is predicted by sequence reconstruction and chaos identification. For the subsequence, the hierarchical clustering method is used to reconstruct it according to the fuzzy entropy value to reduce the computational complexity and prediction error. Then, the chaotic and non-chaotic oil temperature reconstruction components are separated according to the maximum Lyapunov exponent of the reconstructed sequence. The phase space reconstruction (PSR) and partial autocorrelation function (PACF) are used to determine the lag order. The prediction performance of the model is improved through more comprehensive data processing.

[0053] (3) The GRU-Informer hybrid deep learning model is used to process subsequences. The GRU model has strong short-term feature extraction capabilities. When combined with the Informer model that is good at capturing long-term trend features, the advantages of both can be fully utilized, and the model has stronger short-term prediction capabilities.

[0054] (4) The method of the present invention reflects the internal health status of the wind turbine, thereby reducing equipment maintenance costs, improving the stability of the wind turbine gearbox operation and ensuring the operating efficiency of the wind farm.

[0055] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flow chart of the method of the present invention;

[0057] Figure 2 This is a structural diagram of the GRU-Informer hybrid deep learning model of the present invention;

[0058] Figure 3 This is the gearbox oil temperature sequence diagram of the present invention;

[0059] Figure 4 It is a TTAO optimization curve diagram based on oil temperature data of the present invention;

[0060] Figure 5 This is a TVMD decomposition flow chart based on oil temperature data of the present invention;

[0061] Figure 6 It is a hierarchical clustering diagram of IMFs based on oil temperature data decomposition of the present invention;

[0062] Figure 7 It is a lag order selection diagram of the reconstructed sequence SEQ2 based on oil temperature data of the present invention;

[0063] Figure 8 It is a lag order selection diagram of the reconstruction sequence SEQ3 based on oil temperature data of the present invention;

[0064] Fig. 9 It is a lag order selection diagram of the reconstruction sequence SEQ4 based on oil temperature data of the present invention;

[0065] Fig.10 It is a lag order selection diagram of the reconstruction sequence SEQ5 based on oil temperature data of the present invention;

[0066] Fig.11 This is a prediction graph of the deep learning model based on oil temperature data of the present invention;

[0067] Fig.12 This is a comparison chart of ablation experiment prediction based on oil temperature data of the present invention;

[0068] Fig.13 This is a single model prediction comparison chart based on oil temperature data of the present invention;

[0069] Fig.14 This is a comparison chart of the hybrid model prediction based on oil temperature data of the present invention. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0071] Example

[0072] Reference Figure 1 The present invention provides a gearbox oil temperature prediction method based on a hybrid deep learning model, the steps comprising:

[0073] S1. Use the TTAO algorithm to optimize the two adjustment factors of variable mode decomposition VMD, and use the TTAO-VMD model to decompose the gearbox oil temperature data. The decomposed data is represented as IMF1, IMF2, ..., IMFn.

[0074] TTAO algorithm (Triangle Topology Aggregation Optimization Algorithm) As the iteration proceeds, new vertices are continuously generated in the search space and used to construct similar triangles of different sizes. In the TTAO algorithm, each triangle is regarded as a basic evolutionary unit composed of four agents, namely the three vertices of the triangle and one randomly generated vertex inside. More importantly, the aggregation core of the algorithm lies in grouping vertices with superior characteristics. The TTAO algorithm collects high-quality information vertices between different topological units or within the unit through aggregation. The algorithm mainly relies on two stages of generic aggregation and local aggregation for parameter optimization. The update process is as follows:

[0075] 1) Generic aggregation stage

[0076] Generic aggregation emphasizes the exploration phase, by collecting information about excellent individuals in different triangular units and creating new feasible solutions. Information interaction occurs between the best individual in each triangular topological unit and the best individual in any group of randomly selected units. There is a linear combination of different weights between each dimension variable of two positive individuals. New individuals are generated from better connections between two vertices, which can be expressed mathematically as:

[0077]

[0078] In the formula, r 4 is a random number between [0,1], and represents the best individual of unit i and the randomly selected unit at the tth iteration. In addition, the greedy strategy is used to update the optimal agent, and the mathematical expression is:

[0079]

[0080] In the formula, represents the suboptimal individual at the i-th iteration, and f(·) is the function of the given problem.

[0081] 2) Local aggregation

[0082] Local aggregation mainly emphasizes the development stage. In this stage, the triangular topological units are aggregated internally. A triangular topological structure is temporarily formed between the updated optimal or suboptimal individual and the two vertices with good fitness values ​​in the group. In this case, the formed topological structure is not necessarily an equilateral triangle. Based on the motion vector difference between the optimal and suboptimal individuals, the position of the optimal individual is disturbed (including direction and step length) in the local area. Therefore, the development of each topological triangular unit is achieved by re-searching each group in a certain local area. The calculation formula for the new vertex is:

[0083]

[0084] In the formula, represents the new position of the individual in the t+1 generation at the i-th iteration, and α represents the acceleration factor, which can be calculated as:

[0085]

[0086] Where e is the natural logarithm, which is used to describe the growth or decay relationship, and T represents the number of iterations in the optimization algorithm.

[0087] The purpose of using suboptimal individual information is to prevent the optimal individual from falling into a local extreme value. After aggregation, the guide point of the temporary triangle unit is guaranteed to be optimal within the unit. In order to make the convergence move in a promising direction, the fitness values ​​of the two vertices before and after local mining are compared to determine the update of the position. If the new individual is better than the original individual, the position is updated, otherwise no update is performed. The corresponding mathematical expression is as follows:

[0088]

[0089] VMD uses a non-recursive method to decompose the original signal of gearbox oil temperature, including:

[0090] Establish the constraint equation, the formula is:

[0091]

[0092] In the formula, * represents the convolution calculation symbol, x(t) represents the original data, ω k and u K They are respectively represented as the center frequency and frequency band of the Kth intrinsic mode function, represents the partial derivative in time, δ(t) represents the unit impulse function, j is the imaginary unit, and t represents time;

[0093] The penalty parameter α and the Lagrange multiplier λ(t) are introduced to solve the constraint equation. The formula is:

[0094]

[0095] In the formula, L represents the objective function, f(t) represents the input signal, is a complex exponential signal, indicating the spectrum deviation of the modulation, u k (t) represents the kth mode function, ω k represents the center frequency of the kth mode;

[0096] The optimal solution is obtained using the alternating direction multiplier method, the formula is:

[0097]

[0098] In the formula, represents the spectrum of K modes in n+1 iterations, represents the spectrum of the signal x(t), represents the spectrum of the jth mode, represents the frequency-dependent Lagrange multiplier.

[0099] The TTAO algorithm is used to optimize the penalty parameters and the number of modes of the VMD model, and the envelope entropy is used as the fitness function. The specific optimization function is as follows:

[0100]

[0101] In the formula, α and K represent the penalty factor and the number of modes respectively, E k (t) represents the energy distribution of the kth mode, represents the positive part of the variable, represents the optimization objective function. The energy distribution of the kth mode is calculated to form the envelope entropy, which is used to measure the modal component u k (t) complexity and stationarity, λ 1 , 2 Represents the trade-off parameters in the optimization process. The introduction of the optimization algorithm minimizes the subjectivity of VMD parameter selection and improves the quality of data decomposition.

[0102] S2. Use the hierarchical clustering method to reconstruct the gearbox oil temperature time series according to the fuzzy entropy data of different frequency components after decomposition, and obtain the sequence SEQ1, SEQ2, ..., SEQn. Among them, the fuzzy entropy is used to measure the time complexity of the sequence. For the time series U={u 1 ,u 2 ,...,u N}, starting from the first element, select m consecutive u values ​​to form a sequence segment U with a total of N-m+1 k = {u k ,u k+1 ,...,u k+m-1}, and then construct a new time series, expressed as:

[0103] U_new (k) =U k -u 0 (k);

[0104]

[0105] Where: U_new (k) is the new local sequence segment, u 0 (k) is a sequence of m consecutive u k , l represents the index variable of the summation process, and its value range is from 0 to 1.

[0106] Calculate any two m-dimensional vectors U k and U l The maximum absolute distance between The formula is:

[0107]

[0108] In the formula, u 0 (l) represents the smoothed signal value, u l+p-1 Represents the value of signal U at position l+p-1.

[0109] Introducing fuzzy membership function The formula is:

[0110]

[0111] In the formula, r represents the tolerance.

[0112] Calculate the average value for each k and define the relationship dimension function φ under the function m dimension m (r), the formula is:

[0113]

[0114] In the formula, Represents a calculated statistic, and Nm represents the signal length N minus a window size m.

[0115] Increase the dimension, increase the dimension to m+1, repeat the above calculation steps, and get the relationship dimension function φ under m+1 dimension m+1 (r), the formula is:

[0116]

[0117] In summary, the fuzzy entropy of the original sequence can be obtained:

[0118]

[0119] N is generally a finite value, then the fuzzy entropy formula is:

[0120] F E (m,r,N)=lnφ m (r)-lnφ m+1 (r).

[0121] S3. The chaotic characteristics of the reconstructed sequence are identified according to the Lyapunov exponent, and the lag order is selected for the chaotic sequence and the non-chaotic sequence respectively using phase space reconstruction and partial autocorrelation function.

[0122] Specifically, the maximum Lyapunov exponent of the reconstructed sequence is calculated and its chaotic characteristics are identified. For chaotic sequences, phase space reconstruction is used to select the lag order, including:

[0123] The phase space reconstruction condition is set to: d>2D+1, where d is the embedding dimension and D is the system correlation dimension. The phase space of the time series data is reconstructed, and the time delay and embedding dimension are estimated using the correlation integral. Considering the time delay τ and the embedding dimension d, the correlation integral of the embedded time series is used to analyze the correlation of the time series and obtain the statistic S cor (τ) and according to S cor (τ), and τ to obtain the optimal delay time τ d and the embedding window τ w , and finally find the embedding dimension d.

[0124] For non-chaotic sequences, the lag order is selected based on the partial autocorrelation function and the Akaike information criterion.

[0125] S4. Use the GRU-Informer hybrid deep learning model for short-term prediction of gearbox oil temperature.

[0126] GRU (Gated Recurrent Unit) effectively extracts local features of time series through a gating mechanism, including update gate, reset gate, and hidden state.

[0127] Update gate z t : Determines how much of the state of the current time step comes from past information and how much comes from the current input, allowing the network to selectively retain variable information. The formula is:

[0128] z t =σ(W xz x t +W hz h t-1 +b z ).

[0129] Reset Gate t : Determines how the current input is combined with past information. If the value is close to 0, it means that the network will discard the past hidden state and only use the current input for update. The formula is:

[0130] r t =σ(W xr x t +W hr h t-1 +b r ).

[0131] Hidden State Adjust the past state according to the reset gate and update the gate z t Under the control of , the new input is combined with the past hidden state to generate a new hidden state:

[0132]

[0133] Where W xr , W hr , W xz , W hz , W xc , W hc Represent different weight matrices, b r , b z , b c represents different bias vectors, σ(·) represents the sigmoid activation function, is a candidate hidden state, tanh(·) is a hyperbolic sine activation function, represents the Hadamard product, c t 、h t 、x t They represent the cell state, hidden state, and input at time step t respectively.

[0134] The Informer model uses the probabilistic sparse self-attention mechanism to improve operating efficiency and prediction effect. The calculation formula of the probabilistic sparse self-attention mechanism is:

[0135]

[0136] In the formula, is the matrix obtained by probabilistic sparseness of the original query matrix Q, and Softmax is the normalized activation matrix, K, V, D k , K T denote the key matrix, value matrix, dimension of the key matrix, and transpose of the key matrix, respectively.

[0137] The distillation layer is introduced in the encoding stage to remove irrelevant information and generate concentrated self-attention feature maps in the lower layers. The distillation process is:

[0138]

[0139] in,[*] AB The basic operations include the key operations in the probabilistic sparse self-attention mechanism. ELU represents an activation function, Conv1d represents a one-dimensional convolution, and MaxPool represents the maximum pooling operation. Represents a state or representation of the input data at time step t.

[0140] The input vector received by the decoder is represented as:

[0141]

[0142] In the formula, represents the processed input data at time step t, Represents the token-related representation in the input data, usually associated with each token in the sequence. Represents the initial input data at time step t, usually the input features at the beginning of a certain stage. token , L v ,d model They respectively represent the length of the token, the length of a variable (which may be the sequence length or other), and the dimension of the model. A token can refer to a word, character, subword, or other smaller units, depending on how the data is processed and the word segmentation method used.

[0143] The structure diagram of the GRU-Informer hybrid deep learning model is as follows Figure 2 As shown in the figure, the advantages of short-term feature extraction and long sequence modeling of GRU and Informer models are integrated, so that the model provides better short-term prediction performance while maintaining efficient calculation. The steps of using the GRU-Informer hybrid deep learning model include:

[0144] First, the data obtained in step S3 is input into the GRU-Informer hybrid deep learning model;

[0145] Secondly, the GRU model extracts the short-term dependencies of the time series and encodes the dynamic information within the time window into a hidden state sequence;

[0146] Subsequently, the encoding layer of Informer receives the hidden state sequence of GRU and generates intermediate results using the probabilistic sparse attention and distillation modules. The decoding layer performs multi-head attention operations using these intermediate results and the hidden states extracted by GRU.

[0147] Finally, the output dimension is adjusted through the fully connected layer to obtain the prediction result, and the reconstructed sequence components are predicted and superimposed to obtain the final oil temperature prediction result.

[0148] The proposed method is verified using wind turbine gearbox oil temperature data collected by SCADA from a wind farm. After removing null values ​​and extreme outliers in the data, 3000 experimental data are selected and divided into training set and test set in a ratio of 8:2. Table 1 describes the basic statistical information of the obtained oil temperature data. The original oil temperature sequence data is as follows Figure 3 As shown, before the hybrid model prediction, its parameters are set, and the settings are shown in Table 2:

[0149] Table 1 Basic statistical information of gearbox oil temperature data

[0150]

[0151] Table 2 Parameter setting table

[0152]

[0153]

[0154] In the process of optimizing VMD parameters using the TTAO algorithm, the population size is set to 30, the maximum number of iterations is set to 20, the penalty factor is set between [100, 2500], and the envelope entropy is used as the fitness function. When the number of iterations is 4, the fitness function reaches the minimum value of 7.59676425. At this time, the penalty parameter is 2487 and the number of modes is 9. The specific process of optimization iteration is as follows Figure 4 Then, TVMD is used to decompose the original oil temperature data, and the frequency of the components is arranged from low to high. The decomposition results are shown in Figure 5 As shown. Since the prediction result is the superposition of the prediction value of each component, direct prediction will increase the computational complexity, thus affecting the prediction accuracy of the model. In order to reduce the prediction error and workload, the data is reconstructed by calculating the fuzzy entropy value of each IMF. The specific values ​​of the fuzzy entropy of each modal component are shown in Table 3:

[0155] Table 3 Fuzzy entropy of each component of oil temperature

[0156] IMFs IMF1 IMF2 IMF3 IMF4 IMF5 IMF6 IMF7 IMF8 IMF9 FE 0.029 0.131 0.276 0.325 0.331 0.331 0.272 0.104 0.184

[0157] In order to merge the sequences with similar fuzzy entropy values ​​in a reasonable way, the hierarchical clustering method is used to reconstruct the subsequences with similar complexity, and 0.04 is used as the threshold line, as shown in Figure 6 As shown. IMF1 is named SEQ1 as a reconstruction component. IMF2 and IMF8 are merged and named SEQ2 because of their similar sequence complexity. Similarly, IMF3 and IMF7 are merged and named SEQ3, and IMF4, IMF5 and IMF6 are merged and named SEQ4. Finally, IMF9 can be used as a reconstruction component SEQ5 alone. The reconstruction results are shown in Table 4:

[0158] Table 4 Reconstructed sequence list

[0159]

[0160]

[0161] After reconstruction, the CC algorithm is used to calculate the time delay τ and embedding dimension d of the subsequence. Finally, the Rosenstein algorithm in the small data set is used to calculate the maximum Lyapunov exponent of each subsequence. The specific calculation results are shown in Table 5:

[0162] Table 5 Chaotic characteristics identification table

[0163] sequence Time Delay Embedding Dimension Maximum Lyapunov exponent SEQ1 18 10 0.0145 SEQ2 7 24 -0.0060 SEQ3 3 20 -0.0057 SEQ4 17 6 -0.0087 SEQ5 64 3 -0.0007

[0164] In the above table, SEQ1 is a chaotic sequence with significant nonlinear characteristics. It is difficult for the model to capture its complex behavior. Reconstructing the phase space of SEQ1 based on time delay and embedding dimension can show the nonlinear characteristics of the subsequence. For non-chaotic subsequences, the lag order is selected based on the PACF diagram and the Akaike Information Criterion (AIC). Figures 7-10 The PACF information and AIC values ​​of four non-chaotic sequences are displayed. The lagged historical data highly correlated with the original subsequences are selected to capture the real time dependency of the sequence and improve the prediction effect of the model. However, in order to avoid the increase of modeling complexity and multicollinearity problems caused by too many lagged items, the AIC and the number of lagged items are considered. Finally, Figures 10 to 13 The information displayed,determines the lag order of SEQ2, SEQ3, SEQ4, and SEQ5 as 6th, 7th, 4th, and 6th order, respectively, and uses the corresponding lag terms as influencing variables.,After the internal reconstruction of the data is completed, it is input into the GRU-Informer model,,which undergoes short-term feature extraction and long-term dependency modeling by the hybrid model. Fig.11 The final prediction results of the combined model are shown. The predicted values ​​basically coincide with the actual values, which well fits the nonlinearity in the original oil temperature data.

[0165] In order to prove the effectiveness of parameter optimization and data processing used in the present invention, an ablation experiment was conducted. Among them, VMD-FE-PSR-GRU-Informer represents a model that has not been optimized by the TTAO algorithm module, VMD-PSR-GRU-Informer represents a model that has not been optimized by the TTAO algorithm module and the FE reconstruction module, PSR-GRU-Informer represents a model that has not been optimized by the VMD decomposition module, the TTAO algorithm optimization module, and the FE reconstruction module, GRU-Informer represents a model that has not been optimized by the VMD decomposition module, the TTAO algorithm optimization module, the FE reconstruction module, and the chaotic data processing PSR, and Proposed represents the model of the present invention. Fig.12The comparison chart of the predicted curve and the actual curve on the test set after eliminating different modules shows that if the model is used directly to predict the original data without any data processing, it cannot accurately fit the true value and can only roughly simulate the trend of the original data. With the addition of PSR module, VMD decomposition module, FE reconstruction module and TTAO optimization module one by one, the predicted curve gradually converges to the original data. Fig.12 Both Table 6 and Table 7 prove that the mixed strategy can further improve the prediction accuracy of the model compared with the single strategy. This shows that there are different components in the data. Only after the components are identified and processed in a targeted manner can better prediction results be achieved.

[0166] Table 6 Comparison of ablation experiments

[0167]

[0168]

[0169] It can be seen from Table 6 that compared with VMD-FE-PSR-GRU-Informer without algorithm optimization, the RMSE and MAE of the proposed model are reduced by 40.45% and 43.88% respectively, while R 2 The combined model proposed in the present invention has an improvement of 2%. 2 , RMSE and MAE show the best results, reaching 0.9899, ​​0.8915 and 0.6446 respectively. Compared with the direct prediction of the original data, R 2 The results show that the model of the invention processes the oil temperature data more carefully and fully exploits the data characteristics, thus significantly improving the prediction accuracy and performance.

[0170] In order to verify the effectiveness of the model of the present invention, the model of the present invention is compared with a single model. The hybrid model used is compared with the single model KELM, BP, LSTM, GRU, Transformer and Informer models. The experiment proves that the prediction effect of the hybrid model of the present invention is better than other single models. Fig.13 As shown in the figure, a single model cannot capture the nonlinear characteristics of oil temperature data very well. Although the Informer model can fit the general trend of the data, the predicted values ​​of the data at some times are quite different from the original data, indicating that there are components in the data that cannot be directly predicted. Only by separating the components of the original data and passing them into the prediction model through different data processing methods can the prediction effect of the model be improved.

[0171] Table 7 Comparative experiments of single prediction models

[0172]

[0173]

[0174] Table 7 compares the mainstream single model and hybrid model to show their competitiveness. It is concluded that the machine learning model KELM is slightly better than BP neural network and LSTM, but lower than GRU, Transformer and Informer. The Informer model has better prediction performance in the single model comparison because it improves the information extraction mechanism based on the Transformer model, but there is still a gap compared with the hybrid model proposed in this invention. Compared with Informer, the R 2 The improvement was 8.68%, and the RMSE and MAE were reduced by 66.34% and 66.37% respectively. In the benchmark experiment of single model comparison, it was proved that the proposed hybrid model has the best prediction performance.

[0175] In order to further verify the effectiveness of the combined module TVMD-FE-PSR and the prediction model for the original oil temperature processing, the oil temperature data was processed by the same combined data module TVMD-FE-PSR and then nonlinearly fitted using different models. In addition to the proposed model, it also includes the LSTM-Informer model, Transformer model, GRU model and Informer model. Fig.14 Compared with the prediction results of the model without data processing, it is proved that after the data processing of the combined module, the prediction effect of the model is better and closer to the true value.

[0176] Table 8 Hybrid model comparison experiment

[0177] Model <![CDATA[R 2 ]]> RMSE MAE Proposed 0.9899 0.8915 0.6446 TVMD-FE-PSR-LSTM-Informer 0.9862 1.2905 0.8524 TVMD-FE-PSR-Informer 0.9843 1.0911 0.8196 TVMD-FE-PSR-Transformer 0.9793 1.4279 1.1749 TVMD-FE-PSR-GRU 0.9692 1.4786 1.1937

[0178] Table 8 compares the prediction effects of each hybrid model numerically. Combined with Table 7, it can be found that after the same data processing, GRU-Informer has a higher R than LSTM-Informer. 2 The prediction effect of the GRU-Informer hybrid model is better than that of the LSTM-Informer model because the structure of the GRU is simpler and the ability to remember information is stronger, which can capture key local features. In addition, after data processing, the prediction effect of the model is generally improved. Before and after data processing, the R of the Transformer model is 2 The improvement was 10.33%, and the RMSE and MSE were reduced by 51.97% and 51.31% respectively. Therefore, the hybrid model proposed in the present invention is superior to other hybrid models in gearbox oil temperature prediction.

[0179] In summary, there are difficult-to-capture non-linear components in the original oil temperature data, and the direct prediction effect is worse than that after data processing. The results of the ablation experiment prove the necessity of adding each module. Compared with directly using the model for prediction, the data processing by the combined module makes the R 2 increase by 5.31%, and the RMSE and MAE are reduced by 58.98% and 59.14% respectively. The hybrid model can accurately fit the sharp rise and fall of the oil temperature data. Compared with the currently popular benchmark models, it has higher prediction accuracy and stronger generalization ability. In the comparative experiment of the hybrid model with the same data processing, it further proves that the hybrid model proposed in the present invention has better stability and superior performance.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A gearbox oil temperature prediction method based on a hybrid deep learning model, characterized in that the steps include: S1. Use TTAO algorithm to optimize the two adjustment factors of variable mode decomposition VMD, and use TTAO-VMD model to decompose the gearbox oil temperature data; S2. Reconstruct the gearbox oil temperature time series using a hierarchical clustering method based on the fuzzy entropy values ​​of different frequency components after decomposition; S3, identify the chaotic characteristics of the reconstructed sequence according to the Lyapunov index, and use phase space reconstruction and partial autocorrelation function to select the lag order for chaotic and non-chaotic sequences respectively; S4. Use the GRU-Informer hybrid deep learning model for short-term prediction of gearbox oil temperature.

2. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 1 is characterized in that: In step S1, VMD uses a non-recursive method to decompose the original signal of the gearbox oil temperature, including: Establish the constraint equation, the formula is: In the formula, * represents the convolution calculation symbol, x(t) represents the original data, ω k and u K They are respectively represented as the center frequency and frequency band of the Kth intrinsic mode function, represents the partial derivative in time, δ(t) represents the unit impulse function, j is the imaginary unit, and t represents time; The penalty parameter α and the Lagrange multiplier λ(t) are introduced to solve the constraint equation. The formula is: In the formula, L represents the objective function, f(t) represents the input signal, is a complex exponential signal, indicating the spectrum deviation of the modulation, u k (t) represents the kth mode function, ω k represents the center frequency of the kth mode; The optimal solution is obtained using the alternating direction multiplier method, the formula is: In the formula, represents the spectrum of K modes in n+1 iterations, represents the spectrum of the signal x(t), represents the spectrum of the jth mode, represents the frequency-dependent Lagrange multiplier.

3. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 2 is characterized in that: The TTAO algorithm is used to optimize the penalty parameters and the number of modes in the VMD model, with the envelope entropy as the fitness function. The optimization function is: In the formula, α and K represent the penalty factor and the number of modes respectively, λ1 and λ2 represent the trade-off parameters in the optimization process, and E k (t) represents the energy distribution of the kth mode, represents the positive part of the variable, represents the objective function to be optimized.

4. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 1, characterized in that: In step S2, the fuzzy entropy data of different frequency components are used to reconstruct the gearbox oil temperature time series using a hierarchical clustering method.

5. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 4 is characterized in that: In step S3, the lag order selection of the chaotic sequence using phase space reconstruction includes: The phase space reconstruction condition is set as: d>2D+1, where d is the embedding dimension and D is the system correlation dimension; The time delay and embedding dimension are estimated using the correlation integral. Considering the time delay τ and the embedding dimension d, the correlation of the time series is obtained by using the correlation integral analysis of the embedded time series, and the statistic is obtained. S cor (τ) and according to S cor (τ), and τ to obtain the optimal delay time τ d and the embedding window τ w , and finally find the embedding dimension d.

6. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 5 is characterized in that: In step S3, the lag order of the non-chaotic sequence is selected according to the partial autocorrelation function and the Akaike information criterion.

7. The gearbox oil temperature prediction method based on a hybrid deep learning model according to claim 6 is characterized in that: Step S4 specifically includes: Input the data obtained in step S3 into the GRU-Informer hybrid deep learning model; The GRU model extracts short-term dependencies of time series and encodes dynamic information within the time window as a hidden state sequence; The encoding layer of Informer receives the hidden state sequence of GRU and generates intermediate results using the probabilistic sparse attention and distillation modules. The decoding layer uses these intermediate results and the hidden states extracted by GRU to perform multi-head attention operations. Finally, the output dimension is adjusted through the fully connected layer to obtain the prediction result, and the reconstructed sequence components are predicted and superimposed to obtain the final oil temperature prediction result.