Residual life prediction method and system for rotary machinery
Through the Transformer-BiGRU neural network with multi-domain feature extraction and ReLU linear attention mechanism, the problem of sample scarcity and noise interference in rotating machinery is solved, high-precision residual life prediction is achieved, state maintenance strategy is supported, and maintenance costs and failure risks are reduced.
Patent Information
- Application Number
- CN202510400686.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-22
Smart Images

Figure CN120354068A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the remaining life prediction method for rotating machinery, and more specifically, to a remaining life prediction method and system for rotating machinery. Background Art
[0002] Currently, data-driven methods have become the mainstream technical direction in this field by extracting the operation trend characteristics of key components of rotating machinery and relying on neural networks for in-depth training to achieve remaining life prediction. Such methods can effectively quantify the equipment degradation process, overcome the limitations of traditional mechanism models in modeling complex physical laws, and show significant advantages in industrial predictive maintenance. Especially after the introduction of the Transformer model, its powerful long-time series modeling ability and parallel computing advantages have further promoted the improvement of the remaining life prediction accuracy.
[0003] However, the monitoring data of the core components of rotating machinery in industrial scenarios often have problems such as scarce sample size, complex and variable working conditions, and multi-source noise interference, resulting in the phenomenon that a single Transformer model is prone to overfitting or insufficient feature extraction, and it is difficult to construct a robust degradation trend characterization model under high complexity and small sample conditions. Therefore, how to break through the model generalization bottleneck under limited data conditions and improve the accuracy of the remaining life prediction of rotating machinery is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides a remaining life prediction method and system for rotating machinery, which overcomes the above defects.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A remaining life prediction method for rotating machinery, the specific steps are as follows:
[0007] Collect multi-source signals of key components of rotating machinery, and screen the multi-source signals to obtain original data;
[0008] Extract multi-domain features from the original data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features, and time-frequency domain features;
[0009] Screen the initial features based on a comprehensive index of weighted fusion of a monotonicity index and mutual information method to construct a feature subset;
[0010] Construct a Transformer-BiGRU neural network based on the ReLU linear attention mechanism, and use the feature subset to iteratively train the Transformer-BiGRU neural network to obtain a remaining life prediction model;
[0011] Obtain the data to be predicted, input the data to be predicted into the remaining useful life prediction model, and output the predicted value of the remaining useful life.
[0012] Optionally, the original data is a horizontal vibration signal.
[0013] Optionally, after the feature subset is constructed, each feature in the feature subset needs to be normalized.
[0014] Optionally, the expression of the monotonicity index is:
[0015]
[0016] In the formula, F represents the extracted feature; N represents the number of samples in the feature sequence; d / dF represents the difference of the feature sequence; No.(d / dF>0) represents the number of values greater than 0; No.(d / dF<0) represents the number of values less than 0.
[0017] Optionally, the expression of the mutual information method is:
[0018] I(F;t)=H(F)+H(t)-H(F,t);
[0019] In the formula, I(F;t) represents the mutual information relationship between the feature F and the time t; H(F) represents the information entropy of the feature F; H(t) represents the information entropy of the time t; H(F,t) represents the joint information entropy of the feature F and the time t.
[0020] Optionally, the expression of the ReLU linear attention mechanism is:
[0021]
[0022] In the formula, O i is the i-th row of the O matrix; Q i is the i-th query vector; K j is the j-th key vector; V j is the j-th value vector.
[0023] Optionally, the Transformer-BiGRU neural network includes a Transformer model and a BiGRU neural network; the BiGRU neural network is constructed by introducing an update gate and a reset gate on the basis of the recurrent neural network, and the output expressions of the update gate and the reset gate are respectively:
[0024] z t =σ(W z x t +U z h t-1 +bz )
[0025] r t = σ(W r x t + U r h t-1 + b r )
[0026] where z t represents the output value of the update gate at time t; σ represents the Sigmoid function; W z , U z both represent the weight matrices of the update gate; x t represents the input data; h t-1 represents the output of the gated recurrent unit at time t - 1; b z represents the bias vector of the update gate; r t represents the output value of the reset gate at time t; W r , U r both represent the weight matrices of the reset gate; br represents the bias vector of the reset gate.
[0027] A remaining useful life prediction system for a rotating machine, comprising:
[0028] A data acquisition module, configured to acquire multi-source signals of key components of the rotating machine, and screen the multi-source signals to obtain raw data;
[0029] A feature extraction module, configured to perform multi-domain feature extraction on the raw data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features, and time-frequency domain features;
[0030] A data screening module, which screens the initial features based on a comprehensive index obtained by weighted fusion of a monotonicity index and the mutual information method, and constructs a feature subset;
[0031] A model construction module, configured to construct a Transformer-BiGRU neural network based on the ReLU linear attention mechanism, and iteratively train the Transformer-BiGRU neural network using the feature subset to obtain a remaining useful life prediction model;
[0032] A life prediction module, configured to obtain data to be predicted, input the data to be predicted into the remaining useful life prediction model, and output a predicted value of the remaining useful life.
[0033] It can be seen from the above technical solutions that the present invention provides a method and system for predicting the remaining useful life of a rotating machine, which has the following beneficial effects compared with the prior art:
[0034] 1. By fusing the monotonicity of features with mutual information, it is possible to more accurately screen out the feature subset highly relevant to the equipment degradation process, significantly improving the sensitivity and characterization ability of features for life prediction, thereby providing a more reliable data basis for subsequent model training.
[0035] 2. By improving the Transformer model, replacing the traditional attention mechanism with the ReLU linear attention mechanism, it significantly reduces memory occupancy and training time while maintaining the ability to capture temporal features. Further combined with the parallel training of the BiGRU model, using its bidirectional gated structure to deeply model the temporal degradation pattern, fully integrating local feature dependencies and global degradation trends; while maintaining lightweight computing, it significantly improves the prediction accuracy.
[0036] 3. The present invention can accurately predict the remaining life of key components of rotating machinery (such as bearings, gears, etc.), and can be widely applied to fields such as aerospace engines, wind turbine generators, chemical compressors, coal mining machinery, and ship power systems. By real-time monitoring the health status of equipment and predicting the remaining life, it supports condition-based maintenance (CBM) strategies, avoids the deficiencies of traditional regular maintenance, reduces unnecessary shutdowns and component replacements, and reduces maintenance costs. At the same time, by early warning of potential failures, it can effectively prevent safety accidents caused by sudden equipment failures, ensuring production safety and the safety of personnel's lives and property, and having significant social and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0038] Figure 1 It is a schematic diagram of the method flow provided by the present invention;
[0039] Figure 2 It is a schematic diagram of the improved ReLU linear attention mechanism structure provided by the present invention;
[0040] Figure 3 It is a schematic diagram of the Transformer-BiGRU neural network structure provided by the present invention;
[0041] Figure 4(a) is a schematic diagram of the remaining life prediction result of the first rolling bearing in the test set provided by the present invention;
[0042] Figure 4(b) is a schematic diagram of the remaining life prediction result of the second rolling bearing in the test set provided by the present invention;
[0043] Figure 4(c) is a schematic diagram of the remaining life prediction result of the third rolling bearing provided by the present invention. Specific Embodiments
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] The embodiment of the present invention discloses a method for predicting the remaining life of a rotating machine, which can accurately predict the remaining life of key components of the rotating machine. Based on the prediction results, a maintenance plan is set, which can greatly reduce the maintenance cost and reduce the occurrence of catastrophic accidents. The steps are as Figure 1 shown, specifically:
[0046] Step 1: Collect multi-source signals of key components of the rotating machine, and screen the multi-source signals to obtain original data;
[0047] Step 2: Extract multi-domain features from the original data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features, and time-frequency domain features;
[0048] Step 3: Screen the initial features based on a comprehensive index of weighted fusion of the monotonicity index and the mutual information method to construct a feature subset;
[0049] Step 4: Construct a Transformer-BiGRU neural network based on the ReLU linear attention mechanism, and use the feature subset to iteratively train the Transformer-BiGRU neural network to obtain a remaining life prediction model;
[0050] Step 5: Obtain the data to be predicted, input the data to be predicted into the remaining life prediction model, and output the predicted value of the remaining service life.
[0051] In one embodiment, the original data is a horizontal vibration signal.
[0052] Furthermore, an acceleration sensor is used to collect the vibration signal of the key component of the rotating machine, an acoustic sensor is used to collect the acoustic signal, and a temperature sensor is used to collect the temperature signal. The vibration signal includes the vibration signal in the horizontal direction and the vibration signal in the vertical direction; the present invention selects the horizontal vibration signal of the rolling bearing as the original data.
[0053] Furthermore, in this embodiment, the rolling bearing accelerated life test data provided by the PRONOSTIA experimental platform is used to verify this method. This experiment is specifically designed for research on bearing fault diagnosis and prediction. By applying fixed loads and speeds to the bearings, experiments are conducted on 17 rolling bearings under three working conditions, with 7 rolling bearings in the first working condition and the second working condition respectively, and 3 rolling bearings in the third working condition. In this experiment, signals of the rolling bearings are collected by accelerometers fixed on the outer rings of the rolling bearings. When the acceleration of the vibration signal exceeds 20g, it is considered that the bearing has reached its final life and signal collection is stopped. Vibration signals of the rolling bearings are collected through vibration sensors. The vibration signals are collected at a sampling frequency of 25.6 kHz, with a sampling interval of 10 s each time and a sampling time of 0.1 s each time.
[0054] Horizontal vibration signals are selected as the data for the experiments of the present invention. Initial features are extracted from the horizontal vibration signals of the rolling bearings, including 25 features in total, including 12 time-domain features, 4 frequency-domain features, and 9 time-frequency domain features. Among them, the time-domain features include 8 dimensional features: root mean square, peak value, peak-to-peak value, standard deviation, variance, absolute average amplitude, kurtosis, and skewness; 4 dimensionless features: kurtosis factor, waveform factor, pulse factor, and peak factor. The frequency-domain features include average frequency, center frequency, mean square frequency, and frequency variance. The time-frequency domain features include the energy ratios of 8 sub-signals obtained through three-layer wavelet packet decomposition and the wavelet packet energy entropy calculated based on the total energy percentages of each sub-signal.
[0055] In one embodiment, after the feature subset is constructed, each feature in the feature subset needs to be normalized.
[0056] In one embodiment, the expression of the monotonicity index is:
[0057]
[0058] In the formula, F represents the extracted feature, N represents the number of samples in the feature sequence; d / dF represents the difference of the feature sequence; No.(d / dF>0) represents the number of values greater than 0; No.(d / dF<0) represents the number of values less than 0; the range of Mon(F) is [0, 1], and the closer it is to 1, the better the monotonicity of the feature.
[0059] In one embodiment, mutual information represents the amount of information about another random variable contained in a random variable, and the expression is:
[0060] I(F;t)=H(F)+H(t)-H(F,t);
[0061] Wherein, I(F; t) represents the mutual information relationship between feature F and time t; H(F) represents the information entropy of feature F; H(t) represents the information entropy of time t; H(F, t) represents the joint information entropy of feature F and time t.
[0062] Further, in order to select the best feature subset, this embodiment performs weighted fusion based on the above indexes to obtain a comprehensive index for evaluating the initial features. The formula is as follows:
[0063]
[0064] After comprehensive consideration, the features with a threshold greater than 0.6 are selected as the feature subset.
[0065] Subsequently, the [max, min] normalization method is used for the screened feature subset to make its data range between [0, 1]. Its expression is:
[0066]
[0067] Wherein, F max is the maximum eigenvalue; F min is the minimum eigenvalue; Fnor represents the normalized feature sequence.
[0068] In one embodiment, the training set and the test set are divided according to the feature subset, and the data is processed based on the sliding window technique.
[0069] Further, the specific steps for dividing the feature subset are as follows:
[0070] Constructing the training set: To ensure the effectiveness of the experiment, in this embodiment, the data of 6 key rotating machinery components are randomly selected as the training set in the first working condition and the second working condition, and the data of 2 key rotating machinery components are randomly selected as the training set in the third working condition. A total of 14 key rotating machinery component data are used as the training set. The expression of a single data in the training set is:
[0071] {(X m , Y n )}, X m = [x1, x2,..., x N T ; wherein, x = [F1, F2,..., F 10 , N is the number of samples, and m = [1,..., 14]. Y n is the life label of the nth bearing Y n = [y1, y2,..., y N T ; in this embodiment, the expression of the training set is: D Train = {(X1, Y1), (X2, Y2)...(X 14 ,Y 14 )}。
[0072] Set the lifespan label: For example, the total lifespan time length T of a bearing, that is, Y = y k / T, y k represents the current lifespan length.
[0073] Construct a test set: Select the feature subsets of the remaining key components of the rotating machinery as the test set {(X l )}, l = [15, 17], that is, D Test = {(X 15 ), (X 16 ), (X 17 )};
[0074] In this embodiment, the sliding window technique is used to process the data, and the sliding window size is set to 6 and the step size is 1.
[0075] In one embodiment, a Transformer model based on the ReLU linear attention mechanism is constructed: Among them, the improved ReLU linear attention mechanism is as Figure 2 shown, and its expression is:
[0076]
[0077] In the formula, Q = XW Q , K = XW K , V = XW V and W Q / W K / W V are learnable parameters, O i represents the i-th row of the O matrix, Sim(·,·) represents the similarity function, when the above equation is equivalent to the softmax attention; in the formula, d represents the vector dimension, and T represents the transpose of the key value.
[0078] In this embodiment, by using ReLU linear attention to replace the softmax therein, the expression of the similarity function is:
[0079] Sim(Q, K) = ReLU(Q)ReLU(K) T ;
[0080] That is, O i can be rewritten as:
[0081]
[0082] Compared with the previous softmax attention, this attention only needs to calculate and Part. Without changing the attention function, this method keeps the overall framework of Q, K, and V unchanged, replaces the previous dot product operation and softmax operation, and reduces the computational complexity and memory occupancy from quadratic to linear.
[0083] In one embodiment, a BiGRU model is constructed. The gated recurrent unit (GRU) introduces an update gate and a reset gate on the basis of the recurrent neural network (RNN) to control the retention and update of the hidden state information at the previous moment.
[0084] The update gate controls the current input x t and the previous memory information h t-1 , and outputs z t ; the reset gate controls the importance of h t-1 to the result h t . When the memory of h t-1 is completely irrelevant to the new memory, the reset gate removes the previous memory, and the calculation formula is as follows:
[0085] z t = σ(W z x t + U z h t-1 + b z );
[0086] r t = σ(W r x t + U r h t-1 + b r );
[0087] In the formula, z t represents the output value of the update gate at time t; σ represents the Sigmoid function; W z , U z both represent the weight matrix of the update gate; x t represents the input data; h t-1 is the output of the gated recurrent unit at time t-1; bz represents the bias vector of the update gate; r t represents the output value of the reset gate at time t; W r , U r both represent the weight matrix of the reset gate; b r represents the bias vector of the reset gate.
[0088] The candidate hidden state can be obtained according to the update gate The calculation formula is as follows:
[0089]
[0090] Finally, the output h is obtainedt :
[0091]
[0092] Wherein, W and U represent weight matrices, tanh is the hyperbolic tangent function, and b is the bias vector.
[0093] BiGRU considers the problem of backpropagation based on the GRU recurrent unit. By calculating the outputs of forward propagation and backpropagation and splicing them, the final output can be obtained.
[0094] The formulas for forward propagation and backpropagation are as follows:
[0095]
[0096] The formula for the final output is as follows:
[0097]
[0098] Wherein, GRU(·) represents the corresponding GRU hidden layer state; w t , v t respectively represent the weights corresponding to the forward hidden layer state and the output of the backward hidden layer state of the bidirectional GRU at time t, and b t represents the bias corresponding to the hidden layer state at time t.
[0099] In one embodiment, the Transformer model and the BiGRU neural network are in parallel, as Figure 3 shown. This model trains by simultaneously inputting the training set into the Transformer model and the BiGRU, and connecting the outputs of the Transformer model and the BiGRU through a splicing layer. Then, the final output is obtained through two fully connected layers and a regression layer. Inputting the test set data into the trained model, the predicted life value is obtained. Figures 4(a)-4(c) The predicted result for the bearings in the test set shows that as time increases, the degradation trend of the key components of the rotating machinery always fluctuates near the true life, indicating that the model provided by the present invention has good prediction effect.
[0100] In one embodiment, the root mean square error (RMSE) and the coefficient of determination (R 2 ) are used to evaluate the prediction performance of the model, and their calculation formulas are:
[0101]
[0102]
[0103] Wherein, ActRUL is the true degradation value of the key component data of the rotating machinery, and PreRUL is the predicted remaining life value. is the average value of the true life, and N is the number of data samples.
[0104] To verify the effectiveness of the model of the present invention, comparisons are made respectively based on BiGRU, Transformer and LSTM neural networks, and the results are shown in Table 1; the method provided in this embodiment predicts the remaining life based on RMSE and R 2 The evaluation index is smaller than the other three methods. RMSE is reduced by 18.98%, 62.01%, and 57.5% respectively, and R 2 is increased by 0.68%, 2.39%, and 1.86% respectively. It is proved that the model can well predict the remaining service life.
[0105] Table 1
[0106] model RMSE <![CDATA[R 2 > Transformer-BiGRU 0.018263 99.60% BiGRU 0.02254 98.93% Transformer 0.04808 97.27% LSTM 0.043003 97.78%
[0107] On the other hand, this embodiment also discloses a remaining life prediction system for a rotating machine, including:
[0108] A data acquisition module for acquiring multi-source signals of key components of the rotating machine and screening the multi-source signals to obtain raw data;
[0109] A feature extraction module for performing multi-domain feature extraction on the raw data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features and time-frequency domain features;
[0110] A data screening module for screening the initial features based on a comprehensive index of weighted fusion of a monotonicity index and a mutual information method to construct a feature subset;
[0111] A model construction module for constructing a Transformer-BiGRU neural network based on a ReLU linear attention mechanism and iteratively training the Transformer-BiGRU neural network using the feature subset to obtain a remaining life prediction model;
[0112] A life prediction module for obtaining data to be predicted, inputting the data to be predicted into the remaining life prediction model, and outputting a predicted value of the remaining service life.
[0113] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0114] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the remaining useful life of a rotating machine, characterized in that, The specific steps are as follows: Collect multi-source signals of key components of rotating machinery, and screen the multi-source signals to obtain original data; Extract multi-domain features from the original data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features, and time-frequency domain features; Screen the initial features based on a comprehensive index of weighted fusion of the monotonicity index and the mutual information method to construct a feature subset; Construct a Transformer-BiGRU neural network based on the ReLU linear attention mechanism, and iteratively train the Transformer-BiGRU neural network using the feature subset to obtain a remaining life prediction model; Obtain the data to be predicted, input the data to be predicted into the remaining life prediction model, and output the predicted value of the remaining service life.
2. The remaining life prediction method of a rotating machine according to claim 1, characterized in that, The original data is a horizontal vibration signal.
3. A remaining life prediction method for a rotating machine according to claim 1, characterized in that After the feature subset is constructed, each feature in the feature subset needs to be normalized.
4. A remaining life prediction method for a rotating machine according to claim 1, characterized in that, The expression of the monotonicity index is: In the formula, F represents the extracted feature; N represents the number of samples in the feature sequence; d / dF represents the difference of the feature sequence; No.(d / dF>0) represents the number of values greater than 0; No.(d / dF<0) represents the number of values less than 0.
5. A remaining life prediction method for a rotating machine according to claim 1, characterized in that, The expression of the mutual information method is: I(F;t)=H(F)+H(t)-H(F,t); In the formula, I(F;t) represents the mutual information relationship between the feature F and the time t; H(F) represents the information entropy of the feature F; H(t) represents the information entropy of the time t; H(F,t) represents the joint information entropy of the feature F and the time t.
6. A remaining life prediction method for a rotating machine according to claim 1, characterized in that, The expression of the ReLU linear attention mechanism is: Wherein, O i is the i-th row of the O matrix; Q i is the i-th query vector; K j is the j-th key vector; V j is the j-th value vector.
7. A method for predicting the remaining useful life of a rotating machine according to claim 1, characterized in that, The Transformer-BiGRU neural network includes a Transformer model and a BiGRU neural network; the BiGRU neural network is constructed by introducing an update gate and a reset gate on the basis of a recurrent neural network, and the output expressions of the update gate and the reset gate are respectively: z t = σ(W z x t + U z h t-1 + b z ); r t = σ(W r x t + U r h t-1 + b r ); where z t represents the output value of the update gate at time t; σ represents the Sigmoid function; W z , U z both represent the weight matrices of the update gate; x t represents the input data; h t-1 represents the output of the gated recurrent unit at time t - 1; b z represents the bias vector of the update gate; r t represents the output value of the reset gate at time t; W r , U r both represent the weight matrices of the reset gate; br represents the bias vector of the reset gate.
8. A remaining life prediction system for a rotating machine, characterized in that, Include: A data acquisition module, which is used to collect multi-source signals of key components of rotating machinery, and screen the multi-source signals to obtain original data; A feature extraction module, which is used to extract multi-domain features from the original data to obtain initial features; the multi-domain features include time-domain features, frequency-domain features, and time-frequency domain features; A data screening module, which screens the initial features based on a comprehensive index of weighted fusion of the monotonicity index and the mutual information method to construct a feature subset; A model construction module, which is used to construct a Transformer-BiGRU neural network based on the ReLU linear attention mechanism, and iteratively train the Transformer-BiGRU neural network using the feature subset to obtain a remaining life prediction model; A life prediction module, which is used to obtain the data to be predicted, input the data to be predicted into the remaining life prediction model, and output the predicted value of the remaining service life.