Method and system for predicting residual service life of gearbox based on multi-task semi-supervised learning

By adopting a multi-task semi-supervised learning method in the remaining service life prediction of gearboxes, combined with EMD-KPCA and HMM-Viterbi technology, the problem of not fully considering the gearbox degradation stage in the prior art is solved, and a prediction effect with higher accuracy and robustness is achieved.

CN119935550APending Publication Date: 2025-05-06XIAN UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202411717555.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art does not fully consider the impact of the gearbox degradation stage on the residual service life prediction, resulting in low model prediction accuracy.

Method used

Using a multi-task semi-supervised learning method, the health indicators based on EMD-KPCA are constructed through the acquisition and pre-processing of gearbox vibration signals, and the degradation phase is divided using the HMM-Viterbi algorithm. Combining the multi-task semi-supervised learning model, unsupervised pre-training and supervised training are used to improve the model's generalization ability of unknown data.

Benefits of technology

It improves the accuracy and robustness of the remaining service life prediction of the gearbox, reduces the need for a large amount of labeled data, reduces the cost of data preparation, and enhances the model's ability to identify abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119935550A_ABST
    Figure CN119935550A_ABST
Patent Text Reader

Abstract

The invention belongs to but is not limited to the technical field of gearbox residual service life prediction, and discloses a gearbox residual service life prediction method based on multi-task semi-supervised learning, which comprises the following steps of: firstly, decomposing and reconstructing a preprocessed vibration signal by using an empirical mode decomposition technology, extracting sensitive features, and calculating the residual service life of a gearbox; and then carrying out dimensionality reduction on the sensitive features by using a kernel principal component analysis technology, taking a dimensionality reduction result as a gearbox health index, and then dividing a degradation stage of the gearbox based on a hidden Markov model and a dimensionality bit algorithm. And obtaining a residual service life prediction model training set by using the full-life-cycle gearbox data. And a multi-task semi-supervised residual service life prediction model is constructed and trained, and finally the residual service life of the gearbox is predicted through multiple times of iterative prediction. According to the method, multi-task learning and semi-supervised learning are fused and applied to prediction of the remaining service life of the gearbox, and higher prediction precision can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the technical field of gearbox remaining service life prediction, and in particular relates to a gearbox remaining service life prediction method and system based on multi-task semi-supervised learning. Background Art

[0002] As a key transmission component in industrial equipment, the stable and reliable operation of the gearbox is crucial to the safe operation of the entire system. However, the operating environment of the gearbox is complex and changeable, and its health status will gradually decline over time, which may eventually lead to failure. Therefore, accurately predicting the remaining useful life (RUL) of the gearbox is of great significance for preventing failures, reducing maintenance costs, and improving equipment operation efficiency. Efficient and accurate prediction of the remaining useful life of the gearbox is a research hotspot and difficulty in the industrial field.

[0003] The operating condition of a gearbox is usually affected by a variety of factors, which show certain regularity at the time level. CN202310519912.9 discloses a gearbox life prediction method based on sparse Transformer. However, this method does not fully consider the impact of the gearbox degradation stage on the remaining service life prediction, resulting in low model prediction accuracy.

[0004] In the research on the prediction of the remaining useful life of gearboxes, multi-task learning can simultaneously consider multiple pieces of information related to the remaining useful life of gearboxes, such as failure modes, degradation stages, etc. Semi-supervised learning can use a large amount of unlabeled data combined with a small amount of labeled data to improve the model's generalization ability for unknown situations. By combining multi-task learning and semi-supervised learning, the potential information in the data can be more effectively mined, improving the accuracy and robustness of predictions.

[0005] In view of the above analysis, the technical problem that urgently needs to be solved in the existing technology is: the impact of the gearbox degradation stage on the remaining service life prediction is not fully considered, resulting in low model prediction accuracy. Summary of the invention

[0006] In view of the problems existing in the prior art, the present invention provides a method and system for predicting the remaining useful life of a gearbox based on multi-task semi-supervised learning.

[0007] The present invention is implemented as follows: a method for predicting the remaining service life of a gearbox based on multi-task semi-supervised learning, comprising the following steps:

[0008] Step 1: Gearbox vibration signal collection and preprocessing;

[0009] Collect the vibration signals of the gearbox during its entire life cycle, that is, from the initial healthy state to the end of its life. Data preprocessing includes filling missing values ​​and removing outliers to obtain gearbox vibration data with complete timestamps.

[0010] Step 2, construct health indicators based on EMD-KPCA;

[0011] The empirical modal decomposition (EMD) technology is used to decompose the preprocessed gearbox vibration data to obtain multiple intrinsic mode functions (IMFs), reconstruct the components and trend items that are strongly correlated with the fault characteristics, and then extract sensitive features. The kernel principal element analysis (KPCA) technology is then used to reduce the dimension of the sensitive features to obtain the degradation state of the gearbox throughout its life cycle, that is, the health index.

[0012] Step 3, degradation stage division based on HMM-Viterbi algorithm;

[0013] The constructed health index is used as the observation sequence of the Hidden Markov model (HMM). The model is trained to learn the transition probability and observation probability of the health state. Combined with the Viterbi algorithm, based on the trained HMM, the most likely hidden state path, that is, the most likely degradation stage sequence, is found for the new observation sequence to determine the different degradation stages of the gearbox.

[0014] Step 4: Construction and training of multi-task semi-supervised RUL prediction model;

[0015] The input of the model is the historical health index data of the gearbox, and the output is the degradation stage classification result of this time step and the health index prediction result of the next time step;

[0016] Step 5: Model application and RUL prediction.

[0017] Furthermore, in step 2, time domain features such as kurtosis, pulse index, variance, mean, and frequency domain features such as average frequency and centroid frequency are extracted from the reconstructed signal.

[0018] Furthermore, the degradation stage in step 3 is divided into an early degradation stage, a middle degradation stage, and a late degradation stage.

[0019] Furthermore, in step 4, the unlabeled data is pre-trained through the autoencoder, and the training parameters are retained. After that, the trained autoencoder is used as the backbone network of multi-task learning; a long short-term memory neural network (LSTM) layer based on the attention mechanism is set for RUL prediction, and a network structure based on multilayer perceptrons (MLP) is selected for degradation stage division.

[0020] Furthermore, in step 4, the data is divided into labeled data and unlabeled data, and the labeled data includes two labels: degradation stage and remaining service life;

[0021] Using unsupervised pre-training, after completing the construction of health indicators and the classification of degradation stages, first, a large amount of unlabeled data is used for unsupervised pre-training, the model training parameters are retained, and the autoencoder (AE) is used to learn the compact representation of the data; by training the autoencoder, the model can learn an unsupervised feature extractor, so that the subsequent supervised tasks, i.e., RUL prediction and degradation stage classification, can use these learned features for further training;

[0022] Afterwards, using the data with degradation stage labels and health indicator label information for the next time step, multi-task learning with attention mechanism and long short-term memory neural network is used for labeled training, and the gearbox data is simultaneously predicted for RUL and divided into degradation stages based on multi-task learning. By taking degradation stage division as an auxiliary task, it helps the model learn richer feature representations, and the unsupervised feature extractor obtained through pre-training is used as the feature sharing layer of multi-task learning to share internal parameters, thereby improving the accuracy of RUL prediction.

[0023] Furthermore, in step 5, the sliding window method is used to intercept the historical health indicator data of a certain step length, and the health indicator of the future time step is predicted. After multiple iterations, until the health indicator reaches the life termination threshold, the number of time steps from the current moment to the termination moment is the RUL value.

[0024] Another object of the present invention is to provide a gearbox remaining service life prediction method based on multi-task semi-supervised learning and a gearbox remaining service life prediction system based on multi-task semi-supervised learning, comprising:

[0025] The data acquisition and preprocessing unit collects the vibration signals of the gearbox during its entire life cycle, that is, from the initial healthy state to the end of its life. The data preprocessing includes filling in missing values ​​and removing outliers to obtain the gearbox vibration data with complete timestamps.

[0026] The health index construction unit uses the EMD method to decompose the vibration signal of the gearbox, identify and reconstruct the intrinsic mode functions that are strongly related to the degradation process and reconstruct these components, extract the sensitive features of the reconstructed signal, map the data to a high-dimensional space through KPCA, apply linear PCA in this space to achieve dimensionality reduction, and use the final dimensionality reduction result as the health index of the gearbox;

[0027] The degradation stage division unit uses the health indicators of the gearbox as the observation sequence of the HMM, trains the model to learn the transition probability and observation probability of the health state, and combines the Viterbi algorithm to find the most likely hidden state path for the new observation sequence based on the trained HMM, that is, the most likely degradation stage sequence, and determines the three different degradation stages of the gearbox;

[0028] The model building and training unit adopts a multi-task semi-supervised learning method. First, a large amount of unlabeled data is used for unsupervised pre-training, and the model training parameters are retained. Then, the data with degradation stage labels and health indicator label information for the next time step are used to simultaneously perform degradation stage prediction tasks and remaining service life prediction tasks, promoting knowledge transfer between different tasks and enhancing the generalization ability of the model.

[0029] The model prediction unit, after the model is trained, is tested on other preprocessed data. The sliding window method is used to intercept the historical health indicator data of a certain step length, and the health indicator of the future time step is predicted. After multiple iterations, until the health indicator reaches the life termination threshold, the number of time steps from the current moment to the termination moment is the RUL value.

[0030] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the gearbox remaining service life prediction method based on multi-task semi-supervised learning.

[0031] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the gearbox remaining service life prediction method based on multi-task semi-supervised learning.

[0032] Another object of the present invention is to provide an information data processing terminal, which includes the gearbox remaining service life prediction system based on multi-task semi-supervised learning.

[0033] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0034] First, the present invention takes into account the impact of the degradation stage on the prediction of the remaining service life of the gearbox. Through the shared representation layer of multi-task learning, the model can learn common features across tasks. At the same time, the task-specific layer can capture unique information for the prediction of the remaining service life. Joint training reduces the overfitting of the model for specific tasks, thereby improving the generalization ability of the model on unknown data.

[0035] The present invention takes into account that the unlabeled operating data in the gearbox data is usually far more than the labeled fault data. Through semi-supervised learning, these unlabeled data can be effectively utilized to enhance the prediction ability of the model, reduce the demand for a large amount of labeled data, reduce the cost of data preparation, help the model better understand the distribution of data, and improve the ability to identify abnormal situations.

[0036] Second, the expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: the present invention can accurately predict the remaining life of the gearbox, which can improve the reliability and service life of the gearbox, which is of great significance to the industry, because the gearbox is a key component, and its performance directly affects the life of the entire equipment. In addition, the use of semi-supervised learning can effectively utilize unlabeled data, reduce the demand for a large amount of labeled data, and reduce the cost of data preparation.

[0037] The technical solution of the present invention fills the technical gap in the industry at home and abroad: Taking into account that the unlabeled operating data in the gearbox data is far more than the labeled fault data, the present invention adopts semi-supervised learning to effectively utilize these unlabeled data, reduce the demand for a large amount of labeled data, and improve the ability to identify abnormal situations. Through the shared representation layer of multi-task learning, the model can learn common features across tasks. At the same time, the task-specific layer captures unique information for remaining service life prediction, thereby improving the generalization ability of the model on unknown data.

[0038] Third, technical problem solving:

[0039] 1. Solve the problem of poor vibration signal data quality

[0040] In the prior art, the gearbox vibration signal acquisition process may have problems such as noise interference, missing values, and outliers, resulting in low data quality and affecting the accuracy of life prediction. The present invention ensures the integrity and reliability of vibration data through data preprocessing methods, including missing value filling and outlier removal, and provides a high-quality data foundation for the subsequent construction of health indicators and model training.

[0041] 2. Overcoming the problem of high feature dimension and redundant information

[0042] In traditional life prediction methods, feature extraction mostly relies on simple statistics or frequency domain analysis, which easily leads to a large amount of redundant information in high-dimensional feature data, thereby reducing the performance of the model. The present invention uses EMD decomposition and KPCA dimensionality reduction technology to extract key features from multi-frequency signals and construct health indicators, significantly reducing feature redundancy and improving the effectiveness of data description and the accuracy of life prediction.

[0043] 3. Improve the adaptability and robustness of life prediction models

[0044] Traditional life prediction models based on supervised learning are highly dependent on labeled data, and the labeling process is costly and may be insufficient, which limits the practical application of the model. This invention innovatively adopts a multi-task semi-supervised learning method, effectively utilizes unlabeled data, enhances the model's adaptability to different working conditions and degradation modes, and significantly improves the robustness and generalization performance of the prediction.

[0045] 4. Solve the problem of insufficient accuracy in life prediction

[0046] Existing life prediction technologies often have difficulty accurately predicting the remaining service life in the early or mid-stage of equipment degradation, resulting in unreasonable maintenance strategies. This invention builds a full life cycle degradation model based on health indicators, quantifies the health status of the gearbox into a single indicator, and combines it with a multi-task learning model to achieve accurate prediction of life, providing a scientific basis for predictive maintenance and operation optimization of equipment.

[0047] Technological advancements:

[0048] 1. Improve the accuracy of life prediction

[0049] Through an innovative health indicator construction method, the present invention significantly improves the quantification accuracy of the gearbox degradation state, and combines advanced multi-task learning algorithms to achieve high-precision life prediction, providing technical support for early warning of equipment failures.

[0050] 2. Enhance model adaptability and efficiency

[0051] The present invention makes full use of unlabeled data for training, overcomes the limitations of traditional models in data-scarce scenarios, and reduces the reliance on labeled data. In addition, the application of KPCA dimensionality reduction technology effectively simplifies the model calculation complexity and improves the real-time performance and application efficiency of the algorithm.

[0052] 3. Applicable to various industrial scenarios

[0053] The method of the present invention has good versatility and is applicable to the life prediction needs of different types of gearboxes and under various industrial working conditions. It can be promoted and applied to equipment health management in wind power, metallurgy, petrochemical and other fields.

[0054] 4. Promote the development of predictive maintenance technology

[0055] Through high-precision remaining life prediction, the present invention provides a new idea for predictive maintenance of industrial equipment, which helps to reduce the unplanned downtime rate of equipment, optimize maintenance strategies and extend the service life of equipment, further improving the safety and economy of industrial production. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flow chart of a gearbox life prediction method based on multi-task semi-supervised learning provided by an embodiment of the present invention;

[0057] Figure 2 It is a network architecture diagram of a gearbox remaining service life prediction model based on multi-task semi-supervised learning provided by an embodiment of the present invention;

[0058] Figure 3 is a structural diagram of a gearbox life prediction system based on multi-task semi-supervised learning provided by an embodiment of the present invention;

[0059] Figure 4 This is the prediction and classification result on the pitting data set provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] like Figure 1 As shown, the gearbox remaining service life prediction method based on multi-task semi-supervised learning provided by an embodiment of the present invention includes the following steps:

[0062] Step 1: Gearbox vibration signal collection and preprocessing.

[0063] After a series of data preprocessing such as missing value filling and outlier removal, the original gearbox vibration signal is subjected to vibration data with complete timestamps.

[0064] Step 2: Construction of health indicators based on EMD-KPCA.

[0065] The EMD method is used to decompose the gearbox vibration signal into components of different frequencies to obtain multiple intrinsic mode functions (IMFs). Then, the IMFs and trend terms that are strongly correlated with the fault characteristics are reconstructed. Time domain features such as kurtosis, pulse index, variance, and mean, and frequency domain features such as average frequency and center of gravity frequency are extracted from the reconstructed signal. KPCA is used to reduce the dimension of the extracted features to obtain the degradation state of the gearbox during its entire life cycle, that is, the health index.

[0066] Step 3: Degradation stage division based on HMM-Viterbi algorithm.

[0067] After obtaining the health indicators of the gearbox throughout its life cycle, they are used as the input of the model to train the HMM. This paper uses a mixed Gaussian distribution as the observation probability density function, and sets the number of mixed Gaussians to 3, which represents three different health states, early degradation stage, mid-stage degradation stage, and late degradation stage, and gives each state an independent covariance matrix. This allows the model to learn different observed data distribution shapes for each health state, providing greater flexibility and more accurate data fitting. For the selection of the initial values ​​of π and A, a simple and common initialization strategy is selected, that is, π and A are initialized to a uniform distribution, allowing the model to self-adjust and learn more complex state transitions and initial state distributions based on the training data. After the parameters are initialized, the maximum number of iterations is set to 100, and the algorithm convergence error is set to 0.0001.

[0068] After the training is completed, given the current observation sequence, the Viterbi algorithm is applied to decode the trained HMM model to identify the most likely hidden state sequence and divide the current observation sequence into three degradation stages: early degradation stage, mid-stage degradation stage and late degradation stage.

[0069] Step 4: Build and train a multi-task semi-supervised RUL prediction model.

[0070] To achieve the remaining service life prediction, the input of the model is the historical health index data of the gearbox, and the output is the classification result of the degradation stage at this time step and the prediction result of the health index at the next time step. In time series modeling, in order to match the input dimension of the model, it is usually necessary to window the time series, that is, divide the original data into multiple time series segments according to a certain window length. Therefore, the obtained gearbox health index is divided to intercept the window data similar to the image, and the model training and test sample data sets are constructed. The health index sequence is intercepted using a window of a certain length, and the health index and degradation stage at the next moment of each window are used as the labels corresponding to the window subsequence, and finally the reconstructed data matrix and the corresponding labels are obtained. All training data sets contain multiple N×1 time series, which contain two corresponding label sets: one is the classification label for degradation stage classification, and the other is the regression label for trend prediction.

[0071] The collection of gearbox vibration signals is the first step in predicting its remaining service life. Since the signals in industrial environments may be affected by noise interference, sampling anomalies and other issues, the raw data needs to be preprocessed. Preprocessing includes missing value filling and outlier removal to ensure the integrity and accuracy of the data. The processed vibration signal forms continuous vibration data based on the timestamp, laying the foundation for subsequent feature extraction and modeling.

[0072] The Empirical Mode Decomposition (EMD) method is used to decompose the vibration signal into several Intrinsic Mode Functions (IMFs), each of which corresponds to components of different frequencies. By analyzing these IMFs, components that are highly correlated with the gearbox fault characteristics can be identified. These related IMFs and trend terms are reconstructed into new signals, making the signal more focused on the fault characteristics and filtering out excess noise information.

[0073] Multiple key time domain and frequency domain features are extracted from the reconstructed signal, including kurtosis, pulse index, variance, mean (time domain), average frequency, center of gravity frequency (frequency domain), etc. The above features are significantly descriptive of the health status of the gearbox at different degradation stages. Since high-dimensional features may have redundancy and correlation problems, the kernel principal component analysis (KPCA) method is used to reduce the dimension of the features, and finally a single health index is generated, which can concisely and accurately describe the degradation status of the gearbox throughout its life cycle.

[0074] On the basis of establishing health indicators, the health indicators are combined with known life data for training through a multi-task semi-supervised learning model. This model not only uses labeled degradation data, but also effectively uses unlabeled data to improve the accuracy and stability of life prediction. This method can monitor the degradation trend of the gearbox in real time and accurately predict the remaining service life, providing an important basis for equipment maintenance and operation optimization.

[0075] The network architecture and loss function of the multi-task semi-supervised RUL prediction model are defined as follows:

[0076] 1) Backbone network

[0077] First, the unlabeled data is pre-trained through the autoencoder, and the training parameters are retained. After that, the trained autoencoder is used as the backbone network for multi-task learning. By sharing this encoder layer, not only the complexity of the model is reduced, but also information sharing and interaction between different tasks are achieved, providing more effective feature representation for RUL prediction and degradation stage identification tasks.

[0078] 2) Task-Specific Networks

[0079] In the study, two networks were designed for specific tasks: one for the RUL prediction task and the other for the degradation stage division task. For the RUL prediction task, a Long Short Term Memory Neural Network (LSTM) layer based on the attention mechanism was used. This structure can effectively capture long-term dependencies in sequence data and focus on important time steps through the attention mechanism. For the degradation stage division task, a network structure based on multilayer perceptrons (MLP) was selected. Through the design of these two specific task networks, it is possible to flexibly select suitable network structures according to different task requirements, thereby improving the performance and accuracy of the model in RUL prediction and degradation stage division.

[0080] 3) Loss Function

[0081] When designing a network, the choice of loss function is an important aspect. Because the network needs to handle two tasks at the same time, two loss functions need to be set. The loss function of the prediction task is set to the root mean square loss, and the loss function of the classification task is set to the cross entropy loss. The calculation formulas for the root mean square loss function and the cross entropy loss function are as follows:

[0082]

[0083] In the formula, y represents the true value Represents the predicted value.

[0084]

[0085] in P is the target distribution, Q is the estimated distribution.

[0086] In order to make the two tasks co-trained in the network and use the features learned by the degradation phase division to improve the network's RUL prediction performance, it is not optimal to directly add the two loss functions to get the final loss function of the network. This is because it is not possible to directly determine the contribution of the auxiliary task to RUL prediction. If too much weight is given to the auxiliary task, the features learned by the network may be more inclined to solve the auxiliary task rather than accurately predict RUL. In other words, the auxiliary task may interfere with the feature learning of the main task. Therefore, a hyperparameter is introduced α To control the weight of the auxiliary task. The total loss is expressed as follows.

[0087] L MT-AE =L p +L u +αL c

[0088] Where L p is the loss of the RUL prediction task, L c is the loss of the classification task in the degradation stage, α is the weight factor for controlling the auxiliary task, L u is the loss for unsupervised learning via the AE autoencoder.

[0089] The mean square error (MSE) is used as an indicator to evaluate the life prediction performance of the model, and the accuracy (ACC) is used as an indicator to evaluate the degradation stage recognition performance of the model.

[0090] Step 5: Model application and RUL prediction.

[0091] The test data set is preprocessed and the key health index based on EMD-KPCA is constructed. After the health index is constructed, the sliding window is used to construct the data set so that the matrix input to the network model is N×1. The model will directly output the degradation state at this time and the health index value of the next time step.

[0092] Set the gearbox failure threshold to determine whether the health index at the current moment has reached the threshold. If it has reached the threshold, directly calculate the time difference between this time step and the actual failure point to obtain the remaining life prediction result. If it has not reached the threshold, add the latest predicted value to the rightmost end of the model input sequence, remove the leftmost data of the original input sequence, and update the model input sequence. Repeat the above operation to achieve iterative prediction of the degradation trend, and then calculate the time difference between the predicted failure point and the actual failure point to obtain the remaining life prediction result.

[0093] like Figure 3 As shown, an embodiment of the present invention provides a gearbox remaining service life prediction method based on multi-task semi-supervised learning and a gearbox remaining service life prediction system based on multi-task semi-supervised learning, comprising:

[0094] The data acquisition and preprocessing unit collects the vibration signals of the gearbox during its entire life cycle, that is, from the initial healthy state to the end of its life. The data preprocessing includes filling in missing values ​​and removing outliers to obtain the gearbox vibration data with complete timestamps.

[0095] The health index construction unit uses the EMD method to decompose the vibration signal of the gearbox, identify and reconstruct the intrinsic mode functions that are strongly related to the degradation process and reconstruct these components, extract the sensitive features of the reconstructed signal, map the data to a high-dimensional space through KPCA, apply linear PCA in this space to achieve dimensionality reduction, and use the final dimensionality reduction result as the health index of the gearbox;

[0096] The degradation stage division unit uses the health indicators of the gearbox as the observation sequence of the HMM, trains the model to learn the transition probability and observation probability of the health state, and combines the Viterbi algorithm to find the most likely hidden state path for the new observation sequence based on the trained HMM, that is, the most likely degradation stage sequence, and determines the three different degradation stages of the gearbox;

[0097] The model building and training unit adopts a multi-task semi-supervised learning method. First, a large amount of unlabeled data is used for unsupervised pre-training, and the model training parameters are retained. Then, the data with degradation stage labels and health indicator label information for the next time step are used to simultaneously perform degradation stage prediction tasks and remaining service life prediction tasks, promoting knowledge transfer between different tasks and enhancing the generalization ability of the model.

[0098] The model prediction unit, after the model is trained, is tested on other preprocessed data. The sliding window method is used to intercept the historical health indicator data of a certain step length, and the health indicator of the future time step is predicted. After multiple iterations, until the health indicator reaches the life termination threshold, the number of time steps from the current moment to the termination moment is the RUL value.

[0099] The technical solution of the present invention can be widely applied to multiple industrial and technological fields through its innovative method, thereby significantly improving the efficiency and reliability of various industries. For example, in the wind power industry, the present invention can predict and extend the service life of wind turbine gearboxes, reducing maintenance costs and downtime. The automotive manufacturing industry can also benefit from it, because the health of automatic transmissions and differentials directly affects the reliability and safety of the car. The aerospace field has extremely high requirements for the reliability of gearboxes, and the real-time monitoring and prediction functions of the present invention can ensure flight safety. In rail transportation, the gearbox is the core of the power transmission system, and the present invention helps to improve the operating efficiency and safety of railway vehicles.

[0100] The present invention will be further described below in conjunction with specific embodiments.

[0101] Step 1: Gearbox vibration signal collection and preprocessing.

[0102] The experimental data uses the Chongqing University gearbox degradation public dataset. This dataset provides several gearbox full life cycle vibration datasets (the actual file provides data for the last 400 or 600 minutes of each gear life cycle), including two types of faults, broken teeth and pitting. The gear contact fatigue test bench collects vibration signals on the box body, sampling for 10 seconds every 60 seconds, with a sampling frequency of 50000Hz. In this embodiment, the pitting 1 dataset is selected as the training set, and the pitting 2 dataset is selected as the test set.

[0103] Since the sampling frequency of the original data is very high, the data operation cost is too high. Therefore, in this embodiment, the original vibration signal is first downsampled while retaining the main characteristics of the signal to reduce the amount of data and improve the subsequent analysis and modeling processing speed. This operation is not mandatory. After that, the signal is preprocessed with a series of data preprocessing such as missing value filling and outlier removal to obtain vibration data with complete timestamps.

[0104] Step 2: Construction of health indicators based on EMD-KPCA.

[0105] After decomposing the preprocessed Chongqing University gearbox vibration data using the EMD algorithm, the components and trend items that are strongly correlated with the original signal are reconstructed. The time domain and frequency domain features are extracted from the reconstructed signal, and the extracted features are reduced in dimension using KPCA to obtain the degradation state of the gearbox throughout its life cycle, that is, the health index.

[0106] Step 3: Degradation stage division based on HMM-Viterbi algorithm.

[0107] The proposed HMM-Viterbi analysis method is used to divide the constructed health indicators into three significant stages of degradation: early degradation stage, middle degradation stage, and late degradation stage, which are marked as 0, 1, and 2.

[0108] Step 4: Build and train a multi-task semi-supervised RUL prediction model.

[0109] Using unsupervised pre-training, after completing the construction of health indicators and the classification of degradation stages, first, unlabeled data is trained, and the autoencoder AE is used to learn the compact representation of the data. By training the autoencoder, the model can learn an unsupervised feature extractor so that the subsequent supervised tasks, namely RUL prediction and degradation stage classification, can use these learned features for further training.

[0110] Afterwards, the gearbox data is simultaneously predicted for RUL and classified into degradation stages based on multi-task learning. By using degradation stage classification as an auxiliary task, the model can learn richer feature representations, and the unsupervised feature extractor obtained through pre-training is used as the feature sharing layer of multi-task learning to share internal parameters, thereby improving the accuracy of RUL prediction.

[0111] Step 5: Model application and RUL prediction.

[0112] For the pitting 2 test data set, it first undergoes a series of data preprocessing operations such as missing value filling and outlier removal, and then uses the EMD-KPCA method to construct the health index.

[0113] The sliding window method is used to intercept the input health index and predict the health index data of the degradation stage and the next time step. The sliding window length is selected as 10, that is, the health index data of the first 10 time steps are used to predict the health index data of the degradation stage and the next time step.

[0114] The gearbox failure threshold is set to 1, and the time difference between the predicted failure point and the actual failure point is calculated as the remaining life prediction result.

[0115] The prediction and classification results on the pitting dataset are as follows Figure 4 shown.

[0116] In order to verify the effectiveness of the multi-task semi-supervised model, the performance of semi-supervised learning and multi-task learning is compared through ablation experiments on the pitting 2 data. The effects of using unlabeled pre-training and not using unlabeled pre-training, that is, the effect of the multi-task supervised model, and only training a single RUL prediction task without training the degradation stage division task, that is, the effect of the single-task semi-supervised model, are compared. The advantages of the model are deeply explored. The experimental results are shown in Table 1 below, which verifies that the multi-task semi-supervised remaining useful life prediction method proposed in this method can achieve the highest prediction accuracy.

[0117] Table 1 Comparison of prediction effects of different models

[0118]

[0119] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0120] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A gearbox remaining service life prediction method based on multi-task semi-supervised learning, characterized in that: The following steps are involved: Step 1: Gearbox vibration signal collection and preprocessing; Collect the vibration signals of the gearbox during its entire life cycle, that is, from the initial healthy state to the end of its life. Data preprocessing includes filling missing values ​​and removing outliers to obtain gearbox vibration data with complete timestamps. Step 2, construct health indicators based on EMD-KPCA; The Empirical Mode Decomposition (EMD) technique is used to decompose the preprocessed gearbox vibration data to obtain multiple intrinsic mode functions (IMFs), reconstruct the components and trend items that are strongly correlated with the fault characteristics, and then extract the sensitive features. The Kernel Principal Component Analysis (KPCA) technique is then used to reduce the dimension of the sensitive features to obtain the degradation state of the gearbox throughout its life cycle, i.e., the health index. Step 3, degradation stage division based on HMM-Viterbi algorithm; The constructed health index is used as the observation sequence of the hidden Markov model HMM, and the model is trained to learn the transition probability and observation probability of the health state. Combined with the Viterbi algorithm, based on the trained HMM, the most likely hidden state path is found for the new observation sequence, that is, the most likely degradation stage sequence, to determine the different degradation stages of the gearbox; Step 4: Construction and training of multi-task semi-supervised RUL prediction model; The input of the model is the historical health index data of the gearbox, and the output is the degradation stage classification result of this time step and the health index prediction result of the next time step; Step 5: Model application and RUL prediction.

2. The gearbox remaining service life prediction method based on multi-task semi-supervised learning according to claim 1 is characterized in that: In step 2, time domain features are extracted from the reconstructed signal: kurtosis, pulse index, variance, mean, and frequency domain features: Average frequency, center of gravity frequency.

3. The gearbox remaining service life prediction method based on multi-task semi-supervised learning according to claim 1, characterized in that: In step 3, the degradation stage is divided into an early degradation stage, a middle degradation stage, and a late degradation stage.

4. The gearbox remaining service life prediction method based on multi-task semi-supervised learning according to claim 1, characterized in that: In step 4, the unlabeled data is pre-trained through the autoencoder and the training parameters are retained. After that, the trained autoencoder is used as the backbone network for multi-task learning. The long short-term memory neural network LSTM layer based on the attention mechanism is set for RUL prediction, and the network structure based on the multi-layer perceptron MLP is selected for degradation stage division.

5. The method for predicting the remaining useful life of a gearbox based on multi-task semi-supervised learning according to claim 1, characterized in that: In step 4, the data is divided into labeled data and unlabeled data, and the labeled data contains two labels: degradation stage and remaining service life; Using unsupervised pre-training, after completing the construction of health indicators and the classification of degradation stages, first, a large amount of unlabeled data is used for unsupervised pre-training, the model training parameters are retained, and the autoencoder AE is used to learn the compact representation of the data; by training the autoencoder, the model learns an unsupervised feature extractor, so that the subsequent supervised tasks, namely RUL prediction and degradation stage classification, use these learned features for further training; Afterwards, using the data with degradation stage labels and health indicator label information for the next time step, multi-task learning with attention mechanism and long short-term memory neural network is used for labeled training, and the gearbox data is simultaneously predicted for RUL and divided into degradation stages based on multi-task learning. By taking degradation stage division as an auxiliary task, it helps the model learn richer feature representations, and the unsupervised feature extractor obtained through pre-training is used as the feature sharing layer of multi-task learning to share internal parameters, thereby improving the accuracy of RUL prediction.

6. The method for predicting the remaining useful life of a gearbox based on multi-task semi-supervised learning according to claim 1, characterized in that: In step 5, the sliding window method is used to intercept the historical health indicator data of a certain step length, and the health indicator of the future time step is predicted. After multiple iterations, until the health indicator reaches the life termination threshold, the number of time steps from the current moment to the termination moment is the RUL value.

7. A gearbox remaining service life prediction system based on multi-task semi-supervised learning according to any one of claims 1 to 6, characterized in that: include: The data acquisition and preprocessing unit collects the vibration signals of the gearbox during its entire life cycle, that is, from the initial healthy state to the end of its life. The data preprocessing includes filling in missing values ​​and removing outliers to obtain the gearbox vibration data with complete timestamps. The health index construction unit uses the EMD method to decompose the vibration signal of the gearbox, identify and reconstruct the intrinsic mode functions that are strongly related to the degradation process and reconstruct these components, extract the sensitive features of the reconstructed signal, map the data to a high-dimensional space through KPCA, apply linear PCA in this space to achieve dimensionality reduction, and use the final dimensionality reduction result as the health index of the gearbox; The degradation stage division unit uses the health indicators of the gearbox as the observation sequence of the HMM, trains the model to learn the transition probability and observation probability of the health state, and combines the Viterbi algorithm to find the most likely hidden state path for the new observation sequence based on the trained HMM, that is, the most likely degradation stage sequence, and determines the three different degradation stages of the gearbox; The model building and training unit adopts a multi-task semi-supervised learning method. First, a large amount of unlabeled data is used for unsupervised pre-training, and the model training parameters are retained. Then, the data with degradation stage labels and health indicator label information for the next time step are used to simultaneously perform degradation stage prediction tasks and remaining service life prediction tasks, promoting knowledge transfer between different tasks and enhancing the generalization ability of the model. The model prediction unit, after the model is trained, is tested on the preprocessed data. The sliding window method is used to intercept the historical health indicator data of a certain step length, and the health indicator of the future time step is predicted. After multiple iterations, until the health indicator reaches the life termination threshold, the number of time steps from the current moment to the termination moment is the RUL value.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the gearbox remaining service life prediction method based on multi-task semi-supervised learning as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the gearbox remaining service life prediction method based on multi-task semi-supervised learning as described in any one of claims 1 to 6.

10. An information data processing terminal, characterized in that: The information data processing terminal includes the gearbox remaining service life prediction system based on multi-task semi-supervised learning as described in claim 7.

Citation Information

Patent Citations

  • Gearbox life prediction method based on sparse Transform

    CN116465623A

Cited By

  • Method and device for predicting service life of optical module

    CN120357962A

  • Method for predicting fatigue life and evaluating residual life of high-power heavy-duty gearbox

    CN121167677A

  • Fatigue life prediction and residual life evaluation method for high-power heavy-load gearboxes

    CN121167677B

  • High-horsepower tractor gearbox state on-line monitoring method

    CN121384446A