Multi-target domain bearing residual service life prediction method and system
By filling the source domain data and embedding 1-dimensional convolutional layer in the offline stage, training is used with ModernTCN, and target domain weight allocation and distribution difference measurement is used in the online stage, weighted hybrid target domain is constructed for bearing RUL prediction, which solves the problems of poor adaptability and high computing resource consumption in the multi-target domain environment, and achieves more accurate and efficient RUL prediction.
Patent Information
- Application Number
- CN202510433006.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing bearing residual service life (RUL) prediction methods are difficult to adapt in multi-objective environments, and traditional deep learning models face the problems of data distribution differences and high computing resource consumption in practical applications.
A multi-objective domain bearing RUL prediction method is proposed. By filling the source domain data and embedding 1-dimensional convolutional layer in the offline stage, training is carried out using ModernTCN. In the online stage, dynamic time regularization (DTW) and maximum mean difference (MMD) are used for target domain weight allocation and distribution difference measurement, weighted mixed target domains are constructed for RUL prediction, and the prediction results are optimized by fine-tuning model parameters.
The generalization ability of the RUL prediction model is improved, so that it can adapt to diversified industrial application scenarios more accurately, reduce the consumption of online computing resources, and improve the operation efficiency of the prediction system.
Smart Images

Figure CN119939259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remaining useful life (RUL) prediction of mechanical equipment, and in particular to a bearing RUL prediction method in multiple target domains. Background Art
[0002] With the rapid development of industrial big data and industrial Internet of Things, the prediction and health management (PHM) technology in mechanical reliability engineering is becoming increasingly important, among which the remaining useful life (RUL) prediction of equipment has become a key task. Bearings are vital rotating parts in mechanical equipment, and their performance degradation directly affects the overall performance of the mechanical system. Therefore, bearing RUL prediction is of great significance for improving the PHM level, optimizing maintenance strategies, and enhancing the reliability of rotating machinery.
[0003] Currently, many bearing RUL prediction studies focus on data-driven methods based on deep learning technology, but traditional deep learning models face many challenges in practical applications. On the one hand, the distribution of training and test data sets must be consistent when predicting RUL. However, in actual industrial environments, the operating conditions of equipment are complex and changeable, and the data distribution varies greatly, making it difficult to meet this requirement, resulting in limited model performance. On the other hand, most traditional methods use offline modeling, and model parameters cannot be adjusted online, making it difficult to adapt to new working conditions. Although online modeling methods can cope with new working conditions, they consume large amounts of computing resources and have low model application efficiency.
[0004] In addition, most existing studies assume that the test data only comes from a single target domain. However, in actual industrial processes, there is no guarantee that the domain where the test data is located has complete life cycle data, and it is also difficult to determine whether it belongs to the source domain or the target domain. It is difficult to obtain good prediction results directly using existing methods. Summary of the invention
[0005] In order to solve the problems existing in the prior art, the present invention provides a multi-objective domain bearing remaining service life prediction method, the method comprising the following steps: S1, randomly selects data under different working conditions from the industrial process bearing operation related data set as the source domain data , ,……, , n>=1, is the number of source domains, target domain data , ,……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as part of the target domain training data; S2, in the offline stage, the source domain data The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding ; S3, using ModernTCN Training is performed to obtain a model model_1; the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; S4, online stage, compares the test data Test and the target domain data , ,……, Dynamic time warping (DTW) is used to align the sequences and calculate the distance, and the initial weight of the target domain is assigned according to the distance; S5, using the maximum mean difference MMD to measure the distribution difference between the target domain and the test data after DTW processing, and calculating the value of the initial weighted maximum mean difference WMMD; S6, taking the last two layers of model model_1 as the remaining useful life predictor model_2, inputting the weighted mixed target domain, and obtaining the predicted remaining useful life; S7, taking the WMMD value and the sum of the mean square error between the predicted remaining useful life obtained by model_2 and the actual remaining useful life as the loss function, fine-tuning some model parameters of model_1, and then repeating S4, S5 and S6 in sequence to obtain the final remaining useful life prediction model model_1_final, and finally passing the test data Test through model_1_final to obtain the corresponding remaining useful life prediction target value.
[0006] Furthermore, the ModernTCN block includes a deep convolution DWConv and two convolutional feedforward networks ConvFFN, namely ConvFFN1 and ConvFFN2. DWConv is responsible for independently learning the temporal information between tags on each feature channel, ConvFFN1 is responsible for generating new feature representations for each variable, and ConvFFN2 is responsible for capturing the inter-variable dependencies of each feature.
[0007] Furthermore, the specific implementation of step S3 is as follows: (3.1) Input to the backbone network to capture interdependencies across time and variables, thereby learning informative representations , the formula is as follows:
[0008] Among them, M represents the number of variables contained in the source domain data, N represents the number of patches, and D represents the number of channels. It is a stacked ModernTCN block;
[0009] in, , {1, ...,K} is the input of the i-th ModernTCN block, Represents the ModernTCN block. After passing through the backbone network, the final ; (3.2) The final prediction result is obtained using a linear head with a flattened layer:
[0010] in, is the prediction result of a freely set length of T, containing M variables; An operation refers to a layer that reshapes the final representation into a tensor Z with dimension M×(D×N); Represents a linear projection layer that transforms the final representation into the final prediction result; (3.3) In the offline stage, the loss function calculation formula of model_1 is as follows:
[0011] Where n represents the number of source domains, Represents the true label.
[0012] Furthermore, the weight allocation process using dynamic time warping (DTW) in step S4 is as follows: (4.1) Dynamic time warping (DTW) is used to measure the similarity between two sequences. Given two time series, )and ), DTW(P, Q) is defined as the optimal path among all possible paths T; a path T is a sequence ),in , , , a, b represent the sequence number, and for each step k∈ Belongs to the collection , the DTW formula is as follows:
[0013] Where t represents the number of nodes on the optimal path, represents the value of the kth dimension of the i-th data point in the time series P, represents the value of the kth dimension of the jth data point in the time series Q, is the Euclidean distance, which is a distance metric calculated as follows:
[0014] (4.2) The weighting coefficient is obtained by normalizing the DTW distance. The calculation formula is as follows:
[0015] in, represents the number of labeled target domains, r represents the rth labeled target domain, represents the labeled target domain data embedded by the backbone network, Represents the test data after being embedded in the backbone network.
[0016] Furthermore, in step S5, WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW to achieve feature-level transfer learning between the test data and multiple target domains. The calculation formula is as follows:
[0017] in, Indicates the value of WMMD, () is a mapping function that maps data to a high-dimensional feature space. represents the norm of the Hilbert space.
[0018] Furthermore, in step S6, the hybrid target domain is defined as:
[0019] in, Represents a mixed target domain.
[0020] Furthermore, the loss function formula between the true label and the target domain data label in the remaining useful life predictor model_2 is as follows:
[0021] Where m represents the number of target domain data, represents the true label, Represents the predicted value obtained by predictor model_2.
[0022] Further, in step S7, based on the random freezing of some model parameters in step S3, other parameters use pre-trained values, and then fine-tune the model model_1 to complete the model-based transfer learning between the second-level source domain and the multi-target domains. The formula of the overall optimization function of multi-level domain transfer learning is as follows:
[0023] Among them, λ and γ are regularization parameters.
[0024] Furthermore, the root mean square error RMSE and the mean absolute error MAE are used as evaluation indicators to evaluate the final prediction results.
[0025] The present invention also provides a multi-objective domain bearing remaining service life prediction system, comprising the following modules: The data domain acquisition module is used to randomly select data under different working conditions from the industrial process bearing operation related data set as the source domain data , ,……, , n>=1, is the number of source domains, target domain data , ,……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as part of the target domain training data; The source domain data preprocessing module is used to preprocess the source domain data in the offline stage. The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding ; Offline feature extraction module, using ModernTCN Training is performed to obtain a model model_1; the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; The data relationship measurement module is used in the online stage to compare the test data Test and the target domain data , ,……, Dynamic time warping (DTW) is used to align the sequences and calculate the distance, and the initial weight of the target domain is assigned according to the distance; A distribution difference calculation module is used to measure the distribution difference between the target domain and the test data after DTW processing by using the maximum mean difference MMD, and calculate the value of the initial weighted maximum mean difference WMMD; The remaining useful life prediction module is used to use the last two layers of model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life; The final prediction module is used to use the WMMD value and the sum of the mean square error between the predicted remaining useful life obtained through model_2 and the actual remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, distribution difference calculation module and remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value.
[0026] Compared with the prior art, the above technical solution conceived by the present invention can achieve the following beneficial effects: This paper proposes a novel RUL transfer learning method across multiple target domains. In the current research field, the literature on related topics seems to be scarce.
[0027] A multi-level domain adaptation strategy is designed, in which the improved weighted online MMD completes the transfer of feature extraction and degradation knowledge from the source domain to the target domain and test set, improving the generalization ability of the RUL prediction model and enabling it to adapt to diverse industrial application scenarios more accurately.
[0028] The strategy of placing the main training task of the model in the offline stage and then performing online fine-tuning based on the test set not only ensures the accuracy of the prediction, but also avoids excessive occupation of online computing resources, greatly improving the operating efficiency of the entire prediction system. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flow chart provided by an embodiment of the present invention; FIG2 is a schematic diagram of RUL prediction fitting effects of four tasks provided by an embodiment of the present invention; FIG3 is an ablation experiment on Task 1 provided by an embodiment of the present invention, wherein (a) the method proposed by the present invention; (b) no fine-tuning; (c) no WMMD; and (d) only one target domain. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0031] In order to solve the problem of bearing multi-target domain RUL prediction and create a transferable model that can dynamically adapt to online test data from different target domains in real time, the present invention provides a lightweight bearing remaining service life prediction method based on multi-level domain transfer. First, in the offline stage, different working conditions are randomly selected from the industrial process bearing operation related data set as the source domain, target domain and test data, the source domain data is padded and segmented and input into the 1D convolution layer, and then the embedded vector is used to preliminarily model the ModernTCN network. In the subsequent online stage, DTW is used to calculate the distance between the test data and the target domain data and assign weights, and then MMD is used to calculate the distribution difference between the target domain and the test set after DTW processing to obtain the WMMD value. Next, the last two layers of ModernTCN are taken to construct the RUL predictor, and the weighted mixed target domain data is input to obtain the predicted remaining service life. The initial model ModernTCN is fine-tuned with the sum of the mean square deviation of the weighted true RUL and the predicted RUL as the loss function. After iterative optimization, the Test is finally input into the model to obtain the prediction of the remaining service life. Figure 1 The flowchart of the present invention is shown below with specific implementation examples.
[0032] (1) Data collection and processing: Randomly select different working conditions from the industrial process bearing operation related dataset as the source domain , ,……, , n>=1, is the number of source domains, target domain , ,……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as part of the target domain training data. Specifically: Obtaining data on bearing operation in industrial processes , where M represents the number of variable features and L represents the sequence length; Source domain data The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding , specifically: (1.2.1) Input the source domain of M variables of length L into the data sequence Unzip to ; (1.2.2) The step size in the block process is S, and it is repeated The last value of (PS) times and adds it to at the end of , to keep N = L / / S.
[0033] (1.2.3) Fill the Divide into N patches of size P. Then input these patches into a 1D convolutional layer, which maps 1 input channel to D output channels, and obtains the embedding vector , the formula is as follows:
[0034] (2) Model pre-training and online domain adaptation (2.1) Using ModernTCN The training is performed to obtain a model model_1, wherein the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; specifically: (2.1.1) Input to the backbone network to capture interdependencies across time and variables, thereby learning informative representations , the formula is as follows:
[0035] It is a stack of ModernTCN blocks.
[0036]
[0037] in, , {1, ...,K} is the input of the i-th ModernTCN block, represents the ModernTCN block. After passing through the backbone network, we get the final .
[0038] (2.1.2) The ModernTCN block includes a deep convolution DWConv and a convolutional feedforward network ConvFFN. The following are their respective operations and functions: (2.1.2.1) Deep convolution DWConv: responsible for independently learning the temporal information between tokens on each feature channel, mixing information in the temporal dimension, Input channels are mapped to output channels to achieve independence of variable and feature dimensions.
[0039] (2.1.2.2) Convolutional Feedforward Network ConvFFN: Based on the idea of group convolution, ConvFFN is decoupled to form ConvFFN1 and ConvFFN2. ConvFFN1 can learn new feature representations for each variable independently by setting the number of groups to M, while ConvFFN2 can capture cross-variable dependencies in each feature dimension by setting the number of groups to D.
[0040] (2.1.3) Use the linear head with the flatten layer to get the final prediction result:
[0041] in, It is the prediction result of length T which we set arbitrarily, containing M variables. An operation refers to a layer that reshapes the final representation into a tensor Z with dimension M×(D×N). Represents a linear projection layer that transforms the final representation into the final prediction result.
[0042] (2.1.3) In the offline stage, the loss function is calculated as follows:
[0043] Where n represents the number of source domains, Represents the true label. When the loss function converges, the pre-training ends, and when the loss is minimized, the best model is saved. This study freezes and copies the best pre-trained model parameters of the hidden layer in the source domain and uses them as fixed parameters of the model in the subsequent target domain.
[0044] (2.2) In the online phase, the test data Test and the target domain data , ,……, Dynamic time warping (DTW) is used to align the sequences and calculate the distances, and the initial weights of the target domains are assigned according to the distances. Specifically: (2.2.1) Dynamic Time Warping (DTW) is used to measure the similarity between two sequences. Given two time series, )and ), DTW(P, Q) is defined as the optimal path among all possible paths T. A path T is a sequence ),in , , , a, b represent the sequence number, and for each step k∈ Belongs to the collection , the DTW formula is as follows:
[0045] Where t represents the number of nodes on the optimal path, represents the value of the kth dimension of the i-th data point in the time series P, represents the value of the kth dimension of the jth data point in the time series Q, is the Euclidean distance, which is a distance metric calculated as follows:
[0046] (2.2.2) The weighting coefficient is obtained by normalizing the DTW distance. The calculation formula is as follows:
[0047] in, represents the number of labeled target domains, r represents the rth labeled target domain, represents the labeled target domain data embedded by the backbone network, Represents the test data after being embedded in the backbone network.
[0048] (2.3) The maximum mean difference (MMD) is used to measure the distribution difference between the target domain and the test set after DTW processing, and the value of the initial weighted maximum mean difference (WMMD) is calculated. Specifically: WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW to achieve feature-level transfer learning between test data and multiple target domains. The calculation formula is as follows:
[0049] in, Indicates the value of WMMD, () is a mapping function that maps data to a high-dimensional feature space. represents the norm of the Hilbert space. By minimizing , which enables the model to learn how to adjust parameters to reduce the difference in data distribution between the test domain and multiple target domains, and realize feature-level transfer learning between test data and multiple target domains.
[0050] (3) RUL prediction and model iteration optimization (3.1) The last two flat layers and linear layer of ModernTCN are used as RUL predictor model_2, the model parameters trained in the source domain are retained, and the weighted mixed target domain is input to obtain the predicted remaining useful life. Specifically: (3.1.1) The hybrid target domain is defined as follows:
[0051] (3.1.2) In order to obtain the degradation information contained in the target data label, use the RUL prediction results in model_2 Predict the mixed target domain. The loss function formula between the true label and the target data label is as follows:
[0052] Where m represents the number of target domains, Represents the true label.
[0053] (3.2) The WMMD value and the sum of the mean square error between the predicted remaining useful life obtained by model_2 and the actual remaining useful life are used as the loss function to fine-tune some model parameters of model_1. After fine-tuning, the updated model_1 is obtained, and then (2.2), (2.3) and (3.1) are repeated in sequence to obtain the final RUL prediction model model_1_final. Finally, test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value. Specifically: (3.2.1) Based on the random freezing of some model parameters in step S3, the other parameters use pre-trained values, and then fine-tune the model to complete the model-based transfer learning between the second-level source domain and the multiple target domains. The above loss functions are integrated, and the formula of the overall optimization function of multi-level domain transfer learning is as follows:
[0054] Among them, λ and γ are regularization parameters. In this process, the transfer learning between the target domain and the test data, and between the source domain and the target domain continues.
[0055] (3.2.2) When the loss function converges, Test is input into the trained model to calculate the final remaining useful life prediction target value. The formula is as follows:
[0056] in represents the input embedding after padding segmentation, represents the trained model, Represents the input embedding go through The remaining useful life forecast.
[0057] A second aspect of the present invention provides a multi-objective domain bearing remaining service life prediction system, comprising the following modules: The data domain acquisition module is used to randomly select data under different working conditions from the industrial process bearing operation related data set as the source domain data , ,……, , n>=1, is the number of source domains, target domain data , ,……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as part of the target domain training data; The source domain data preprocessing module is used to preprocess the source domain data in the offline stage. The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding ; Offline feature extraction module, using ModernTCN Training is performed to obtain a model model_1; the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; The data relationship measurement module is used in the online stage to compare the test data Test and the target domain data , ,……, Dynamic time warping (DTW) is used to align the sequences and calculate the distance, and the initial weight of the target domain is assigned according to the distance; A distribution difference calculation module is used to measure the distribution difference between the target domain and the test data after DTW processing by using the maximum mean difference MMD, and calculate the value of the initial weighted maximum mean difference WMMD; The remaining useful life prediction module is used to use the last two layers of model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life; The final prediction module is used to use the WMMD value and the sum of the mean square error between the predicted remaining useful life obtained through model_2 and the actual remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, distribution difference calculation module and remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value.
[0058] The specific implementation method of each module corresponds to each step and is not described in detail in the present invention.
[0059] The beneficial effects that can be achieved by the present invention are further explained below in conjunction with specific application scenarios.
[0060] In order to verify the effectiveness of the proposed method in predicting RUL under various working conditions using online insufficient unseen data, a series of experiments were conducted on the IEEE PHM Challenge 2012 bearing dataset and the XJTU-SY bearing dataset. The IEEE PHM Challenge 2012 bearing dataset (hereinafter referred to as PHM) was collected from the PRONOSTIA platform. The vibration signals of the bearings were measured by two accelerometers in the horizontal and vertical directions with a sampling frequency of 25.6 kHz. The dataset contains 17 running failure bearings under three different working conditions with different speeds and radial forces. The XJTU-SY bearing dataset (hereinafter referred to as XJTU) was collected by Xi'an Jiaotong University and Changxing Sumitomo Technologies. The sampling content and frequency of this dataset are the same as those of PHM. This dataset contains 15 running bearings under three different working conditions. Table 1 lists the detailed test conditions of these data.
[0061] Table 1 Experimental conditions of IEEE PHM Challenge 2012 and XJTU-SY bearing datasets
[0062] To illustrate the versatility of our proposed method, we randomly select a working condition from the PHM dataset as source domain 1 (S1) and a working condition from the XJTU dataset as source domain 2 (S2) to perform the transfer prediction task. The target domain is another working condition randomly selected from the PHM and XJTU datasets. The test data does not belong to the same domain as the training data in the source domain, but belongs to the same domain as the training data in some target domains. The data of the two datasets are cross-selected for training in the same task, and the two datasets are from different devices. The differences between devices are generally larger than the differences within the devices, so it is more challenging to verify the effectiveness of the model in this case. Table 2 lists the detailed experimental settings. In each task, we used two target domains and considered the multi-target domain problem to extract features. The data of the test set comes from the same device as one of the target domains but with different bearings. The experiment proves that effective information can be extracted from data of different devices and different bearings.
[0063] Table 2 Experimental settings
[0064] For the feature extractor, i.e., the backbone network, it adapts all the same parameter settings. The patch size is 8, the patch step is 4, the dropout rate is 0.3, the learning rate is OneCycleLR dynamic learning rate, and the initial value is 0.0001. In order to verify the prediction accuracy of the proposed method, the root mean square error (RMSE) and the mean absolute error (MAE) are selected as the evaluation indicators.
[0065]
[0066]
[0067] In the formula, n represents the number of samples, represents the true value of the i-th sample, is the predicted value of the ith sample.
[0068] Figure 2 shows the RUL prediction values and their 95% confidence intervals for the four tasks in Table 2. The proposed method produces good RUL prediction results, which are very close to the decreasing trajectory of the actual RUL. From the perspective of the 95% confidence interval, the prediction uncertainty at each moment is very small. These results verify the effectiveness of the proposed method in extracting domain-invariant features and promoting the transfer of fault degradation diversity, which is crucial for industrial maintenance.
[0069] The proposed method is compared with three representative RUL prediction methods to analyze the tasks listed in Table 2. These methods include LSTM, transferable CNN + MMDs, and TPDAN.
[0070] LSTM: A method for predicting remaining useful life using deep feature representation and employing a long short-term memory (LSTM) neural network.
[0071] CNN+MMD: Transferable Convolutional Neural Networks with Multi-kernel Maximum Mean Divergence TPDAN: A multi-source ensemble adversarial adaptation strategy with stage-weighted maximum mean divergence and ranking-based feature regularization.
[0072] All compared methods used the best parameters provided in the original paper. Performance comparison is shown in Table 3. It is worth noting that the RUL transfer prediction in this paper is a supervised prediction task, which is trained using the full life cycle labeled data in the target domain. As shown in Table 3, the method proposed in this paper obtains the smallest RMSE and MAE in the four tasks, indicating that this method significantly improves the prediction performance.
[0073] Table 3 Error result data
[0074] As can be seen from the table, the proposed method has the smallest prediction error in all tasks. The error of our method is less than 50% of that of LSTM, which may be because this method cannot obtain satisfactory prediction results when there are differences in data distribution between different machines. Although CNN+MMD uses domain adaptation to extract domain-invariant features, different working conditions have different degradation trends, and using only a single prediction model is not enough to mine the specific degradation information of each source domain. Therefore, the performance of CNN+MMD is not as good as that of the proposed method, and the error is more than four times that of ours. TPDAN performs the worst in most tasks. The reason is that it trains multiple discriminators without considering multi-source and multi-target conditions.
[0075] Furthermore, this paper discusses the impact of some important hyperparameters in the model on the prediction performance through experiments. There are three main hyperparameters in our proposed model: , and The weights in , and the learning rate. They all have some impact on the results and are adjusted for each task in this experiment. Select hyperparameters with smaller errors for testing. Here, we take Task 1 as an example to illustrate how to obtain the parameters.
[0076] Table 4 lists and The hyperparameter adjustment results show that under combination 1, the RMSE and MAE of the model reach the best, and our model finally selects this combination.
[0077] Table 4 Hyperparameter adjustment
[0078] In addition, to verify the importance of certain modules in the method, ablation experiments were designed in the example, including the entire model, the model without fine-tuning, the model without WMMD, and the model using only one target domain. In order to further verify the function and importance of multi-level transfer learning and multi-target in the proposed method, we conducted three ablation experiments. The detailed description of the ablation experiment is shown in Table 5.
[0079] Table 5 Detailed description of ablation experiment
[0080] Still taking Task 1 as an example, the remaining life prediction results are shown in Table 6. Each component in the proposed method is beneficial to RUL prediction.
[0081] Table 6 Performance indicators of Task 1 in different ablation experiments
[0082] Multi-level domain adaptation plays an important role in the proposed method and improves the ability of ordinary feature extraction. To verify this contribution, we compute and visualize the experimental results of multi-source and multi-target domains. Due to limited space, only the experimental results of Task 1 are plotted, as shown in Figure 3.
[0083] As shown in (a) of Figure 3, they are the experimental results of the features of two source domains and one mixed target domain after passing through the feature extractor and before fine-tuning. The difference between these experimental results is very obvious. (b) of Figure 3 shows the situation after the network is fine-tuned. Now, the experimental results of the source domain and mixed target domain features are more consistent.
[0084] As shown in (c) of Figure 3, they are the experimental results of features from the two target domains and the test set without WMMD, and the differences between these experimental results are quite significant. (d) of Figure 3 is the distribution diagram of the features of the mixed target domain and the test set, showing the consistency of their distribution.
[0085] Multi-level domain adaptation reduces the differences in data distribution in different domains, and the extracted features are more universal, which is why the model achieves excellent results.
[0086] In this patent document, the embodiments and detailed descriptions provided are intended to illustrate the concepts, principles and applications of the present invention, rather than to limit the scope of protection of the present invention. The scope of protection of the present invention shall be determined according to the claims, and it should be understood that any technical improvement, substitution or change that conforms to the innovative spirit and essence of the present invention, as long as it does not exceed the scope of protection defined by the claims, shall be deemed to be within the scope of protection of the present invention. Under the enlightenment of the present invention, technicians in this field can make various forms of replacement or adjustment to adapt to different application scenarios and needs without departing from the core idea and scope of protection of the present invention.
Claims
1. A multi-objective domain bearing remaining service life prediction method, characterized in that: The steps include: S1, randomly selects data under different working conditions from the industrial process bearing operation related data set as the source domain data , target domain data , and test data Test, ensure that the test data is in a different domain from the source domain data and in the same domain as part of the target domain data; S2, in the offline stage, the source domain data The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding ; S3, using ModernTCN Training is performed to obtain a model model_1; the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; S4, online stage, compares the test data Test and the target domain data Dynamic time warping (DTW) is used to align the sequences and calculate the distance, and the initial weight of the target domain is assigned according to the distance; S5, using the maximum mean difference MMD to measure the distribution difference between the target domain and the test data after DTW processing, and calculating the value of the initial weighted maximum mean difference WMMD; S6, taking the last two layers of model model_1 as the remaining useful life predictor model_2, inputting the weighted mixed target domain, and obtaining the predicted remaining useful life; S7, taking the WMMD value and the sum of the mean square error between the predicted remaining useful life obtained by model_2 and the actual remaining useful life as the loss function, fine-tuning some model parameters of model_1, and then repeating S4, S5 and S6 in sequence to obtain the final remaining useful life prediction model model_1_final, and finally passing the test data Test through model_1_final to obtain the corresponding remaining useful life prediction target value.
2. The method for predicting the remaining useful life of a bearing in a multi-objective domain according to claim 1, characterized in that: The ModernTCN block includes a deep convolution DWConv and two convolutional feedforward networks ConvFFN, namely ConvFFN1 and ConvFFN2. DWConv is responsible for independently learning the temporal information between tags on each feature channel. The role of ConvFFN1 is to generate a new feature representation for each variable, while the task of ConvFFN2 is to capture the inter-variable dependencies of each feature.
3. The method for predicting the remaining useful life of a bearing in a multi-objective domain according to claim 1, characterized in that: The specific implementation of step S3 is as follows: (3.1) Input to the backbone network to capture interdependencies across time and variables, thereby learning informative representations , the formula is as follows: ; Among them, M represents the number of variables contained in the source domain data, N represents the number of patches, and D represents the number of channels. It is a stacked ModernTCN block; ; in, , {1, ...,K} is the input of the i-th ModernTCN block, Represents the ModernTCN block. After passing through the backbone network, the final ; (3.2) The final prediction result is obtained using a linear head with a flattened layer: ; in, is the prediction result of a freely set length of T, containing M variables; An operation refers to a layer that reshapes the final representation into a tensor Z with dimension M×(D×N); Represents a linear projection layer that transforms the final representation into the final prediction result; (3.3) In the offline stage, the loss function calculation formula of model_1 is as follows: ; Where n represents the number of source domains, Represents the true label.
4. The method for predicting the remaining useful life of a bearing in a multi-objective domain according to claim 1, characterized in that: The weight allocation process using dynamic time warping (DTW) in step S4 is as follows: (4.1) Dynamic time warping (DTW) is used to measure the similarity between two sequences. Given two time series, )and ), DTW(P, Q) is defined as the optimal path among all possible paths T; a path T is a sequence ),in , , , a, b represent the sequence number, and for each step k∈ Belongs to the collection , the DTW formula is as follows: ; Where t represents the number of nodes on the optimal path, represents the value of the kth dimension of the i-th data point in the time series P, represents the value of the kth dimension of the jth data point in the time series Q, is the Euclidean distance, which is a distance metric calculated as follows: ; (4.2) The weighting coefficient is obtained by normalizing the DTW distance. The calculation formula is as follows: ; in, represents the number of labeled target domains, r represents the rth labeled target domain, represents the labeled target domain data embedded by the backbone network, Represents the test data after being embedded in the backbone network.
5. The method for predicting the remaining useful life of a bearing in a multi-objective domain according to claim 4, characterized in that: In step S5, WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW to achieve feature-level transfer learning between the test data and multiple target domains. The calculation formula is as follows: ; in, Indicates the value of WMMD, () is a mapping function that maps data to a high-dimensional feature space. represents the norm of the Hilbert space.
6. A multi-objective domain bearing remaining service life prediction method as claimed in claim 4, characterized in that: In step S6, the hybrid target domain is defined as: ; in, Represents a mixed target domain.
7. The method for predicting the remaining useful life of a bearing in a multi-objective domain according to claim 5, characterized in that: The loss function formula between the true label and the target domain data label in the remaining useful life predictor model_2 is as follows: ; Where m represents the number of target domain data, represents the true label, Represents the predicted value obtained by predictor model_2.
8. The method for predicting the remaining useful life of a bearing in multiple objective domains according to claim 7, characterized in that: In step S7, based on the random freezing of some model parameters in step S3, other parameters use pre-trained values, and then fine-tune the model model_1 to complete the model-based transfer learning between the second-level source domain and the multi-target domains. The formula of the overall optimization function of multi-level domain transfer learning is as follows: ; Among them, λ and γ are regularization parameters.
9. The method for predicting the remaining useful life of a bearing in multiple objective domains according to claim 1, characterized in that: It also includes using root mean square error RMSE and mean absolute error MAE as evaluation indicators to evaluate the final prediction results.
10. A multi-objective domain bearing remaining service life prediction system, characterized in that: Includes the following modules: The data domain acquisition module is used to randomly select data under different working conditions from the industrial process bearing operation related data set as the source domain data Target domain data , and test data Test, ensure that the test data is in a different domain from the source domain data and in the same domain as part of the target domain training data; The source domain data preprocessing module is used to preprocess the source domain data in the offline stage. The patch obtained by padding segmentation is input into a 1D convolutional layer to obtain the input embedding ; Offline feature extraction module, using ModernTCN Training is performed to obtain a model model_1; the ModernTCN includes a stacked ModernTCN block, namely a backbone network and a linear head with a flat layer; The data relationship measurement module is used in the online stage to compare the test data Test and the target domain data Dynamic time warping (DTW) is used to align the sequences and calculate the distance, and the initial weight of the target domain is assigned according to the distance; A distribution difference calculation module is used to measure the distribution difference between the target domain and the test data after DTW processing by using the maximum mean difference MMD, and calculate the value of the initial weighted maximum mean difference WMMD; The remaining useful life prediction module is used to use the last two layers of model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life; The final prediction module is used to use the WMMD value and the sum of the mean square error between the predicted remaining useful life obtained through model_2 and the actual remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, distribution difference calculation module and remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value.
Citation Information
Patent Citations
Gas consumption behavior anomaly detection method and system based on transfer learning
CN115017970A
Equipment state evaluation method and system based on domain self-adaption and federated learning
CN116257972A
Remaining service life prediction method of multi-source domain transfer learning based on dynamic distribution self-adaption
CN116415485A
Method for predicting residual service life of rotary machine
CN119129364A
Vehicle power consumption prediction device and method using the same
KR1020220001915A
Cited By
Method and system for predicting residual life of alkaline electrolytic cell and medium
CN120874556A