A remaining service life prediction method for cross-operating condition scenario applications

By extracting and separating the degraded features of the source domain and the target domain, combining deep feature adaptive alignment and InfoNCE loss function, the problem of poor prediction accuracy of mechanical equipment in cross-working scenarios is solved, and higher prediction accuracy and feature representativeness are achieved.

CN118171240BActive Publication Date: 2025-06-13UNIV OF ELECTRONICS SCI & TECH OF CHINA ZHONGSHAN INST

Patent Information

Application Number
CN202410441652.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-06-13
Estimated Expiration
2044-04-12

AI Technical Summary

Technical Problem

The residual service life prediction accuracy of the prior art mechanical equipment in cross-working scenarios is poor, mainly because the private characteristics under different working conditions are ignored.

Method used

A method for predicting the residual service life across working conditions is proposed. By extracting the degraded features of the source domain and the target domain and separating them into private and common features, the InfoNCE loss function in deep feature adaptive alignment and comparison learning is used to ensure that the extracted features are representative.

Benefits of technology

By considering the private and common features between different domains, the prediction accuracy across operating conditions is improved, ensuring that the extracted features are representative, and significantly improving the performance of RUL prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118171240B_ABST
    Figure CN118171240B_ABST
Patent Text Reader

Abstract

The present invention discloses a remaining useful life prediction method for meeting cross-operating condition scenario applications, comprising the following steps: 1). Given existing source domain and target domain degradation data; 2). Normalize the collected data; 3). Extract the degradation features of the source domain and the target domain, and separate them into two parts: private features and common features; 4). Align the common features of the source domain and the target domain from two perspectives of global distribution and local distribution; 5). Use the InfoNCE loss function in contrastive learning to maximize the mutual information between the input data and the extracted features, ensuring that the extracted features are representative; 6). Regression prediction, using the model trained in the above steps to perform RUL prediction on the test data, greatly improving the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of prediction, and in particular to a remaining useful life prediction method that meets the application requirements of cross-operating conditions scenarios. Background Art

[0002] With the application of technologies such as artificial intelligence and machine learning, supported by a large amount of data and algorithms, the fault and prediction management system has developed rapidly, greatly improving the reliability of mechanical equipment operation and reducing maintenance costs. As one of the key issues of PHM, the prediction of remaining useful life has always been a research hotspot in PHM. The remaining useful life prediction technology uses the state monitoring data and fault mechanism of the equipment to establish a degradation model, analyze the degradation trend, and predict the remaining time from the current state to the fault threshold under the current operating conditions. This technology has greatly reduced the maintenance cost of the equipment and improved the reliability of equipment operation. However, the prediction accuracy of RUL is easily affected by various factors, such as uncertainty factors such as working environment and operating speed. Therefore, the RUL prediction of mechanical equipment has always been a challenging task.

[0003] In recent years, many RUL prediction methods have been proposed, which are mainly divided into three categories: model-based, data-driven, and the combination of both. Model-based methods establish the relationship between equipment failures and physical and chemical causes such as component wear through failure physics analysis and physical and chemical analysis, so as to predict the life of the equipment. However, such methods rely too much on powerful expert knowledge and it is difficult to obtain the mechanism model of the equipment. Data-driven methods use a large amount of historical data to extract the degradation characteristics of the equipment, establish a mapping relationship between the degradation characteristics and the remaining useful life, and fit the degradation curve to achieve the purpose of predicting RUL. The fusion method refers to the combination of model-driven and data-driven. Although it can make full use of the advantages of these two methods, the process is relatively complex, so this method is rarely used. Therefore, data-driven RUL prediction methods are becoming more and more popular because they can not only process a large amount of data, but also have low requirements for expert knowledge. The commonly used methods in data-driven methods are machine learning, deep learning, and transfer learning. The characteristic of machine learning that it can process data in batches has been widely applied in various fields. However, it is difficult for machine learning to capture complex relationships using shallow networks, and it requires manual extraction of important features, resulting in low generalization ability. Deep learning has achieved good performance in RUL prediction because it does not rely on manual feature selection, but there is still a problem: in actual industry, the same machine can work under different working conditions, but the labeled data can only be obtained under specific working conditions. When the working conditions of the test data and the training data are different, their data distributions are different, which will lead to poor RUL prediction accuracy of the trained model. To solve the above problems, transfer learning is introduced into RUL prediction. Transfer learning can transfer knowledge from the domain with sufficient labeled data to different but related domains with scarce labels, so RUL prediction under different working conditions can be realized. However, most of the current RUL predictions based on transfer learning are established on aligning the common feature distributions under different working conditions, ignoring the private features under different common working conditions. The present invention proposes a method for predicting the remaining useful life applied to cross-working-condition scenarios, which solves the problem of poor prediction accuracy caused by ignoring the private features under different common working conditions. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a remaining useful life prediction method that meets the cross-operating condition scenario application, which is used for a challenging but practical unsupervised RUL cross-domain prediction method for machines. This method can make full use of a large amount of source domain data for knowledge transfer under different operating conditions and does not require annotation of target domain data. The proposed method separates the extracted features into common features and private features, considering both the common features between different domains and the private features between different domains. The proposed method reduces the global and local distribution differences of common features by introducing deep feature adaptive alignment, and at the same time uses the InfoNCE loss function in contrastive learning to maximize the mutual information between the input data and the extracted features, ensuring that the extracted features are representative and greatly improving the prediction accuracy.

[0005] To solve the above problems, the present invention adopts the following technical solutions.

[0006] A remaining useful life prediction method that meets the cross-operating condition scenario application includes the following steps:

[0007] 1). Given existing source domain and target domain degradation data;

[0008] 2). Normalize the collected data;

[0009] 3). Extract the degradation features of the source domain and the target domain and separate them into two parts: private features and common features;

[0010] 4). Align the common features of the source domain and the target domain from two perspectives: global distribution and local distribution;

[0011] 5). Use the InfoNCE loss function in contrastive learning to maximize the mutual information between the input data and the extracted features to ensure that the extracted features are representative;

[0012] 6). Perform regression prediction, and use the model trained in the above steps to predict the RUL of the test data.

[0013] In the step 1), the existing source domain and target domain degradation data are given as shown in formula (1):

[0014]

[0015] The sensor readings use a sliding time window to convert the data into the time series input required by the model, where N tw is the time window size and d is the number of sensors.

[0016] In the step 2), due to the difference in the order of magnitude of the monitoring data of different sensors, feature selection and data normalization are performed on the data in the step 1), and the calculation formula is as follows:

[0017]

[0018] wherein, represents the j-th reading of the k-th selected sensor in the data sample, and max(x k ), min(x k ) are the maximum and minimum values respectively in the readings of the k-th selected sensor, is the normalized data sample.

[0019] The feature extraction and separation module in step 3): It is divided into a feature extraction module and a feature separation module, which are used to extract and separate the common features and private features from the source domain and target domain data through a common feature extractor and a private feature extractor;

[0020] Feature extraction: The function of this part is to extract degradation features through degraded data. A double-layer LSTM is used as the feature extractor. BiLSTM, that is, bidirectional long short-term memory network, consists of two independent LSTM networks and is interconnected at adjacent depths to achieve information extraction from two time directions. BiLSTM can improve the accuracy of machine remaining life prediction by capturing relevant information from both the past and the future simultaneously. When using BiLSTM for feature extraction, in BiLSTM, two independent hidden layers are used to process the sequence features in two directions, and its calculation formula is as follows:

[0021]

[0022] Feature separation: The function of this part is to separate the features extracted in the previous step into two parts: private features and common features. To ensure that the model can learn the common features and private features of the source domain and target domain, BiLSTM networks with the same number of layers in both the source domain and target domain are used as the common feature extractor and private feature extractor. Among them, the common feature extractors in the source domain and target domain have the same parameters. Let be the private feature extractors in the source domain and target domain respectively. After inputting the source domain data X S and the target domain data X T into the feature extractor, the features of each part are obtained, which are represented by the following formula:

[0023]

[0024]

[0025] wherein, represents the common features of the source domain and target domain, represents the private features of the source domain and target domain;

[0026] To make the common features and private features different, a soft subspace orthogonality constraint is added between them. Let be a matrix, and its row elements are The difference loss between the common features and the private features can be expressed as:

[0027]

[0028] denotes the second norm. In addition, there are differences between the common features and private features in the same domain, and there are also differences between the private features in different domains. The difference loss between the private features in the source domain and the target domain is added:

[0029]

[0030] The total difference loss can be expressed as:

[0031]

[0032] By minimizing L diff the features extracted by the common feature extractor and the private feature extractor can be constrained to be different, achieving the purpose of feature separation.

[0033] The deep feature adaptation module in step 4): The role of this part is to align the common features of the source domain and the target domain from two perspectives of the global distribution and the local distribution, and dynamically adjust the influence of the two distribution differences. Since the data distributions of different domains are different, only aligning the marginal distribution or the conditional distribution will result in the extracted degenerate features being difficult to align. In addition, when aligning the marginal distribution and the conditional distribution simultaneously, the reduction of the marginal difference and the conditional distribution difference is not equally important for the alignment of the degenerate features, and the weights of the two need to be dynamically adjusted. A deep feature adaptation module is used to adaptively align the global and local features of each domain pair from the perspectives of the marginal probability distribution and the conditional probability distribution. This method can adaptively adjust the importance of the marginal distribution and the conditional distribution in the distribution adaptation process according to the distribution of the actual degenerate data.

[0034] Global feature distribution alignment: The maximum mean discrepancy (MMD) is used to achieve global feature distribution alignment. The MMD metric can quantify the difference between the two domain distributions by subtracting the average value of the features after mapping to the reproducing kernel Hilbert space (RKHS). MMD can be expressed as:

[0035]

[0036] where K(·,·) represents the kernel function, and the Gaussian kernel function is selected in the present invention σ is the width of the kernel function to calculate MMD(Xs ,X t ).

[0037] Local feature distribution alignment: While reducing the global feature distribution difference, further calculate the conditional distribution difference and achieve the purpose of aligning the local feature distribution by minimizing this difference. For the specific implementation of the conditional distribution difference, first, the source domain degradation data samples are divided into 10 sub-domains according to the RUL label, and are sequentially labeled from 1 to 10. The local features can be interpreted as the features of each sub-domain. The sub-domain labels are constructed as follows,

[0038]

[0039] where y cls is the classification label, t r is the remaining useful life, t u is the used life, and T u is the maximum life cycle. This classification method takes the degradation data samples in the healthy state as one class, and divides the data samples in different degradation stages into 10 classes, with the degradation degree gradually increasing;

[0040] Since the target domain degradation data is unlabeled, during the training process, the model first uses the source domain pre-trained sub-domain classifier, and then the target domain data obtains sub-domain pseudo-labels through the classifier. As the iterative training progresses, the accuracy of the classifier gradually improves, so as to obtain accurate sub-domain labels. The cross-entropy loss function is used to calculate the sub-domain classification error of the source domain:

[0041]

[0042] where and respectively represent the true label and the classification label of the j-th source domain sample. According to the classification results of the sub-domains, the local feature distribution difference CMMD is calculated as follows,

[0043]

[0044] where C = {c|0, 1, 2,..., n cls} represents the sub-domain label of the degradation feature, and respectively represent the features belonging to the c sub-domain, and represent the conditional weights of the sub-domain feature distribution difference, and the calculation method is as follows,

[0045]

[0046] where and Represents the sub-domain label in one-hot encoding. One-hot encoding converts the sub-domain label into a binary form, increasing the dimension of the label;

[0047] Adaptive coefficient: Since the influence of the differences between the marginal distribution and the conditional distribution on the alignment of the degenerate feature distribution is different, and as the model is iteratively trained, the distribution difference between the source features and the target features will also change. To better adjust the ratio of the two distribution distances, it is necessary to design an adaptive coefficient to adaptively adjust the importance of the global and local feature differences in different training stages. The calculation method of the adaptive coefficient δ can be expressed as

[0048]

[0049]

[0050] where L MMD represents the measure of the global feature alignment state, and L CMMD represents the measure of the local feature alignment state. ε represents the error of the binary classifier based on the support vector when distinguishing whether the degenerate feature comes from the source domain or the target domain;

[0051] Combining formula (13) and formula (14), the objective function of DDA can be expressed as follows

[0052] L DDA =δL MMD +(1 - δ)L MMD (15)

[0053] It can be seen from equation (15) that the range of δ is [0, 1]. When δ is close to 0, the global feature difference is small, and the local feature difference should be mainly considered; when δ is close to 1, the local feature difference is small, and the global feature difference should be mainly considered. In short, the deep feature adaptive alignment mechanism can dynamically adjust the weights of the two according to the sizes of the global and local feature differences, and obtain cross-domain invariant degenerate feature representations.

[0054] The contrastive learning module in step 5):

[0055] To ensure that the extracted common features and private features are representative, the contrastive learning module applies the InfoNCE loss function in contrastive learning to maximize the mutual information between the extracted features and the input data. The meaning of mutual information is the amount of information about another random variable contained in a random variable. If the extracted features contain more information about the input data, that is, the mutual information is greater, then the representativeness of the extracted features is stronger;

[0056] For simplicity, the domain label of this subsection is ignored. When the input is X j ∈X, the common feature f can be obtained from the equationC,j and the private feature f P,j , since the feature is divided into the common feature f C,j and the private feature f P,j into two parts, the sum of the two is used as the input of the contrast metric module, that is:

[0057] f j = f C,j + f P,j (16)

[0058] To maximize the mutual information between the input data X j and the feature f j , a density ratio function μ j is defined as:

[0059]

[0060] To calculate the density ratio, the dimensions of the feature f j and the input data should be the same. Use a fully connected network θ of one layer to transform the dimension of the feature f j to the dimension of the input data. Use the transformed feature f θ,j = θ(f j ) to estimate the density ratio function by the dot product with the original input, which can be expressed as follows:

[0061]

[0062] To maximize the mutual information, the InfoNCE loss in contrastive learning is adopted. InfoNCE maximizes the mutual information between the feature and the input data by contrasting positive pairs and negative pairs. The calculation method of the InfoNCE loss is as follows:

[0063]

[0064] where (X j , f j ) represents a positive sample pair, and (X i , f j ) (i≠j) represents a negative sample pair. According to previous research, we can use the InfoNCE loss to represent the mutual information between the feature, f j and the input data:

[0065] I(X j , f j ) = log(N) - L InfoNCE (19)

[0066] where I(X j , f j ) represents the mutual information between the input data X jThe mutual information with its corresponding feature f j , where N represents the number of samples in a mini-batch. It can be seen from Equation (19) that by minimizing the InfoNCE loss, the mutual information I(X j , f j ) can be maximized, thus ensuring that the extracted features have sufficient representativeness. The regression prediction module in step 6):

[0067] The domain features are input into the label predictor P to generate the RUL prediction value. The root mean square error (RMSE) is used as the prediction loss function L label :[[]]

[0068]

[0069] The objective function of the model consists of three parts: the regression prediction loss L label , the dynamic distribution adaptation loss L DDA , the InfoNCE loss L InfoNCE , the sub-domain classification loss L cls , and the difference loss L diff ;

[0070] L(θ p , θ c , θ g ) = L label + L cls + αL InfoNCE + ξL DDA + βL diff (25)

[0071] In the formula, α and β are trade-off coefficients used to control the calculation ratio of L InfoNCE and L diff . ξ = 2 / (1 + e -10×(i+1) / epochs ) - 1 is a time-varying coefficient that changes with each training iteration. i is the current iteration number, and epochs is the total number of iterations.

[0072] Advantages of the present invention

[0073] Compared with the prior art, the advantages of the present invention are as follows:

[0074] The prior art does not consider the influence of private features on the alignment of common features between different domains. Since the feature separation method is adopted, the common features and private features are separated first and then aligned, reducing the interference of private features in the alignment of common features;

[0075] A deep feature adaptive alignment method is adopted, reducing the distribution difference between the source domain and the target domain from different angles, and greatly improving the alignment effect between the source domain and the target domain;

[0076] The contrastive learning method is introduced, which can ensure that the extracted features are representative and retain the degradation trends of different devices. Brief Description of the Drawings

[0077] Figure 1 It is a schematic flow diagram of the present invention.

[0078] Figure 2 It is the overall structure diagram of the present invention.

[0079] Figure 3 It is the structure diagram of BiLSTM of the present invention. Detailed Embodiments

[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0081] Please refer to Figures 1 to 3 , a remaining useful life prediction method that meets the application requirements of cross-operating conditions scenarios, including the following steps:

[0082] 1). Given existing degradation data of the source domain and the target domain;

[0083] 2). Normalize the collected data;

[0084] 3). Extract the degradation features of the source domain and the target domain and separate them into two parts: private features and common features;

[0085] 4). Align the common features of the source domain and the target domain from two perspectives: global distribution and local distribution;

[0086] 5). Use the InfoNCE loss function in contrastive learning to maximize the mutual information between the input data and the extracted features to ensure that the extracted features are representative;

[0087] 6). Perform regression prediction, and use the model trained in the above steps to predict the RUL of the test data.

[0088] In step 1) above, the existing degradation data of the source domain and the target domain are given, as shown in formula (1):

[0089]

[0090]

[0091] The sensor readings use a sliding time window to convert the data into the time series input required by the model, where N tw is the time window size and d is the number of sensors.

[0092] In step 2), due to the difference in the order of magnitude of the monitoring data of different sensors, feature selection and data normalization are performed on the data in step 1). The calculation formula is as follows:

[0093]

[0094] where represents the j-th reading of the k-th selected sensor in the data sample, and max(x k ) and min(x k ) are the maximum and minimum values of the readings of the k-th selected sensor respectively. is the normalized data sample.

[0095] The feature extraction and separation module in step 3): It is divided into a feature extraction module and a feature separation module, which are used to extract and separate the common features and private features from the source domain and target domain data through a common feature extractor and a private feature extractor;

[0096] Feature extraction: The role of this part is to extract degradation features through degraded data. A double-layer LSTM is used as the feature extractor. BiLSTM, that is, the bidirectional long short-term memory network, is composed of two independent LSTM networks and is interconnected at adjacent depths to achieve information extraction from two time directions. BiLSTM can improve the accuracy of machine remaining life prediction by capturing relevant information from both the past and the future. When using BiLSTM for feature extraction, in BiLSTM, two independent hidden layers are used to process the sequence features in two directions. The calculation formula is as follows:

[0097]

[0098]

[0099] Feature separation: The role of this part is to separate the features extracted in the previous step into two parts: private features and common features. To ensure that the model can learn the common features and private features of the source domain and target domain, BiLSTM networks with the same number of layers in the source domain and target domain are used as the common feature extractor and the private feature extractor. Among them, the common feature extractors of the source domain and target domain have the same parameters. Let be the private feature extractors of the source domain and target domain respectively. The source domain data X S and the target domain data X TAfter being input into the feature extractor, the features of each part are obtained, which are represented by the following formula:

[0100]

[0101] Among them, represents the common features of the source domain and the target domain, represents the private features of the source domain and the target domain;

[0102] To constrain the difference between the common features and the private features, a soft subspace orthogonality constraint is added between them. Let be a matrix, and its row elements are The difference loss between the common features and the private features can be expressed as:

[0103]

[0104] represents the second norm. In addition, there are differences between the common features and the private features of the same domain, and there are also differences between the private features of different domains. The difference loss between the private features of the source domain and the target domain is added:

[0105]

[0106] The total difference loss can be expressed as:

[0107]

[0108] By minimizing L diff the features extracted by the common feature extractor and the private feature extractor can be constrained to be different, and the purpose of feature separation can be achieved.

[0109] The deep feature adaptation module in step 4) above: The role of this part is to align the common features of the source domain and the target domain from two perspectives of the global distribution and the local distribution, and dynamically adjust the influence of the differences between the two distributions. Since the data distributions of different domains are different, only aligning the marginal distribution or the conditional distribution will result in the extracted degenerate features being difficult to align. In addition, when aligning the marginal distribution and the conditional distribution simultaneously, the reduction of the marginal difference and the conditional distribution difference is not equally important for the alignment of the degenerate features, and the weights of the two need to be dynamically adjusted. A deep feature adaptation module is used to adaptively align the global and local features of each domain pair from the perspectives of the marginal probability distribution and the conditional probability distribution. This method can adaptively adjust the importance of the marginal distribution and the conditional distribution in the distribution adaptation process according to the distribution of the actual degenerate data.

[0110] Global feature distribution alignment: The maximum mean discrepancy (MMD) is used to achieve global feature distribution alignment. The MMD metric can quantify the difference between two domain distributions by subtracting the mean of the features after mapping them to the reproducing kernel Hilbert space (RKHS). MMD can be expressed as:

[0111]

[0112] where \(K(\cdot,\cdot)\) represents the kernel function, and the Gaussian kernel function is selected in the present invention \(\sigma\) is the width of the kernel function) to calculate \(MMD(X s ,X t ).

[0113] Local feature distribution alignment: While reducing the global feature distribution difference, the conditional distribution difference is further calculated, and the local feature distribution is aligned by minimizing this difference. For the specific implementation of the conditional distribution difference, first, the source domain degradation data samples are divided into 10 sub-domains according to the RUL label, and are sequentially labeled from 1 to 10. The local features can be interpreted as the features of each sub-domain. The sub-domain labels are constructed as follows,

[0114]

[0115] where \(y cls is the classification label, \(t r is the remaining useful life, \(t u is the used life, and \(T u is the maximum life cycle. This classification method takes the degradation data samples in the healthy state as one class, and divides the data samples in different degradation stages into 10 classes, with the degradation degree increasing gradually;

[0116] Since the target domain degradation data is unlabeled, during the training process, the model first uses the source domain pre-trained sub-domain classifier, and then the target domain data obtains sub-domain pseudo-labels through the classifier. As the iterative training progresses, the accuracy of the classifier gradually improves, so as to obtain accurate sub-domain labels. The cross-entropy loss function is used to calculate the sub-domain classification error of the source domain:

[0117]

[0118] where and respectively represent the true label and the classification label of the \(j\)-th source domain sample. According to the classification results of the sub-domains, the local feature distribution difference \(CMMD\) is calculated as follows,

[0119]

[0120]

[0121] Among them, C = {c|0, 1, 2,..., n cls} represents the sub-domain label of the degenerate feature, and respectively represent the features belonging to the c sub-domain, and represent the conditional weights of the sub-domain feature distribution differences, and the calculation method is as follows,

[0122]

[0123] Among them, and represent the one-hot encoded sub-domain labels. The one-hot encoding converts the sub-domain labels into binary form, increasing the dimension of the labels;

[0124] Adaptive coefficient: Since the influence of the marginal distribution and conditional distribution differences on the alignment of the degenerate feature distribution is different, and as the model is iteratively trained, the distribution difference between the source features and the target features will also change. In order to better adjust the ratio of the two distribution distances, it is necessary to design an adaptive coefficient to adaptively adjust the importance of the global and local feature differences at different training stages. The calculation method of the adaptive coefficient δ can be expressed as,

[0125]

[0126]

[0127] Among them, L MMD represents the measure of the global feature alignment state, and L CMMD represents the measure of the local feature alignment state. ε represents the error of the binary classifier based on the support vector in distinguishing whether the degenerate feature comes from the source domain or the target domain;

[0128] Combining formula (13) and formula (14), the objective function of DDA can be expressed as follows,

[0129] L DDA = δL MMD +(1 - δ)L MMD (15)

[0130] It can be seen from formula (15) that the range of δ is [0, 1]. When δ is close to 0, the global feature difference is small, and the local feature difference should be mainly considered; when δ is close to 1, the local feature difference is small, and the global feature difference should be mainly considered. In short, the deep feature adaptive alignment mechanism can dynamically adjust the weights of the two according to the sizes of the global and local feature differences, and obtain cross-domain invariant degenerate feature representations.

[0131] The contrastive learning module in step 5) described above:

[0132] To ensure that the extracted common features and private features are representative, the contrastive learning module applies the InfoNCE loss function in contrastive learning to maximize the mutual information between the extracted features and the input data. The meaning of mutual information is the amount of information contained in one random variable about another random variable. If the extracted features contain more information of the input data, that is, the mutual information is greater, then the representativeness of the extracted features is stronger;

[0133] For simplicity, the domain labels of this subsection are ignored. When the input is X j ∈ X, the common feature f C,j and the private feature f P,j can be obtained from the equation. Since the features are divided into the common feature f C,j and the private feature f P,j two parts, the sum of the two is used as the input of the contrast metric module, that is:

[0134] f j = f C,j + f P,j (16)

[0135] To maximize the mutual information between the input data X j and the feature f j , a density ratio function μ j :

[0136]

[0137] To calculate the density ratio, the dimensions of the feature f j and the input data should be the same. Use a fully connected network θ to transform the dimension of the feature f j to the dimension of the input data. Use the dot product between the transformed feature f θ,j = θ(f j ) and the original input to estimate the density ratio function, which can be expressed as follows:

[0138]

[0139] To maximize the mutual information, the InfoNCE loss in contrastive learning is adopted. InfoNCE maximizes the mutual information between the feature and the input data by contrasting positive pairs and negative pairs. The calculation method of the InfoNCE loss is as follows:

[0140]

[0141] where (X j , f j ) represents the positive sample pair, (Xi, f j)(i≠j) represents negative sample pairing. According to previous research, we can use the InfoNCE loss to represent the mutual information between the feature, f j and the input data:

[0142] I(X j , f j ) = log(N) - L InfoNCE (19)

[0143] where I(X j , f j ) represents the mutual information between the input data X j and its corresponding feature f j . N represents the number of samples in a mini - batch. It can be seen from Equation (19) that by minimizing the InfoNCE loss, the mutual information I(X j , f j ) can be maximized, thus ensuring that the extracted features are sufficiently representative. The regression prediction module in step 6):

[0144] The domain features are input into the label predictor P to generate the RUL prediction value. The root mean square error (RMSE) is used as the prediction loss function L label :

[0145]

[0146] The objective function of the model consists of three parts: the regression prediction loss L label , the dynamic distribution adaptation loss L DDA , the InfoNCE loss L InfoNCE , the sub - domain classification loss L cls , and the difference loss L diff ;

[0147] L(θ p , θ c θ g ) = L label + L cls + αL InfoNCE + ξL DDA + βL diff (25)

[0148] In the formula, α and β are trade - off coefficients used to control the calculation ratio of L InfoNCE and L diff . ξ = 2 / (1 + e -10×(i+1) / epochs ) - 1 is a time - varying coefficient that changes with each training iteration. i is the current iteration number, and epochs is the total number of iterations.

[0149] In summary, the present invention proposes a remaining useful life prediction method applied to cross-operating conditions scenarios. The source domain and target domain data extract common features and private features through a common feature extractor and a private feature extractor, and impose orthogonality constraints on the common features and private features to achieve the effect of feature separation. Then, the global features and local features of the source domain and target domain are aligned by introducing deep feature adaptive distribution. To ensure that the extracted features are representative, the InfoNCE loss function in contrastive learning is used to maximize the mutual information between the input data and the extracted features, improving the performance of RUL prediction.

[0150] As described above, the above is only a preferred specific embodiment of the present invention; however, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications, characterized in that: The following steps are involved: 1) Given the existing source domain and target domain degradation data; 2). Normalize the collected data; 3) Extract the degraded features of the source domain and the target domain, and separate them into private features and common features; 4) Align the common features of the source domain and the target domain from the perspectives of global distribution and local distribution; 5) Adopt the InfoNCE loss function in contrastive learning to maximize the mutual information between the input data and the extracted features to ensure that the extracted features are representative; 6). Regression prediction: Use the model trained in the above steps to predict RUL of the test data; A deep feature adaptation module is used to adaptively align the global and local features of each domain pair from the perspectives of marginal probability distribution and conditional probability distribution. Global feature distribution alignment: The maximum mean difference (MMD) is used to achieve global feature distribution alignment. The MMD metric can quantify the difference between the distributions of two domains by subtracting the mean value of the features after mapping to the reproducing kernel Hilbert space (RKHS). MMD can be expressed as: Among them, K(·,·) represents the kernel function, and the Gaussian kernel function ( σ is the width of the kernel function) to calculate MMD(X s ,X t ); Local feature distribution alignment: While the global feature distribution difference is reduced, the conditional distribution difference is further calculated, and the purpose of aligning the local feature distribution is achieved by minimizing the difference. For the specific implementation of the conditional distribution difference, first, the source domain degraded data samples are divided into 10 subdomains according to the RUL label, and are labeled from 1 to 10 in sequence. The local features can be interpreted as the features of each subdomain. The subdomain label is constructed as follows: Among them, y cls is the classification label, t r is the remaining useful life, t u is the service life, T u is the maximum life cycle. This classification method takes the degraded data samples in the healthy state as one category and divides the data samples in different degradation stages into 10 categories, with the degradation degree gradually increasing; The local feature distribution difference CMMD is calculated as follows, Where C = {c|0,1,2,...,n cls } represents the subdomain label of the degenerate feature, and They represent the features belonging to subdomain c, and The conditional weight representing the difference in subdomain feature distribution is calculated as follows, in, and Represents the subdomain label of one-hot encoding. One-hot encoding converts the subdomain label into binary form, which increases the dimension of the label; Adaptive coefficient: Since the difference between edge distribution and conditional distribution has different effects on the alignment of degenerate feature distribution, and the distribution difference between source and target features will change with the iterative training of the model, in order to better adjust the ratio of the two distribution distances, it is necessary to design an adaptive coefficient to adaptively adjust the importance of global and local feature differences at different training stages. The calculation method of the adaptive coefficient δ can be expressed as, Among them, L MMD The metric representing the global feature alignment status, L CMMD It represents the metric of the local feature alignment state, and ε represents the error of the support vector-based binary classifier in distinguishing whether the degraded features come from the source domain or the target domain; Combining formula (13) and formula (14), the objective function of DDA can be expressed as follows: L DDA =d LMMD +(1-δ)L MMD (15)。 2. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 1, characterized in that: In step 1), the existing source domain and target domain degradation data are given, as shown in formula (1): The sensor readings are converted into the time series input required by the model using a sliding time window, where N tw is the time window size, and d is the number of sensors.

3. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 2, characterized in that: In step 2), since the monitoring data of different sensors have different orders of magnitude, feature selection and data normalization are performed on the data in step 1), and the calculation formula is as follows: in, represents the jth reading of the kth selected sensor in the data sample, max(x k )、min(x k ) are the maximum and minimum values ​​in the k-th selected sensor readings, respectively, for Normalized data samples.

4. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 3, characterized in that: The feature extraction and separation module in step 3) is divided into two modules, feature extraction and feature separation, for extracting and separating common features and private features from source domain and target domain data through a common feature extractor and a private feature extractor; feature Extraction: The function of this part is to extract degradation features through degradation data. A two-layer LSTM is used as a feature extractor. BiLSTM, or bidirectional long short-term memory network, consists of two independent LSTM networks, which are interconnected at adjacent depths to achieve information extraction from two time directions. BiLSTM improves the accuracy of machine remaining life prediction by simultaneously capturing relevant information from the past and the future. In BiLSTM, two independent hidden layers are used to process sequence features in two directions. The calculation formula is as follows: Feature separation: This part is used to separate the features extracted in the previous step into two parts, private features and common features. In order to ensure that the model can learn the common features and private features of the source domain and the target domain, the BiLSTM network with the same number of layers in the source domain and the target domain is used as the common feature extractor and the private feature extractor. The parameters are the same, are private feature extractors for the source domain and the target domain respectively, and the source domain data X S and target domain data X T After inputting into the feature extractor, the features of each part are obtained, which can be expressed by the following formula: in, represents the common features of the source domain and the target domain, Represents the private features of the source and target domains; In order to constrain the common features and private features to be different, a soft subspace orthogonality constraint is added between them. is a matrix whose row elements are The difference loss between common features and private features is expressed as: In addition, there are differences between the common features and private features of the same domain, and there are differences between the private features of different domains. The difference loss between the private features of the source domain and the target domain is added: The total discrepancy loss is expressed as: By minimizing L diff The features extracted by the common feature extractor and the private feature extractor are constrained to be different, so as to achieve the purpose of feature separation.

5. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 4, characterized in that: The deep feature adaptation module in step 4): the function of this part is to align the common features of the source domain and the target domain from the perspectives of global distribution and local distribution, and dynamically adjust the impact of the differences between the two distributions. Due to the differences in data distribution in different domains, aligning only the edge distribution or the conditional distribution will make it difficult to align the extracted degraded features. In addition, when aligning the edge distribution and the conditional distribution at the same time, the reduction of the edge difference and the conditional distribution difference is not equally important for the alignment of the degraded features, and the weights of the two need to be dynamically adjusted. This method can adaptively adjust the importance of the edge distribution and the conditional distribution in the distribution adaptation process according to the distribution of the actual degraded data; Since the target domain degraded data is unlabeled, during the training process, the model first uses the source domain to pre-train the subdomain classifier, and then the target domain data obtains the subdomain pseudo label through the classifier. With iterative training, the classifier accuracy gradually improves, thereby obtaining accurate subdomain labels, and the cross entropy loss function is used to calculate the subdomain classification error of the source domain: in and They represent the true label and classification label of the j-th source domain sample, respectively, according to the classification results of the subdomain; It can be seen from formula (15) that the range of δ is [0,1]. When δ is close to 0, the global feature difference is small, and the local feature difference should be focused on; when δ is close to 1, the local feature difference is small, and the global feature difference should be focused on. In short, the deep feature adaptive alignment mechanism can dynamically adjust the weights of the global and local features according to the difference between the two, and obtain cross-domain invariant degradation feature representation.

6. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 5, characterized in that: The contrastive learning module in step 5): In order to ensure that the extracted common features and private features are representative, the contrastive learning module applies the InfoNCE loss function in contrastive learning to maximize the mutual information between the extracted features and the input data. The mutual information means the amount of information contained in one random variable about another random variable. If the extracted features contain more information about the input data, that is, the greater the mutual information, the more representative the extracted features are. For simplicity, the field labels in this section are ignored. j ∈X, the common feature f can be obtained from the equation C,j and private feature f P,j , since the features are divided into common features f C,j and private feature f P,j The two parts are added together as the input of the contrast measurement module, which is: f j =f C,j +f P,j (16) To maximize the input data X j and feature f j The mutual information between them defines a density ratio function μ j : To calculate the density ratio, the feature f j The dimension should be the same as the input data. Use a fully connected network θ to transform the feature f j The dimension of the input data is transformed to the dimension of the input data, and the transformed feature f is used θ,j =θ(f j ) and the original input to estimate the density ratio function, which can be expressed as follows: In order to maximize the mutual information, the InfoNCE loss in contrastive learning is used. InfoNCE maximizes the mutual information between features and input data by comparing positive and negative pairs. The InfoNCE loss is calculated as follows: Where (X j ,f j ) represents the positive sample pairing, (X i ,f j )(i≠j) represents negative sample pairing, and uses InfoNCE loss to represent features, f j Mutual information with input data: I(X j ,f j )=log(N)-L InfoNCE (19) Among them, I(X j ,f j ) represents the input data X j Its corresponding feature f j The mutual information of N represents the number of samples in a mini-batch. From formula (19), we can see that the mutual information I(X) can be maximized by minimizing the InfoNCE loss. j ,f j ), thereby ensuring that the extracted features are representative enough.

7. A method for predicting remaining useful life that meets the requirements of cross-operating scenario applications according to claim 6, characterized in that: The regression prediction module in step 6): The domain features are input into the label predictor P to generate RUL prediction values, and the root mean square error is used as the prediction loss function L label : The objective function of the model consists of three parts: regression prediction loss L label , Dynamic distribution adaptive loss L DDA , InfoNCE loss L InfoNCE , subdomain classification loss L cls , difference loss L diff ; L(θ p ,i c ,i g )=L label +L cls +αL InfoNCE +ξL DDA +βL diff (25) Where α and β are trade-off coefficients used to control L InfoNCE and L diff The calculation ratio is ξ=2 / (1+e -10×(i+1) / epochs )-1 is the time-varying coefficient that changes with each training iteration, i is the current iteration number, and epochs is the total number of iterations.

Citation Information

Patent Citations

  • Bearing fault intelligent diagnosis method and diagnosis system, computer equipment and medium

    CN112633339A

  • Unsupervised cross-domain prediction method for residual service life of aero-engine

    CN116502123A

Cited By

  • Ship engine component cross-working-condition life prediction method and system

    CN121301790A