A Fault Diagnosis Method for Rotating Components Based on Co-Training
Through collaborative training and overall loss function optimization, the problem of identifying unknown fault types in unknown working conditions is solved, cross-condition generalization and efficient fault diagnosis without target working conditions are achieved, and the intelligent operation and maintenance capabilities of rotating equipment are improved.
Patent Information
- Application Number
- CN202510542963.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art cannot effectively deal with the identification of unknown fault types of rotating equipment under unknown operating conditions, especially in complex operating conditions such as variable speed and sudden load, which leads to the emergence of new unknown fault modes. In addition, traditional methods require target operating conditions data to participate in training, and it is impossible to achieve cross-condition generalization and unknown fault identification without target operating conditions data.
Using a rotating component fault diagnosis method based on collaborative training, an overall loss function is constructed through fault type grouping and vibration data collaborative gradient update strategy, combining auxiliary binary classifiers and multi-classifier losses, model parameters are optimized to adapt to unknown working conditions and unknown fault types, and reduce data acquisition costs.
It realizes accurate identification of unknown fault types under unknown working conditions, improves the intelligent operation and maintenance capabilities of rotating equipment, reduces data acquisition costs, improves the accuracy and reliability of diagnosis, and solves the problem of fuzzy classification boundary problems in complex working conditions by traditional methods.
Smart Images

Figure CN120067919B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of specific computing models, and particularly relates to a fault diagnosis method for rotating components based on co-training. Background Art
[0002] In the operation and maintenance of industrial equipment, rotating equipment (such as centrifugal pump units) generally serves as the core power equipment for oil and gas transportation, chemical processes, and municipal water supply systems. The fault diagnosis of rotating components (such as bearings, impellers, couplings, etc.) is a key link to ensure the safe operation of rotating equipment.
[0003] With the intelligent development of industrial equipment, fault diagnosis methods based on deep learning have gradually matured. Existing methods mainly rely on deep learning models to identify fault types by training classifiers through vibration data. However, existing methods rely on training data with known working conditions and known fault types and cannot effectively handle newly emerging fault types under unknown working conditions (such as sudden load changes, speed fluctuations). Patent CN118643324A improves the cross-condition generalization ability through domain adaptation, but it still requires target condition data to participate in training and cannot identify unknown fault types. Patent CN118643346B introduces uncertainty measurement to improve the ability to identify unknown faults, but it also requires target condition data to participate in training and does not solve the problem of unknown fault diagnosis under unknown working conditions.
[0004] In the actual operation of rotating equipment, frequent speed changes (such as dynamic adjustment of pump groups at 2950 - 3500 rpm), sudden load changes (such as sudden increase in the viscosity of chemical media), etc., combined with the change of shafting dynamic characteristics caused by equipment aging (such as the coupling of abnormal shaft deflection and early bearing faults), often lead to the sudden emergence of new unknown fault modes (such as adhesive wear of sliding bearings, microcrack propagation of thrust bearings, low-temperature shrinkage leakage of mechanical seals). Therefore, the fault diagnosis of rotating components under unknown working conditions and unknown faults becomes more important. Although Patent CN117786489A improves the generalization through co-training of multi-sensor data, it does not solve the problem of blurred classification boundaries under the combined action of working condition shift and fault type shift, which will cause a certain degree of misclassification for the identification of unknown fault types under unknown working conditions.
[0005] To solve the above problems, there is an urgent need for a fault diagnosis method that does not require target condition data to participate in training and simultaneously has cross-condition generalization ability and unknown fault identification ability. Summary of the Invention
[0006] The present invention provides a fault diagnosis method for rotating components based on co-training to solve the problem of joint identification of unknown working conditions and unknown fault types of rotating components, improve the diagnostic accuracy and reliability, and realize the intelligent operation and maintenance of rotating equipment under complex working conditions.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A fault diagnosis method for rotating components based on co-training, comprising:
[0009] Obtaining a first fault type group and a second fault type group according to the fault type, and combining with the collected model training data set to obtain a first meta-training set and a second meta-training set, wherein the model training data set contains vibration data of the rotating component under multiple working conditions and multiple fault types;
[0010] Based on the preset model parameters, according to the first meta-training set, obtaining an auxiliary binary classification loss through an auxiliary binary classifier, obtaining a multi-classifier loss through a cross-entropy function, and combining to obtain an overall loss;
[0011] According to the overall loss, updating the model parameters, combining with the second meta-training set to obtain an overall loss, and obtaining an objective function according to the overall losses obtained from the first meta-training set and the second meta-training set for training the rotating component fault diagnosis model.
[0012] Obtaining a first fault type group and a second fault type group according to the fault type, specifically:
[0013] Obtaining a first fault type group and a second fault type group through relevance analysis according to the characteristics of the fault type; and / or obtaining a first fault type group and a second fault type group according to the vibration intensity corresponding to the fault type.
[0014] The first meta-training set and the second meta-training set are specifically:
[0015] Both the first meta-training set and the second meta-training set contain vibration data under multiple working conditions;
[0016] The vibration data of the first meta-training set under at least one working condition belongs to one of the first fault type group and the second fault type group, and the vibration data under at least another working condition belongs to the other of the first fault type group and the second fault type group;
[0017] The fault type group to which the vibration data of the second meta-training set under the corresponding working condition belongs is different from that of the first meta-training set.
[0018] The overall loss is specifically:
[0019] Combining the multi-classifier loss and the auxiliary binary classification loss to obtain the overall loss;
[0020]
[0021] Wherein, represents the overall loss, represents the multi-classifier loss, represents the auxiliary binary classification loss.
[0022] Update the model parameters according to the overall loss, specifically:
[0023] Update the model parameters according to the gradient of the overall loss obtained from the first meta-training set under the preset model parameters, combined with the set learning rate;
[0024]
[0025] where, and represent the vibration data of the first meta-training set, represents that the vibration data corresponding to the first working condition in the first meta-training set belongs to the first fault type group, represents that the vibration data corresponding to the second working condition in the first meta-training set belongs to the second fault type group; and respectively represent and the gradients of the overall losses; represents the set learning rate; represents the preset model parameters; represents the updated model parameters.
[0026] Obtain an objective function for rotating component fault diagnosis according to the overall losses obtained from the first meta-training set and the second meta-training set, specifically:
[0027]
[0028] where, and represent the vibration data of the second meta-training set, represents that the vibration data corresponding to the first working condition in the second meta-training set belongs to the second fault type group, represents that the vibration data corresponding to the second working condition in the second meta-training set belongs to the first fault type group; and respectively represent and the overall losses under the preset model parameters; and respectively represent and the overall losses under the updated model parameters.
[0029] The method for rotating component fault diagnosis based on co-training further includes:
[0030] Based on the collected vibration data, through preprocessing, a model training data set is obtained; wherein, the preprocessing at least includes data anomaly detection, sliding window segmentation, and time-frequency feature conversion processing, and the data anomaly detection is used to identify and eliminate abnormal data.
[0031] The sliding window segmentation is specifically as follows:
[0032] Based on the vibration data after anomaly elimination, through a non-overlapping sliding window segmentation method, a plurality of vibration data segments are obtained; wherein the window step size is determined according to the vibration data acquisition length, acquisition frequency, and data vibration period.
[0033] The time-frequency feature conversion processing is specifically as follows:
[0034] Based on the vibration data after sliding window segmentation processing, the negative frequency part in the signal is removed through an instantaneous feature extraction method;
[0035] Based on the processed vibration data, through a time-frequency analysis method, the instantaneous power spectral density is obtained;
[0036] Wherein, the frequency resolution of the time-frequency analysis method is determined according to the highest frequency of the collected vibration data, and the time resolution is determined according to the minimum bandwidth of the collected vibration data.
[0037] The rotating component fault diagnosis model is specifically as follows:
[0038] Based on the vibration data of the rotating component collected in real time, a confidence score is obtained through an auxiliary binary classifier in the rotating component fault diagnosis model, and the known fault type and unknown fault type are judged through the confidence score;
[0039] For the known fault type, the probability distribution of the known fault type is obtained through a multi-classifier in the rotating component fault diagnosis model to diagnose the rotating component fault.
[0040] Due to the adoption of the above technical solution, the beneficial effects obtained by the present invention are as follows:
[0041] 1. In the present invention, according to the fault type, a first fault type group and a second fault type group are obtained, and combined with the collected model training data set, a first meta-training set and a second meta-training set are obtained, wherein the model training data set contains the vibration data of the rotating component under multiple working conditions and multiple fault types. By using the collaborative gradient update strategy of the working condition and the fault type, after the model is trained with the known working condition data, it can still effectively capture the feature commonalities under unknown working conditions. That is to say, the model can adapt to unknown working conditions without relying on the target working condition data, and solves the problem of feature space mismatch caused by working condition offset in the traditional method.
[0042] In addition, an auxiliary binary classification loss is obtained through an auxiliary binary classifier, a multi-classifier loss is obtained through a cross-entropy function, and the two are combined to obtain an overall loss. By introducing the auxiliary binary classifier, a clear classification boundary is generated by distinguishing known and unknown fault types (determined by a confidence threshold). An overall loss function is constructed by combining the cross-entropy loss and the auxiliary binary classification loss to suppress the interference of unknown faults on known classifications. This enables the model to recognize unknown fault types that do not appear in the training data, avoiding misjudgment problems of traditional closed-set classifiers.
[0043] A target function is obtained based on the overall loss obtained from the first meta-training set and the second meta-training set for training a rotating component fault diagnosis model. A gradient matching mechanism is adopted to ensure that the parameter update directions at the working condition level and the fault type level are consistent, avoiding adversarial interference in cross-working condition training. The model's ability to represent fine-grained features is optimized, reducing feature confusion between known and unknown faults.
[0044] Therefore, the model only needs to be trained with multi-fault type data under known working conditions, without covering all possible combinations of working conditions, significantly reducing the data acquisition cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0046] Figure 1 is a schematic flow chart of the method for diagnosing faults of rotating components based on co-training according to an embodiment of the present invention;
[0047] Figure 2 is a logical architecture diagram of the method for diagnosing faults of rotating components based on co-training according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to more clearly illustrate the overall concept of the present invention, the following will be described in detail by way of examples in conjunction with the drawings of the specification.
[0049] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.
[0050] As Figure 1 shown, a method for diagnosing faults of rotating components based on co-training includes:
[0051] S100: Obtain the first fault type group and the second fault type group according to the fault type, and combine the collected model training data set to obtain the first meta-training set and the second meta-training set, where the model training data set includes vibration data under multiple working conditions and multiple fault types of the rotating component.
[0052] The core objective of this step is to construct a training data set that can adapt to unknown working conditions and unknown fault types. Through the collaborative design of fault type grouping and working condition conditions, the cross-working condition generalization ability of subsequent model training is improved to ensure that the model can still accurately identify known fault types under unknown working conditions; the unknown fault identification ability, by suppressing the interference of unknown faults on known classifications through the grouping strategy, generates clear classification boundaries; the data efficient utilization ability, only known working condition data is required to train the model, reducing the dependence on target working condition data.
[0053] It can be understood that the model training data set includes vibration data under multiple working conditions and multiple fault types of the rotating component, providing the model with the joint feature distribution of working conditions and fault types to support cross-domain generalization. The meta-training set is a training set containing data of known fault types under multiple working conditions, which is used to update the initial model parameters. The gradient calculation is used to guide the optimization direction of the model parameters to ensure the collaborative feature learning of working conditions and fault types.
[0054] According to the fault type grouping, in the present invention, based on the feature similarity of the fault type (such as the vibration frequency spectrum difference between the inner ring damage and the outer ring damage of the bearing), the fault types are divided into two groups (such as the first group is the rolling element fault and the second group is the lubrication fault) through statistical methods (such as the Pearson correlation coefficient).
[0055] Similarly, the fault types can be divided into a high-risk group and a low-risk group according to the vibration energy intensity corresponding to the fault type (such as high-amplitude faults and low-amplitude faults). In a specific embodiment, in the bearing fault of a centrifugal pump, the inner ring damage (high-frequency impact feature) and the grease contamination (low-frequency modulation feature) may be divided into different groups. This step reduces the feature confusion of fault types within the group through grouping, providing a distinguishable feature space for subsequent gradient matching.
[0056] According to the fault type grouping, the model training data set is divided into the first meta-training set and the second meta-training set. Both the first meta-training set and the second meta-training set include vibration data under multiple working conditions, and for the same working condition, the vibration data corresponding to the first meta-training set and the second meta-training set belong to different fault type groups. Ensure that there are differences in both working conditions and fault types between the first meta-training set and the second meta-training set, simulating unknown fault scenarios under unknown working conditions.
[0057] Through the collaborative training of the first meta-training set and the second meta-training set, when updating the model parameters, the model needs to adapt to both the working condition shift and the fault type shift simultaneously, significantly improving the cross-domain generalization ability. It should be noted that both the first meta-training set and the second meta-training set need to include multiple working conditions (such as rotational speeds of 735 RPM, 1102.5 RPM, and 1470 RPM) to ensure that the model learns the working condition invariance features. The fault type groups to which the vibration data of the first meta-training set and the second meta-training set belong are different under the same working condition, avoiding the overlap of feature distributions. This enables the model to focus on the joint features of the working condition and the fault type, reducing the misjudgment of known faults under unknown working conditions. The first meta-training set and the second meta-training set constructed in this step provide a data basis for the subsequent update of the model parameters, ensuring that the feature representations of the working condition and the fault type are optimized synchronously when updating the model parameters.
[0058] Generally speaking, in this step, through fault type grouping and meta-training set design, a training data framework that supports generalization across working conditions and fault types is constructed, providing a key data basis for the collaborative training of the model, solving the bottleneck problem that traditional methods cannot identify unknown faults under unknown working conditions, and significantly improving the practicality and reliability of rotating component fault diagnosis.
[0059] S200: Based on the preset model parameters, according to the first meta-training set, obtain the auxiliary binary classification loss through the auxiliary binary classifier, obtain the multi-classifier loss through the cross-entropy function, and combine them to obtain the overall loss.
[0060] The core objective of this step is to improve the classification accuracy of known faults and the recognition ability of unknown faults under unknown working conditions by jointly optimizing the multi-classification task and the auxiliary binary classification task. Specifically, the multi-classification task optimizes the classification accuracy of known fault types through cross-entropy loss. The auxiliary binary classification task uses the auxiliary binary classifier to distinguish between known faults and unknown faults, generating a clear classification boundary and suppressing the interference of unknown faults on known classifications. The overall loss function combines the gradient information of the two losses to synchronously optimize the model parameters and ensure the consistency of the feature representations of the working condition and the fault type.
[0061] For the binary classifier, input the vibration data of the first meta-training set, and output a binary judgment (known fault or unknown fault) through the auxiliary binary classifier. The auxiliary binary classification loss function is as follows:
[0062]
[0063] Among them, for the input sample , the probability that the model outputs the correct category , for all non-correct categories , calculate their predicted probabilities , and take the one corresponding to the minimum value among them. Combine the two items to generate an auxiliary binary classification loss.
[0064] By maximizing the prediction probability of the correct class to reduce missed judgments. Minimize the confidence of the model in the most confident incorrect class to avoid misjudgments (such as misjudging grease contamination as bearing wear). Through the dual-objective optimization, the model forms a clearer decision boundary between known / unknown faults or critical classes.
[0065] For the multi-classifier, input the vibration data of the first primary training set, and output the probability distribution of known fault types (such as inner ring damage, outer ring damage of the bearing, etc.) through the multi-classifier. The loss of the multi-classifier is as follows,
[0066]
[0067] where the model outputs the prediction probability of all classes for the sample Output all classes of the prediction probability . According to the true label (one-hot encoding), calculate the difference between the model prediction and the true distribution.
[0068] In this step, by minimizing the cross-entropy, the prediction probability distribution of the model for known fault types is closer to the true label, improving the classification accuracy of known faults. The overall loss is obtained by combining the auxiliary binary classification loss and the multi-classifier loss. Compared with only using the cross-entropy loss, the overall loss function improves the classification accuracy of known faults. Moreover, by jointly optimizing the two types of tasks, the recognition ability of known and unknown faults is improved synchronously. The overall loss in this step provides a direction for the parameter optimization of the subsequent second primary training set, and finally forms an objective function for the hierarchical coordination of working conditions and fault types.
[0069] Generally speaking, in this step, through the collaborative optimization of the auxiliary binary classification loss and the multi-classifier loss, the problems of fuzzy classification boundaries and overfitting in traditional methods under complex working conditions are solved, providing an efficient and robust model training method for the fault diagnosis of rotating components.
[0070] S300: Update the model parameters according to the overall loss, combine the second primary training set to obtain the overall loss, and obtain the objective function according to the overall losses obtained from the first primary training set and the second primary training set for the training of the rotating component fault diagnosis model.
[0071] The core objective of this step is to optimize the model parameters through a collaborative training strategy, enabling the model to have stronger generalization ability under unknown working conditions and unknown fault types. Specifically, by using the loss functions of the first meta-training set and the second meta-training set, the model parameters are synchronously optimized to achieve cross-condition parameter updates. The losses of the two meta-training sets are jointly formed into an objective function to ensure that the model can adapt to both working condition shifts and fault type shifts simultaneously. Here, the objective function refers to the loss function ultimately used for model training, which synthesizes the task requirements of multiple working conditions and multiple fault types, and is used to guide the model to learn the features of working condition invariance and fault type discrimination.
[0072] Based on the first meta-training set, the overall loss is calculated through the obtained overall loss function (including the auxiliary binary classification loss and the cross-entropy loss). The parameter gradients are calculated according to the overall loss, and the model parameters are updated. It should be noted that the parameter update directions of the working condition and the fault type are ensured to be consistent through a regularization term (such as maximizing the dot product of gradients).
[0073] In this step, the features learned by the model under known working conditions (such as bearing vibration spectra) are generalized to unknown working conditions (such as scenarios with sudden changes in rotational speed), and the known fault classification accuracy is obtained. Moreover, the gradient direction consistency constraint reduces adversarial updates and improves the model convergence speed.
[0074] The updated model parameters are applied to the second meta-training set (including different fault types under the same working condition as the first meta-training set), and its overall loss is calculated. Based on the overall losses of the first meta-training set and the second meta-training set, the final objective function is formed. In this step, the different fault data in the second meta-training set force the model to learn general features (such as vibration energy distribution), improving the unknown fault recognition accuracy and the cross-domain generalization ability. The losses of the two meta-training sets are weighted and averaged to form the final objective function, which is used to guide the model training. The model can be directly deployed under unknown working conditions without additional data, reducing the operation and maintenance costs. This step solves the generalization ability bottleneck of traditional methods under unknown working conditions and unknown fault types through a collaborative training strategy (joint optimization of the first meta-training set and the second meta-training set) and the objective function.
[0075] As a preferred implementation manner of the present invention, a first fault type group and a second fault type group are obtained according to the fault type, specifically:
[0076] According to the characteristics of the fault type, through relevance analysis, a first fault type group and a second fault type group are obtained; and / or, according to the vibration intensity corresponding to the fault type, a first fault type group and a second fault type group are obtained.
[0077] The core objective of this embodiment is to optimize the division of fault type groups through a scientific grouping strategy, enhancing the generalization ability and classification accuracy of the model for rotating components under multiple working conditions and multiple fault types. Specifically, through correlation analysis or vibration intensity grouping, it is ensured that the characteristic differences between fault types in different groups are significant, avoiding misjudgment caused by feature similarity in the model to reduce feature confusion. Through the grouping strategy, the model can be trained only with known working condition data, reducing the dependence on target working condition data and achieving efficient data utilization.
[0078] Example 1: Correlation analysis grouping method. Extract time-domain (such as root mean square value, kurtosis), frequency-domain (such as main frequency energy, frequency band power ratio), or time-frequency domain (such as transient characteristics of Wigner-Ville distribution) features from vibration data. Calculate the Pearson correlation coefficient or cosine similarity of the feature vectors of each pair of fault types. For example, the vibration spectrum correlation coefficient between inner ring damage and outer ring damage of a bearing may be relatively high (such as 0.85), while the correlation coefficient between grease contamination and rolling element failure is relatively low (such as 0.32).
[0079] Set a correlation coefficient threshold (such as 0.6), and classify fault types with correlation coefficients lower than the threshold into different groups. For example, classify high-correlation faults (such as inner ring damage, outer ring damage) into the first group, and low-correlation faults (such as grease contamination, shaft eccentricity) into different groups.
[0080] In this embodiment, the characteristic differences between fault types in different groups are significant, reducing feature confusion, enhancing the clarity of the classification boundary, and reducing the misjudgment rate.
[0081] Example 2: Vibration intensity grouping method. Calculate the vibration energy (such as RMS value), peak factor, or kurtosis of each type of fault. For example, the kurtosis value of bearing spalling fault is significantly higher than that of grease contamination (such as the kurtosis of spalling fault is 15 and that of grease contamination is 3.2). Set a vibration intensity threshold (such as the RMS threshold is 0.5g), classify high-amplitude faults (such as ball fracture) into the first group, and low-amplitude faults (such as slight wear) into the second group.
[0082] In this embodiment, by enhancing the energy sensitivity, the classification accuracy of the model for faults corresponding to different amplitudes is improved.
[0083] It should be noted that correlation analysis and vibration intensity grouping can be adopted simultaneously to generate multi-dimensional grouping results. For example: The first group is high-correlation and high-amplitude faults (such as inner ring damage, outer ring damage of the bearing); the second group is low-correlation and low-amplitude faults (such as grease contamination, shaft misalignment). It can be understood that in the present invention, the rationality of the grouping results can be verified through clustering analysis (such as K-means) to ensure that the characteristic distributions of fault types within the group are concentrated.
[0084] In a specific embodiment, there are five types of faults in a centrifugal pump bearing: inner ring damage, outer ring damage, ball fracture, grease contamination, and shaft eccentricity. By analyzing the relevance and calculating the spectral correlation coefficients of each fault, it is found that the correlation coefficient between inner ring damage and outer ring damage is 0.85, and that with grease contamination is 0.32. The first group (high correlation) is divided, including inner ring damage and outer ring damage; the second group includes the other three types.
[0085] Grouping according to vibration intensity, calculating the RMS value, ball fracture (0.8g), grease contamination (0.2g). The first group (high amplitude) is divided, including ball fracture; the second group includes the other four types.
[0086] Carrying out combined grouping of relevance and vibration intensity, finally the first group includes inner ring damage, outer ring damage, and ball fracture (high correlation and high amplitude); the second group includes grease contamination and shaft eccentricity (low correlation and low amplitude).
[0087] This preferred embodiment scientifically divides the fault type groups through the combined strategy of relevance analysis and vibration intensity grouping, significantly improving the classification accuracy and robustness of the model under complex working conditions.
[0088] As a preferred embodiment of the present invention, the first meta-training set and the second meta-training set are specifically as follows:
[0089] Both the first meta-training set and the second meta-training set contain vibration data under multiple working conditions;
[0090] The vibration data of the first meta-training set under at least one working condition belongs to one of the first fault type group and the second fault type group, and the vibration data under at least another working condition belongs to the other of the first fault type group and the second fault type group;
[0091] The fault type group to which the vibration data of the second meta-training set belongs is different from that of the vibration data of the first meta-training set under the corresponding working condition.
[0092] The core objective of this preferred embodiment is to improve the cross-domain generalization ability of the model under unknown working conditions and unknown fault types through the carefully designed first meta-training set and second meta-training set. Specifically, the orthogonal design of working conditions and fault types ensures that the model simultaneously learns the working condition invariance features and fault type discrimination features during the training process. By alternately exposing the working condition data of different fault type groups, the dependence of the model on specific working conditions or fault types is reduced. The reverse fault type group design of the second meta-training set forces the model to distinguish known and unknown faults, improving the clarity of the classification boundary and enhancing the robustness to unknown faults.
[0093] It can be understood that the first meta-training set and the second meta-training set cover multiple working conditions, each containing multiple working conditions (such as rotational speeds of 735 RPM, 1102.5 RPM, and 1470 RPM). Through the alternating allocation of fault type groups, the vibration data of at least one working condition (such as working condition A) in the first meta-training set belongs to the first fault type group (such as bearing damage faults). The vibration data of at least another working condition (such as working condition B) belongs to the second fault type group (such as lubrication fault types). The fault type groups of the corresponding working conditions in the second meta-training set are opposite to those in the first meta-training set. The data of working condition A in the second meta-training set belongs to the second fault type group (such as lubrication faults), and the data of working condition B in the second meta-training set belongs to the first fault type group (such as bearing damage).
[0094] This step ensures that there is data for each working condition in the two meta-training sets, and the distribution of fault type groups is complementary. The model learns general features from the bearing damage data (first meta) of working condition A and the lubrication fault data (first meta) of working condition B, enabling it to accurately classify under unknown working condition C and enhancing the cross-working condition generalization ability. The reverse fault type group of the second meta-training set forces the model to distinguish unknown faults to identify unknown faults.
[0095] In a specific embodiment, taking the sample data sets of two working conditions and as an example, and are correspondingly divided into 、 and 、 according to the fault types. Among them, and represent different working conditions, and 1 and 2 represent the first fault type group and the second fault type group respectively.
[0096] That is to say, and have different fault types, and have different fault types, and have the same fault type and belong to the first fault type group, and have the same fault type and belong to the second fault type group. Then is used as the first meta-training set, is used as the second meta-test set.
[0097] Similarly, the present invention can also obtain a test set using the sample data sets under multiple working conditions. Taking the sample data sets under three working conditions as an example, for the sample data set , similarly is correspondingly divided into , . Then use as the first meta-training set, as the second meta-test set; or use as the first meta-training set, as the second meta-test set. The present invention does not limit this, and only needs to satisfy the settings of the first meta-training set and the second meta-training set in this embodiment.
[0098] In this embodiment, through the orthogonal design of working conditions and fault types, a first meta-training set and a second meta-test set that support cross-domain generalization are constructed, solving the bottleneck of the generalization ability of traditional methods under unknown working conditions and unknown fault types.
[0099] As a preferred embodiment of the present invention, the overall loss is specifically: combining the multi-classifier loss and the auxiliary binary classification loss to obtain the overall loss;
[0100]
[0101] wherein, represents the overall loss, represents the multi-classifier loss, represents the auxiliary binary classification loss.
[0102] The core objective of this embodiment is to improve the classification accuracy and robustness of the rotating component fault diagnosis model by jointly optimizing the multi-classification task and the auxiliary binary classification task. Specifically, by maximizing the prediction confidence of the model for known fault types through the multi-classification cross-entropy loss, the accurate classification of known faults is improved. By the auxiliary binary classification loss, the model is forced to distinguish between known and unknown fault types, reducing misjudgment to enhance the robustness of unknown faults.
[0103] The overall loss is . By directly adding, the two types of tasks are jointly optimized to ensure the balance of the global loss. Through the joint optimization of the multi-classification cross-entropy loss and the auxiliary binary classification loss in this embodiment, the problem of insufficient generalization ability of traditional fault diagnosis models under unknown working conditions and unknown fault types is solved.
[0104] As a preferred embodiment of this embodiment, according to the overall loss, the model parameters are updated, specifically: according to the gradient of the overall loss obtained from the first meta-training set under the preset model parameters, combined with the set learning rate, the model parameters are updated;
[0105]
[0106] wherein, and Represents the vibration data of the first meta-training set, Indicates that the vibration data corresponding to the first working condition in the first meta-training set belongs to the first fault type group, Indicates that the vibration data corresponding to the second working condition in the first meta-training set belongs to the second fault type group; And Respectively represent And The gradients of the overall loss; Represents the set learning rate; Represents the preset model parameters; Represents the updated model parameters.
[0107] The core objective of this embodiment is to improve the robustness and cross-domain generalization ability of model parameters by jointly optimizing the gradient information of multiple working conditions and multiple fault type groups. Specifically, the collaborative optimization of working conditions and fault types, by integrating the gradient information of different fault type groups under different working conditions, ensures that the model parameter update direction adapts to both working condition changes and fault type differences simultaneously. By integrating the gradient information of different fault type groups under different working conditions, it ensures that the model parameter update direction adapts to both working condition changes and fault type differences simultaneously, achieving the collaborative optimization of working conditions and fault types. Through the way of gradient summation, it avoids the conflict of gradient directions of different working conditions or fault type groups, reduces adversarial updates, and realizes the gradient direction consistency constraint.
[0108] The parameter update formula is , where, for the learning rate , the initial value is set to 0.001 and dynamically adjusted according to the performance of the second meta-training set (such as using the cosine annealing strategy).
[0109] This embodiment solves the problem of insufficient generalization ability of traditional models under complex working conditions and unknown fault types through the joint optimization of gradients of multiple working conditions and multiple fault type groups. In addition, it improves the model convergence speed through the gradient direction consistency constraint.
[0110] Specifically, an objective function for the fault diagnosis of rotating components is obtained according to the overall loss obtained from the first meta-training set and the second meta-training set, specifically:
[0111]
[0112] Where, And Represents the vibration data of the second meta-training set, Indicates that the vibration data corresponding to the first working condition in the second meta-training set belongs to the second fault type group, Indicates that the vibration data corresponding to the second working condition in the second meta-training set belongs to the first fault type group; And respectively represent and the overall loss under the preset model parameters; and respectively represent and the overall loss under the updated model parameters.
[0113] The core objective of this embodiment is to improve the cross - domain generalization ability of the rotating component fault diagnosis model under unknown working conditions and unknown fault types by jointly optimizing the losses of the first meta - training set and the second meta - training set. Specifically, by jointly optimizing the losses of the two meta - training sets, it is ensured that the model parameters can simultaneously adapt to the characteristic differences of different fault type groups under different working conditions, realizing the collaborative optimization of multiple working conditions and multiple fault types. Through the design of the reverse fault type group of the second meta - training set, the model's dependence on specific working conditions or fault types is suppressed, the overfitting risk is reduced, and the adversarial loss is inhibited.
[0114] For the first meta - training set, the obtained overall loss is . For the second meta - training set, the obtained overall loss is . The model learns the working condition invariance characteristics during the alternating exposure of the first meta - training set (the first fault type group → the second fault type group) and the second meta - training set (the second fault type group → the first fault type group). The reverse fault type group of the second meta - training set (such as lubrication fault at 735 RPM working condition) forces the model to distinguish unknown faults and suppresses overfitting.
[0115] Construct the objective function . The goal is to find the model parameters that minimize the sum of the losses of the first meta - training set and the second meta - training set. The iterative optimization process of the model parameters is as follows: using the loss of the first meta - training set to update the parameters to , using the updated parameters to calculate the loss of the second meta - training set , and jointly optimize.
[0116] It can be understood that for the above objective function formula, the first - order Taylor formula is used for the following expansion to obtain
[0117]
[0118] where the second term is the sum of each gradient dot product. Maximizing the gradient dot product can regularize this training process, making the update directions of different tasks match, and realizing the synchronous matching of working conditions and fault types. This step solves the problem of insufficient generalization ability of traditional models under unknown working conditions and unknown fault types by jointly optimizing the losses of the first meta - training set and the second meta - training set.
[0119] As a preferred embodiment of the present invention, the method for diagnosing faults of rotating components based on co-training further includes:
[0120] According to the collected vibration data, through preprocessing, a model training data set is obtained; wherein, the preprocessing at least includes data anomaly detection, sliding window segmentation, and time-frequency feature conversion processing, and the data anomaly detection is used to identify and eliminate abnormal data.
[0121] The core objective of this embodiment is to improve the quality and applicability of vibration data through data preprocessing, and provide reliable, structured and feature-rich input data for subsequent model training. Specifically, noise or abnormal data is eliminated through anomaly detection to reduce noise interference in model training and ensure data quality. The continuous vibration signal is converted into time series segments of a fixed length through sliding window segmentation to adapt to the model input format and realize the structuring of time series data. The joint time-domain and frequency-domain features of the vibration signal are extracted through time-frequency feature conversion to enhance the ability to distinguish fault types and enhance features and separability.
[0122] For data anomaly detection, the method includes at least any one of statistical methods and machine learning methods. The statistical method is threshold detection based on the mean and standard deviation (such as ). The machine learning method is to use Isolation Forest or LOF (Local Outlier Factor) to detect outliers.
[0123] For abnormal data, directly delete the abnormal data points or repair them through interpolation methods (such as linear interpolation). In addition, the anomaly detection threshold is dynamically adjusted according to working condition parameters (such as rotational speed, load) (for example, the noise threshold is increased by 20% under high rotational speed conditions). This embodiment reduces noise interference, reduces the proportion of abnormal data, improves the stability of model training, and reduces the false positive rate.
[0124] As an example of this embodiment, the sliding window segmentation is specifically:
[0125] According to the vibration data after anomaly elimination, through a non-overlapping sliding window segmentation method, a plurality of vibration data segments are obtained; wherein the window step size is determined according to the vibration data acquisition length, acquisition frequency, and data vibration period.
[0126] The core objective of this embodiment is to convert continuous vibration data into independent segments of a fixed length through a non-overlapping sliding window segmentation method, so as to meet the requirements of the model input format and improve the efficiency of feature extraction. Specifically, the continuous vibration signal is segmented into time-series segments of a fixed length to adapt to the input format of a data structuring adaptation model (such as CNN or RNN). The non-overlapping window reduces data redundancy and computational complexity, is applicable to real-time or resource-constrained scenarios, and optimizes the computational efficiency. By matching the window step with the vibration period, it is ensured that each segment contains a complete fault feature period, improving the classification accuracy.
[0127] To determine the window parameters, the spectrum of the vibration signal is analyzed through Fourier transform or wavelet transform to determine the main fault feature frequencies (such as the impact pulse frequency of bearing faults). The vibration period . Among them, is the peak frequency.
[0128] According to the sampling frequency and the vibration period T , the window length is set. For example, if , , then (250 sampling points). The non-overlapping window step is used to ensure that the segments are completely independent.
[0129] In this embodiment, each window contains at least one complete vibration period, improving the capture rate of the vibration pulses of faults. Compared with overlapping windows, the amount of data is reduced, and the inference speed and computational efficiency are improved. This embodiment solves the problems of computational redundancy and insufficient working condition adaptability brought by traditional overlapping windows through the non-overlapping sliding window segmentation method.
[0130] As another embodiment under this implementation manner, the time-frequency feature conversion processing is specifically as follows: according to the vibration data after the sliding window segmentation processing, the negative frequency part in the signal is removed through an instantaneous feature extraction method; according to the processed vibration data, the instantaneous power spectral density is obtained through a time-frequency analysis method; among them, the frequency resolution of the time-frequency analysis method is determined according to the highest frequency of the collected vibration data, and the time resolution is determined according to the minimum bandwidth of the collected vibration data.
[0131] The core objective of this embodiment is to generate a high-resolution and low-redundancy instantaneous power spectral density through time-frequency analysis and feature extraction, so as to enhance the time-frequency joint features of vibration signals and improve the model's ability to distinguish fault types. Specifically, through the set frequency and time resolutions, high-resolution time-frequency analysis is performed to capture the transient features of vibration signals (such as the impact pulses of bearing damage). The negative frequency components are removed, the negative frequencies are suppressed, the signal redundancy is reduced, and the computational complexity is lowered. The instantaneous power spectral density of the positive frequency part is retained, which intuitively reflects the time-frequency distribution of fault features and enhances the interpretability of features.
[0132] In this step, a real signal is input , with a length of . Through an instantaneous feature extraction method (such as the Hilbert transform ), an analytic signal is generated. The spectrum of the analytic signal only contains positive frequency components.
[0133] It should be noted that before the Hilbert transform, is low-pass filtered (cutoff frequency ) to ensure that the signal bandwidth satisfies the Nyquist sampling theorem. In this step, the negative frequency components of the analytic signal are completely removed, avoiding cross-term aliasing in the WVD calculation, improving the signal fidelity, and reducing the sampling error of high-frequency components.
[0134] According to the processed vibration data, through a time-frequency analysis method (such as the WVD method),
[0135]
[0136] where represents the complex conjugate of the analytic signal. The value range of m is (only positive frequencies are retained). By restricting the range of , it is ensured that the calculation is only for positive frequency components, avoiding spectral folding when it exceeds N / 2. is the standard form of the discrete WVD and is compatible with the positive frequency characteristics of the analytic signal.
[0137] In this step, through aliasing elimination, the time-frequency energy concentration during faults is improved, and the cross-term interference is reduced. In addition, the positive frequency characteristics of the analytic signal reduce the computational amount and optimize the computational efficiency. Generally speaking, this embodiment solves the problems of noise interference, missing time-series patterns, and insufficient feature separability existing in the original vibration data through the joint preprocessing of data anomaly detection, sliding window segmentation, and time-frequency feature conversion.
[0138] As a preferred embodiment of the present invention, the rotating component fault diagnosis model is specifically as follows:
[0139] Based on the vibration data of the rotating component collected in real time, the confidence score is obtained through the auxiliary binary classifier in the rotating component fault diagnosis model, and the known fault types and unknown fault types are judged through the confidence score.
[0140] For the known fault types, the probability distribution of the known fault types is obtained through the multi-classifier in the rotating component fault diagnosis model to diagnose the faults of the rotating component.
[0141] The core objective of this embodiment is to achieve accurate diagnosis of rotating component faults through the collaborative work of the auxiliary binary classifier and the multi-classifier, and at the same time improve the model's detection ability for unknown fault types. Specifically, for known / unknown fault classification, the auxiliary binary classifier is used to distinguish known faults from unknown faults to reduce the risk of misdiagnosis. For the fine-grained diagnosis of known faults, the multi-classifier is used to predict the probability distribution of known fault types to improve the classification accuracy.
[0142] In this embodiment, the trained rotating component fault diagnosis model is applied to judge the fault type of the rotating component according to the vibration data collected in real time. The vibration data of the rotating component is obtained through sensors (such as bearing or gearbox vibration signals). It can be understood that the vibration data collected in real time also needs to be preprocessed to generate feature vectors. The preprocessed feature vectors are input into the auxiliary binary classifier to output the binary classification results (known faults / unknown faults) and the confidence score. If the confidence score , it is determined as a known fault; otherwise, it is an unknown fault. Samples with abnormal fluctuations in the confidence score (such as oscillating around 0.8) are marked as potential noise or marginal cases.
[0143] The samples determined as "known faults" by the auxiliary classifier are input into the multi-classifier for processing to output the probability distribution of the known fault types (such as the probability of inner ring damage of the bearing is 0.85, and the probability of outer ring damage is 0.12). The category with the highest probability is selected as the final diagnosis result. If the highest probability < 0.7, trigger secondary verification (such as combining historical data or expert systems).
[0144] Such as Figure 2 shown, in a specific embodiment, an induction motor drives a centrifugal pump, fresh water is taken from a water storage tank and transported through a pipeline system, and valves are installed before and after the pump to control the flow. The common categories of motor bearing faults at different operating speeds are analyzed. High-frequency vibration data is used as the original input, and this data is collected by a Wilconxon 786B-10 100mV / g uniaxial accelerometer with a sampling frequency of 20KHz.
[0145] The corresponding number of samples at different operating speeds (multiple working conditions) is shown in Table 1.
[0146] Table 1 Corresponding number of samples under multiple working conditions
[0147]
[0148] In addition, the bearing fault categories are shown in Table 2.
[0149] Table 2 Bearing fault categories
[0150]
[0151] To verify the performance of the model in diagnosing known and unknown faults under unknown working conditions, the generalization tasks shown in Table 3 are obtained. In each task, one more bearing fault type is added to the target working condition compared with the known working condition, and only the data in the known working condition is involved in the model training. The data in the target working condition is only used for model performance evaluation.
[0152] Table 3 Generalization tasks for model verification
[0153]
[0154] In all the above generalization tasks, the data of the target working condition does not participate in the model training and is only used for testing. The evaluation metrics are the classification accuracy of known faults (Acc_k) and the classification accuracy of unknown faults (Acc_u) in the unknown target domain.
[0155] The performance metrics of the model on 12 groups of generalization tasks are shown in Table 4. The average metrics of the proposed method in 12 task groups reach a known accuracy of 89.413% and an unknown accuracy of 92.481%. For different target working conditions, the average accuracies of known fault classification at 735 RPM, 1102.5 RPM, and 1470 RPM working conditions are 82.208%, 91.118%, and 91.915% respectively, and the average accuracies of unknown fault identification are 87.953%, 97.333%, and 92.158% respectively.
[0156] Table 4 Model performance metrics
[0157]
[0158] Furthermore, deleting the relevant components is denoted as M2, which is used to evaluate the auxiliary multi-binary classifier and Impact on model performance. In addition, MLDG (Meta-Learning for Domain Generalization) is denoted as M3, which is used to evaluate the impact of the binary learning strategy on model performance. The accuracies corresponding to M2 and M3 are shown in the above table. Compared with M2, the average Acc_k and Acc_u of the model in this application on 12 sets of generalization tasks are increased by 12.476% and 16.98% respectively, indicating that the auxiliary multi-binary classifiers and provide a more accurate classification boundary for fault identification under unknown working conditions.
[0159] In addition, compared with M3, based on achieving an unknown accuracy of 92.481%, the proposed method improves the known fault identification accuracy by 9.645%. This application of the present invention realizes a more fine-grained category-level model optimization.
[0160] What is not described in the present invention can be realized by adopting or referring to the existing technologies.
[0161] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0162] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A fault diagnosis method for rotating components based on co-training, characterized in that, Including: Based on the fault type, obtain the first fault type group and the second fault type group, and combine the collected model training data set to obtain the first meta-training set and the second meta-training set, where the model training data set includes vibration data under multiple working conditions and multiple fault types of the rotating component, both the first meta-training set and the second meta-training set include vibration data under multiple working conditions, the vibration data of the first meta-training set under at least one working condition belongs to one of the first fault type group and the second fault type group, and the vibration data under at least another working condition belongs to the other of the first fault type group and the second fault type group, the fault type group to which the vibration data of the second meta-training set belongs is different from that of the vibration data corresponding to the first meta-training set under the same working condition; Based on the preset model parameters, according to the first meta-training set, obtain the auxiliary binary classification loss through the auxiliary binary classifier, obtain the multi-classifier loss through the cross-entropy function, and combine to obtain the overall loss; According to the overall loss, update the model parameters, combine the second meta-training set, obtain the overall loss, and obtain the objective function according to the overall losses obtained from the first meta-training set and the second meta-training set for training the rotating component fault diagnosis model.
2. The method for diagnosing faults of rotating components based on co-training according to claim 1, wherein Obtain the first fault type group and the second fault type group according to the fault type, specifically: Based on the characteristics of the fault type, obtain the first fault type group and the second fault type group through correlation analysis; and / or Obtain the first fault type group and the second fault type group according to the vibration intensity corresponding to the fault type.
3. The method for diagnosing faults of rotating components based on co-training according to claim 1, wherein The overall loss is specifically: Combine the multi-classifier loss and the auxiliary binary classification loss to obtain the overall loss; wherein, represents the overall loss, represents the multi-classifier loss, represents the auxiliary binary classification loss.
4. The method for diagnosing faults of rotating components based on co-training according to claim 3, characterized in that, Update the model parameters according to the overall loss, specifically: According to the gradient of the overall loss obtained from the first meta-training set under the preset model parameters, combine the set learning rate to update the model parameters; Among them, and represent the vibration data of the first meta-training set, indicates that the vibration data corresponding to the first working condition in the first meta-training set belongs to the first fault type group, indicates that the vibration data corresponding to the second working condition in the first meta-training set belongs to the second fault type group; and respectively represent and[[ID=X]] the gradients of the overall loss; represents the set learning rate; represents the preset model parameters; represents the updated model parameters. Note: There seems to be a missing tag in the original text for the part marked as in the English translation. It should be corrected in the original text for a more accurate translation.
5. The method for diagnosing faults of rotating components based on co-training according to claim 4, wherein Obtain the objective function for rotating component fault diagnosis according to the overall losses obtained from the first meta-training set and the second meta-training set, specifically: Among them, and represent the vibration data of the second meta-training set, indicating that the vibration data corresponding to the first working condition in the second meta-training set belongs to the second fault type group, indicating that the vibration data corresponding to the second working condition in the second meta-training set belongs to the first fault type group; and respectively represent and the overall loss under the preset model parameters; and respectively represent and the overall loss under the updated model parameters.
6. The method for diagnosing faults of rotating components based on co-training according to claim 1, characterized in that It also includes: According to the collected vibration data, through preprocessing, obtain the model training data set; Among them, the preprocessing at least includes data anomaly detection, sliding window segmentation, and time-frequency feature conversion processing, and the data anomaly detection is used to identify and eliminate abnormal data.
7. The method for diagnosing faults of rotating components based on co-training according to claim 6, characterized in that, The sliding window segmentation is specifically: According to the vibration data after anomaly elimination, through the non-overlapping sliding window segmentation method, obtain multiple vibration data segments; Among them, the window step size is determined according to the vibration data acquisition length, acquisition frequency, and data vibration period.
8. The method for diagnosing faults of rotating components based on co-training according to claim 6, characterized in that, The time-frequency feature conversion processing is specifically: According to the vibration data after sliding window segmentation processing, remove the negative frequency part in the signal through the instantaneous feature extraction method; According to the processed vibration data, through the time-frequency analysis method, obtain the instantaneous power spectral density; Among them, the frequency resolution of the time-frequency analysis method is determined according to the highest frequency of the collected vibration data, and the time resolution is determined according to the minimum bandwidth of the collected vibration data.
9. The method for diagnosing faults of rotating components based on co-training according to claim 1, wherein The rotating component fault diagnosis model, specifically: Based on the vibration data of the rotating component collected in real time, a confidence score is obtained through the auxiliary binary classifier in the rotating component fault diagnosis model, and the known fault types and unknown fault types are judged through the confidence score. For the known fault types, the probability distribution of the known fault types is obtained through the multi-classifier in the rotating component fault diagnosis model to diagnose the faults of the rotating component.
Citation Information
Patent Citations
Federal generalization-based tail drive system health state evaluation method
CN117786489A
Adaptive cross-working-condition fault diagnosis method for rotary machinery based on depth discrimination and unsupervised field
CN118643324A
An open set domain adaptive fault diagnosis method and system based on complementary decision making
CN118643346B
Rolling bearing fault diagnosis method based on convolutional neural network (CNN) model and transfer learning
CN110220709A
Industrial equipment fault diagnosis method and device based on cross-domain generalization label
CN116956048A