Mechanical fault intelligent diagnosis method based on progressive transfer learning network
Through a progressive transfer learning network combined with data from laboratory and industrial environments, the problem of fault diagnosis in small samples without labels is solved, and the precise identification of mechanical faults in complex industrial environments is achieved.
Patent Information
- Application Number
- CN202510537774.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
AI Technical Summary
In actual industrial production environment, the vibration signal characteristics distribution of mechanical equipment are complex and changeable, and the cost of labeling data is high, which makes it difficult to diagnose faults in small samples without labels, and the problem of category imbalance affects the accuracy of diagnosis.
A progressive transfer learning network is adopted to perform progressive adversarial learning through labeled data in laboratory environment and labelless data in industrial environments. Pseudo-labels are obtained using multi-core maximum mean difference and cluster distance evaluation, feature distribution is gradually mapped, and fault categories are identified through state classifiers.
It improves the model's fault adaptability and generalization ability in complex industrial environments, enhances the learning efficiency of unlabeled data, and realizes accurate identification of mechanical equipment faults.
Smart Images

Figure CN120449006A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mechanical component monitoring and fault diagnosis, and relates to an intelligent mechanical fault diagnosis method based on a progressive transfer learning network. Background Art
[0002] Gears and bearings, as core components in mechanical equipment transmission systems, not only carry the critical task of transmitting power and motion, but are also the most vulnerable and failure-prone links in the entire equipment system. Their operating status directly affects the performance, efficiency, and safety of mechanical equipment. Therefore, effective monitoring and fault diagnosis of gears and bearings are crucial for ensuring the continuity and safety of industrial production. Against this backdrop, intelligent fault diagnosis technology has emerged as a key technology for improving the operation and maintenance management of industrial machinery.
[0003] In recent years, with the rapid advancement of artificial intelligence (AI), particularly the widespread application of deep learning algorithms, data-driven intelligent fault diagnosis methods have begun to emerge in the industrial sector, demonstrating enormous potential and application value. By analyzing the vast amounts of data generated by mechanical equipment during operation, such as vibration, sound, and temperature signals, these methods can automatically learn and identify the characteristics of different fault types, enabling early warning and accurate diagnosis. However, to fully realize their effectiveness, the training process for deep learning models places stringent requirements on the dataset.
[0004] First, the training dataset must be clearly labeled. Each sample must correspond to a specific fault type or normal state. This is fundamental to the model's ability to distinguish between different states. Furthermore, sufficient fault samples are crucial to ensuring the model's generalization capabilities. In practical applications, this means collecting sufficient samples covering a wide range of possible fault modes to ensure the model can accurately judge even unknown faults.
[0005] Secondly, the number of samples from different categories in the training set should be balanced to avoid model bias caused by too many or too few samples from a certain category. Class imbalance may lead to a decrease in the model's ability to recognize samples from the minority class, thereby affecting the accuracy of the overall fault diagnosis.
[0006] Finally, to ensure the effectiveness and reliability of the model, the training and test sets should have the same feature distribution. This means that the two should be consistent in terms of data collection conditions, environmental noise, device status, etc., to ensure that the features learned by the model during the training phase can be effectively verified on new samples.
[0007] However, in real-world industrial production environments, these ideal conditions are often difficult to fully meet. Frequent changes in the speed, load, and torque of mechanical equipment, coupled with the influence of uncontrollable factors such as mechanical wear and external noise, result in complex and variable distributions of the characteristics of the collected raw vibration signals. Furthermore, due to the high cost of labeling data, including the investment of manpower, time, and material resources, most actual data is unlabeled, further increasing the difficulty of constructing a high-quality training set. Summary of the Invention
[0008] In light of this, the present invention aims to provide an intelligent mechanical fault diagnosis method based on a progressive transfer learning network. This method combines labeled laboratory fault data with a small amount of unlabeled industrial environment fault data through progressive adversarial learning. By progressively constraining the fault characteristics of different scales in the source and target domains, the source domain data feature distribution is gradually mapped to the target data distribution domain. Pseudo-labels of actual fault samples are then obtained through multi-kernel maximum mean difference (MK-MMD) and cluster distance evaluation. Finally, a state classifier is used to identify the fault category of the mechanical equipment.
[0009] In order to achieve the above object, the present invention provides the following technical solutions:
[0010] A method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network, the method comprising the following steps:
[0011] S1. Collect raw signal data in a laboratory environment and a real industrial environment, and preprocess the data;
[0012] S2. Establish a fault identification model based on a progressive transfer learning network architecture, wherein corresponding features are extracted by at least several source domain feature extractors and target domain feature extractors, and the source domain data features are gradually mapped to the target data feature distribution domain for fault classification;
[0013] S3. Combining the loss function in the propagation process, the optimal network parameters are obtained with the goal of minimizing the classification error of source and target domain samples and minimizing the inter-domain difference of sample distribution;
[0014] S4. Collect real-time signal data in a real industrial environment and input it into the optimized fault identification model based on the progressive transfer learning network architecture. The fault identification model outputs the fault type.
[0015] Furthermore, in step S1, the data collected in the laboratory is defined as source domain data, and the data collected in the real industrial environment is defined as target domain data, wherein the source domain data has fault labels, and the target domain data does not have fault labels;
[0016] The data preprocessing process includes: denoising and filtering the original signal, removing outliers, and completing the original signal, and then intercepting the signal through a sliding window to produce the source domain dataset and the target domain dataset respectively.
[0017] Further, in step S2, the fault identification model based on the progressive transfer learning network includes several feature extractors F, several domain discriminators D and at least one state classifier C, wherein the feature extractor is divided into a source domain feature extractor F Sm and target domain feature extractor F Tm , m=1,2,3, including the following process:
[0018] S21, through the source domain feature extractor F Sm and target domain feature extractor F Tm Extract features of different scales from the source domain dataset and the target domain dataset respectively;
[0019] S22, domain identifier D m For the source domain feature extractor F Sm and target domain feature extractor F Tm Each layer of extracted features is used to identify the source and constrain the distribution differences of features in different dimensions of the source or target domain;
[0020] S23, through the flattening operation, the source domain feature extractor F Sm and target domain feature extractor F Tm Finally, the extracted multi-dimensional features are aligned;
[0021] S24, aligned target domain features and source domain features V S and target domain features V T Input to the state classifier C for fault type identification.
[0022] Further, in step S21, the source domain feature extractor F Sm and target domain feature extractor F Tm They all include three-level feature mapping neural networks. Each level of feature mapping neural network extracts features of different scales. Each level of feature mapping neural network includes convolution layer, activation layer, normalization layer and pooling layer. For the input source domain dataset and target domain dataset The extraction process is expressed as:
[0023]
[0024] In the formula Represents the source domain feature extractor F at the mth level Sm and target domain feature extractor F Tm Extracted features, Represents the m-th level source domain feature extractor F Sm and target domain feature extractor F Tm network parameters.
[0025] Further, in step S22, the source domain feature extractor F Sm and target domain feature extractor F Tm Each layer outputs f Sm and f Tm Input into the domain identifier D respectively m To identify the source of the sample, its output data type is Boolean. If the sample comes from the source domain, the domain identifier D m The output is 0; if the sample comes from the target domain, the domain discriminator D m The output is 1; the scope and contribution of each domain classifier constraint are different, and different weights μ are assigned to each domain classifier m .
[0026] Further, in step S23, the source domain feature extractor F Sm and target domain feature extractor F Tm The final extracted multidimensional feature f S3 and f T3 Use the flatten operation Flatten(·) to align the source domain features V S and target domain features V T :
[0027]
[0028] The flattening operation unifies the dimensions of the extracted multidimensional features.
[0029] Furthermore, in step S24, the state classifier C is composed of a fully connected layer, an activation layer, and a normalization layer. In the state classifier C, the distance between the feature distributions is measured by the maximum average deviation, and the target domain samples are classified into the state classifier C. Tags Make a prediction:
[0030]
[0031] In the formula is the source domain sample The fault label, represents the set clustering range, B(·) represents Boolean operation, If the sample If the feature is mapped to the set cluster range, its value is 1; if it is not mapped to the cluster range, its value is 0. K represents the number of all fault types.
[0032] Furthermore, in step S3, the network parameters of the initial source domain feature extractor are adjusted to minimize the classification error of source domain and target domain samples and minimize the inter-domain difference of sample distribution. Network parameters of the target domain feature extractor Network parameters of the domain discriminator The network parameters θ of the state classifier C The optimization is performed by minimizing the loss function L of the maximum average deviation MMD , minimize the total loss function L of the discriminator D_total Get the optimal network parameters:
[0033]
[0034] In the above formula, we have:
[0035]
[0036] In the formula, Non means that this item does not exist. The optimization methods of each network parameter are as follows:
[0037]
[0038] In the formula, n=0,1,2, which is the parameter The training process is to first train the discriminator D1 and then train the feature extractor at the same time. and Retrain the discriminator D2 and then train the feature extractor and And so on, guiding the discriminator and feature extractor to complete the training;
[0039] After training, the optimal network parameters are obtained Update weight μ m The method is as follows:
[0040]
[0041] The fault category of the target sample is estimated by clustering with the maximum mean deviation, and the optimization parameter θ is trained according to the following formula C :
[0042]
[0043] In the above process, L all is the total loss function of the fault identification model, where L MMD The loss function representing the maximum mean deviation, L D_total is the overall loss function of the domain discriminator, L C is the overall loss function of the state classifier.
[0044] Furthermore, the total loss function L of the fault identification model is all Expressed as:
[0045] L all =L C +L D_totall
[0046] Among them, the total loss function L of the domain discriminator is D_total The definition is as follows:
[0047]
[0048] μ m Represents the weight of each scale domain discriminator, and the loss function of each scale domain discriminator The calculation formula is as follows:
[0049]
[0050] Where N S and N T Represent the number of source domain samples and the number of target domain samples, d i d j are the labels of the source domain and the target domain, respectively, and the domain label d of the source sample i = 0, the domain label d of the target sample j =1, then:
[0051]
[0052] The state classifier C uses cross entropy as the loss function L C , which is expressed as follows:
[0053]
[0054] Where ψ S , ψ T Represent the source domain samples {X S ,Y S} and target domain samples {X T ,Y T}, where Y T is the initial predicted label of the target sample, H(·) represents the cross entropy loss function;
[0055] The loss function L of the maximum mean deviation MMD The calculation method is as follows:
[0056]
[0057] κ(·) represents the Gaussian kernel.
[0058] In step S4, after obtaining the optimal network parameters After that, the target domain fault samples without labels are input into the optimized model to identify the fault type, and the final output is expressed as:
[0059]
[0060] where Y test is the predicted label of the mechanical state, X test represents the fault sample in the target domain, and max(·) extracts the label with the highest fault probability.
[0061] The beneficial effects of the present invention are:
[0062] This paper proposes an innovative fault diagnosis method that combines labeled laboratory fault data with a small amount of unlabeled industrial fault data for progressive adversarial learning. This method effectively addresses the difficulty of fault diagnosis in rotating machinery with small, unlabeled samples, providing strong support for the safe and efficient operation of equipment.
[0063] First, the present invention leverages the differences in data features between the source domain (in a laboratory setting) and the target domain (in an industrial setting) using a progressive adversarial learning strategy. This strategy imposes constraints on fault signatures at different scales, gradually mapping the source domain data feature distribution to the target domain data distribution. This process not only considers the common characteristics between the two domains, but also accounts for any subtle differences between them, thereby improving the model's adaptability and generalization capabilities to fault patterns in real-world environments.
[0064] Secondly, based on feature distribution mapping, this paper introduces the Multi-Kernel Maximum Mean Difference (MK-MMD) as an evaluation metric to further optimize the distribution matching effect within the feature space. Furthermore, cluster distance evaluation is used to obtain pseudo-labels for actual fault samples. This combination effectively improves the model's learning efficiency for unlabeled data and enhances its ability to recognize new types of faults.
[0065] Finally, the resulting feature vector is fed into a state classifier to accurately identify the fault type of the mechanical equipment. The state classifier's design takes into account the operating characteristics and common failure modes of rotating machinery, ensuring it can accurately distinguish between different fault types in complex industrial environments.
[0066] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0068] Figure 1 This is an architecture diagram of the fault identification model based on the progressive transfer learning network of the present invention;
[0069] Figure 2 This is a flow chart of an example of identifying a mechanical intelligent diagnosis method based on a progressive migration network according to the present invention. DETAILED DESCRIPTION
[0070] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0071] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0072] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0073] See also Figures 1 and 2 , which is an intelligent diagnosis method for mechanical faults based on progressive transfer learning network.
[0074] Example
[0075] This embodiment proposes a method for intelligent diagnosis of mechanical faults based on a progressive migration network, which includes the following steps:
[0076] S1. Collect raw signal data in a laboratory environment and a real industrial environment, and preprocess the data;
[0077] S2. Establish a fault identification model based on a progressive transfer learning network architecture, wherein corresponding features are extracted by at least several source domain feature extractors and target domain feature extractors, and the source domain data features are gradually mapped to the target data feature distribution domain for fault classification;
[0078] S3. Use the stochastic gradient descent back-propagation algorithm, combined with the loss function in the propagation process, to obtain the optimal network parameters with the goal of minimizing the classification error of source and target domain samples and minimizing the inter-domain difference in sample distribution;
[0079] S4. Collect real-time signal data in a real industrial environment and input it into the optimized fault identification model based on the progressive transfer learning network architecture. The fault identification model outputs the fault type.
[0080] In step S1, the data collected in the laboratory is defined as the source domain data, and the data in a real-world industrial environment is defined as the target domain data. The source domain data collected in the laboratory is labeled with faults, while the target domain data is unlabeled. The data preprocessing process involves first performing noise reduction filtering, outlier removal, and completion on the original signal. Then, the signal is intercepted using a sliding window to create the source and target domain datasets.
[0081] In step S2, the progressive transfer learning network architecture proposed in this embodiment is as follows Figure 1 As shown, the progressive transfer learning network consists of a feature extractor F, a domain discriminator D, and a state classifier C, where the feature extractor is divided into a source domain feature extractor F Sm (m=1,2,3) and target domain feature extractor F Tm (m=1,2,3), where the data processing process of the fault identification model based on the progressive transfer learning network architecture is:
[0082] S21. Input the source and target domain samples into the source and target domain feature extractors, respectively. The source and target domain feature extractors are used to extract features of the source and target domain samples, respectively. Traditional adversarial transfer networks only contain a single feature extractor to simultaneously map the features of both source and target domain samples, resulting in poor feature mapping performance. Progressive transfer learning networks use separate feature extractors for source and target domain feature mapping, effectively overcoming the drawbacks of traditional transfer adversarial networks with a single feature extractor.
[0083] Both the source domain and target domain feature extractors are composed of three-level feature mapping neural networks, which are used to extract features of samples at different scales. Each source domain feature extractor F Sm and target domain feature extractor F Tm They are composed of convolutional layers, activation layers, normalization layers (BN) and pooling layers, with θ F represents the network parameters of the feature extractor, and f represents its output. A major issue facing the unipolar feature extractor in traditional adversarial transfer networks is that it can only constrain the high-dimensional features of samples, but cannot constrain the sample mapping features in other dimensions. This can result in similar feature distributions across the source and target domains, but differences between samples. This can lead to negative transfer and reduce fault classification accuracy. In a progressive transfer learning network, a three-level feature extractor is employed to map the low-, medium-, and high-dimensional features of source and target domain samples at different scales, gradually reducing the distribution differences between source and target domain samples. This step-by-step feature constraint approach, unlike traditional single-level feature mapping methods, effectively supervises the sample feature mapping process, ensuring that the distributions of target and source domain sample features are as consistent as possible across different dimensions. By assigning different weights to domain classifiers at different levels, the contribution of sample features at different scales to the inter-domain constraints is adjusted. This step-by-step feature mapping constraint achieves a smooth mapping from labeled source domain features to unlabeled target domain features.
[0084] The source domain and target domain feature extractors are respectively Sm (·)(m=1,2,3),F Tm (·)(m=1,2,3) means that the network input source domain and target domain datasets are in Represents the target domain sample label set, which is unknown before model training and needs to be initially predicted through clustering. and The characteristic expression of is:
[0085]
[0086] In the formula Source domain feature extractor F Sm and target domain feature extractor F Tm The final extracted multidimensional features and Respectively expressed as:
[0087]
[0088] S22, source domain feature extractor F Sm and target domain feature extractor F TmEach layer outputs f Sm and f Tm Input into the domain identifier D respectively m (m=1,2,3) is used to identify the source of the sample, and its output data type is Boolean. The sample comes from the source domain, and the domain identifier D m The output is 0; the sample comes from the target domain, the domain discriminator D m The output is 1. Through different levels of domain discriminators D m (m=1,2,3) constrains the distribution difference of different dimensional features of the source domain or target domain, using θ Dm Denotes the domain discriminator D m The scope and contribution of each domain classifier constraint are different, and different weights μ are assigned to each domain classifier. m (m=1,2,3).
[0089] The weights of the domain classifier are optimized and learned during training, and the total loss function L of the discriminator is D_total The definition is as follows:
[0090]
[0091] Loss function of discriminators at each scale The calculation formula is as follows:
[0092]
[0093] Where N S and N T Represent the number of source domain samples and the number of target domain samples respectively, where d i d j are the labels of the source domain and the target domain, respectively, and the domain label d of the source sample i = 0, and the domain label of the target sample d j =1, we can conclude that:
[0094]
[0095] S23, source domain feature extractor F Sm and target domain feature extractor F Tm The final extracted multidimensional feature f S3 and f T3 Use the flatten operation Flatten(·) to get the feature V S and V T Then, the feature distribution distance is evaluated and constrained by minimizing the maximum average deviation between the source domain features and the target domain features, so that the distribution of the target sample and the source domain sample features is better aligned and the distribution difference between the domains is reduced. Sample mapping feature V S and V TThe calculation formula is defined as follows:
[0096]
[0097] S24, flattened one-dimensional source domain feature V S and one-dimensional target domain feature V T The input is sent to the state classifier C for fault type identification. The state classifier C consists of a fully connected layer, an activation layer, and a normalization layer, which is used to classify sample faults. In the state classifier C, the distance between feature distributions is measured by MK-MMD to predict the target domain samples. Tags Purpose, target domain sample Tags The prediction method is as follows:
[0098]
[0099] In the formula is a source domain sample The fault label, represents the set clustering range, B(·) represents Boolean operation, If the sample If a feature map falls within the specified cluster range, its value is 1; if it falls outside the specified cluster range, its value is 0. K represents the total number of fault types. Clustering is used to preliminarily predict the fault type of target domain samples and label them. This facilitates training the state classifier using target samples with predicted labels.
[0100] The loss of maximum mean deviation is calculated as follows:
[0101]
[0102] κ(·) represents the Gaussian kernel, and its calculation formula is as follows:
[0103]
[0104] Where, ||x1-x2|| 2 It represents the square of the Euclidean distance between vectors x1 and x2. σ is the main parameter of the Gaussian kernel, which is usually the standard deviation or bandwidth.
[0105] The state classifier C uses cross entropy as the loss function L C , which is expressed as follows:
[0106]
[0107] Where ψ S , ψ T Represent the source domain samples {X S ,YS} and target domain samples {X T ,Y T}, where Y T is the initial predicted label of the target sample, and H(·) represents the cross entropy loss function.
[0108] L all =L C +L D_totall
[0109] The network structure and parameter settings of the feature extractor F, domain discriminator D, and state classifier C are shown in Table 1. The following examples illustrate the meaning of the parameters in Table 1. For example, the convolutional network parameters 64×1×2 / 3 in Table 1 represent a convolution kernel length of 64, a convolution kernel depth of 1, a number of convolution kernels of 2, and a convolution step of 3; the network parameters 2×1 / 2 of the maximum pooling layer represent taking the maximum value between two adjacent values, with a step of 2; the network parameters 632×16 of the fully connected layer represent the input data length of the upper fully connected layer of 632, the output length of 16, and the number of network weight parameters of 632×16+16; ReLU represents the linear rectification function; Smax represents the Softmax function; and the meanings of other network parameters are similar.
[0110] Table 1
[0111]
[0112] In step S3, the parameter definitions of each module in the progressive transfer learning network architecture proposed in this embodiment are shown in Table 2.
[0113] Table 2
[0114]
[0115] The optimization objectives are: 1) minimize the classification error of source domain and target domain samples; 2) minimize the difference between the sample distributions. By minimizing the loss function L of MK-MMD MMD , minimize the total loss function L of the discriminator D_total Get the optimal network parameters:
[0116]
[0117] In the above formula, we have:
[0118]
[0119] In the formula, Non means that this item does not exist. The optimization methods of each network parameter are shown in the following formula:
[0120]
[0121] In the formula, n=0,1,2, which is the parameter The training process is to first train the discriminator D1 and then train the feature extractor at the same time. and Retrain the discriminator D2 and then train the feature extractor and And so on, the training of the discriminator and feature extractor is completed.
[0122] After training, the optimal network parameters are obtained Update weight μ m The method for (m=1,2,3) is as follows:
[0123]
[0124] The fault category of the target sample is estimated through MK_MMD clustering, and the training optimization parameter θC is based on the following formula:
[0125]
[0126] The overall training process is shown in the pseudo code in Table 3:
[0127] Table 3
[0128]
[0129] In step S4, the model parameters are iteratively updated according to Table 3, and the proposed method model is trained to obtain the optimal network parameters. Finally, the target domain fault samples without labels (different from the target domain samples in the test data) are input into the optimized model to predict the fault type, and the performance of the model is evaluated by calculating the predicted fault diagnosis accuracy. The fault diagnosis formula is as follows:
[0130]
[0131] where Y test is the predicted label of the mechanical state, X test represents the fault sample in the target domain, and max(·) extracts the label with the highest fault probability.
[0132] Specifically, if Figure 2 As shown in Figure 2, the process of real-time diagnosis of mechanical faults is as follows:
[0133] (1) Peripheral initialization: First, initialize the peripherals of the device to ensure that all external devices are ready.
[0134] (2) Collect real-time vibration data: Collect vibration data of mechanical equipment in real time through sensors and other equipment.
[0135] (3) Remove standby data: Set a threshold to identify and remove vibration data when in standby state to reduce interference.
[0136] (4) Eliminate idling data: Use RMS (root mean square) and other methods to identify and eliminate idling vibration data to improve data accuracy.
[0137] (5) Eliminate amplitude mutation data: Identify and eliminate amplitude mutation parts in vibration data through methods such as autocorrelation analysis to avoid misjudgment.
[0138] (6) Eliminate the impact of speed changes: Use methods such as calculation order tracking to eliminate the impact of vibration data caused by speed changes of mechanical equipment.
[0139] (7) After the above steps, more accurate and reliable vibration data is obtained and input into the fault identification model based on the progressive transfer learning network: the transfer learning network is used to conduct in-depth analysis of the pre-processed vibration data to identify the type of mechanical fault.
[0140] (8) Output the probability of mechanical fault type: The network outputs the probability of each mechanical fault type, providing a quantitative basis for diagnosis.
[0141] (9) Serial port output displays the fault type: The diagnostic results are output through the serial port to display the specific mechanical fault type for subsequent processing.
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network, characterized by: The method comprises the following steps: S1. Collect raw signal data in a laboratory environment and a real industrial environment, and preprocess the data; S2. Establish a fault identification model based on a progressive transfer learning network architecture, wherein corresponding features are extracted by at least several source domain feature extractors and target domain feature extractors, and the source domain data features are gradually mapped to the target data feature distribution domain for fault classification; S3. Combining the loss function in the propagation process, the optimal network parameters are obtained with the goal of minimizing the classification error of source and target domain samples and minimizing the inter-domain difference of sample distribution; S4. Collect real-time signal data in a real industrial environment and input it into the optimized fault identification model based on the progressive transfer learning network architecture. The fault identification model outputs the fault type.
2. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 1, characterized in that: In step S1, the data collected in the laboratory is defined as source domain data, and the data collected in the real industrial environment is defined as target domain data. The source domain data has fault labels, while the target domain data does not have fault labels. The data preprocessing process includes: denoising and filtering the original signal, removing outliers, and completing the original signal, and then intercepting the signal through a sliding window to produce the source domain dataset and the target domain dataset respectively.
3. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 2, characterized in that: In step S2, the fault recognition model based on the progressive transfer learning network includes several feature extractors F, several domain discriminators D and at least one state classifier C. Among them, the feature extractor is divided into source domain feature extractor F Sm and target domain feature extractor F Tm , m=1,2,3, including the following process: S21, through the source domain feature extractor F Sm and target domain feature extractor F Tm Extract features of different scales from the source domain dataset and the target domain dataset respectively; S22, domain identifier D m For the source domain feature extractor F Sm and target domain feature extractor F Tm Each layer of extracted features is used to identify the source and constrain the distribution differences of features in different dimensions of the source or target domain; S23, through the flattening operation, the source domain feature extractor F Sm and target domain feature extractor F Tm Finally, the extracted multi-dimensional features are aligned; S24, aligned target domain features and source domain features V S and target domain features V T Input to the state classifier C for fault type identification.
4. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 3, characterized in that: In step S21, the source domain feature extractor F Sm and target domain feature extractor F Tm They all include three-level feature mapping neural networks. Each level of feature mapping neural network extracts features of different scales. Each level of feature mapping neural network includes convolution layer, activation layer, normalization layer and pooling layer. For the input source domain dataset and target domain dataset The extraction process is expressed as: In the formula Represents the source domain feature extractor F at the mth level Sm and target domain feature extractor F Tm Extracted features, Represents the m-th level source domain feature extractor F Sm and target domain feature extractor F Tm network parameters.
5. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 4, characterized in that: In step S22, the source domain feature extractor F Sm and target domain feature extractor F Tm Each layer outputs f Sm and f Tm Input into the domain identifier D respectively m To identify the source of the sample, its output data type is Boolean. If the sample comes from the source domain, the domain identifier D m The output is 0; If the sample comes from the target domain, the domain discriminator D m The output is 1; the scope and contribution of each domain classifier constraint are different, and different weights μ are assigned to each domain classifier m .
6. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 5, characterized in that: In step S23, the source domain feature extractor F Sm and target domain feature extractor F Tm The final extracted multidimensional feature f S3 and f T3 Use the flatten operation Flatten(·) to align the source domain features V S and target domain features V T : The flattening operation unifies the dimensions of the extracted multidimensional features.
7. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 6, characterized in that: In step S24, the state classifier C is composed of a fully connected layer, an activation layer, and a normalization layer. In the state classifier C, the distance between the feature distributions is measured by the maximum mean deviation, and the target domain samples are classified into the state classifier C. Tags Make a prediction: In the formula is the source domain sample The fault label, represents the set clustering range, B(·) represents Boolean operation, If the sample If the feature is mapped to the set cluster range, its value is 1; if it is not mapped to the cluster range, its value is 0. K represents the number of all fault types.
8. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 7, characterized in that: In step S3, the network parameters of the initial source domain feature extractor are adjusted to minimize the classification error of source domain and target domain samples and minimize the inter-domain difference of sample distribution. Network parameters of the target domain feature extractor Network parameters of the domain discriminator The network parameters θ of the state classifier C The optimization is performed by minimizing the loss function L of the maximum average deviation MMD , minimize the total loss function L of the discriminator D_total Get the optimal network parameters: In the above formula, we have: In the formula, Non means that this item does not exist. The optimization methods of each network parameter are as follows: Step_(2n+1): Step_(2n+2): In the formula, n=0,1,2, which is the parameter The training process is to first train the discriminator D1 and then train the feature extractor F at the same time. S1 and F T1 , then train the discriminator D2, and then train the feature extractor F S2 and F T2 , and so on, guiding the discriminator and feature extractor to complete the training; After training, the optimal network parameters are obtained Update weight μ m The method is as follows: Step_7: The fault category of the target sample is estimated by clustering with the maximum mean deviation, and the optimization parameter θ is trained according to the following formula C : Step_8: In the above process, L all is the total loss function of the fault identification model, where L MMD The loss function representing the maximum mean deviation, L D_total is the overall loss function of the domain discriminator, L C is the overall loss function of the state classifier.
9. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 8, characterized in that: The total loss function L of the fault identification model all Expressed as: L all =L C +L D_totall Among them, the total loss function L of the discriminator D_total The definition is as follows: μ m Represents the weight of each scale domain discriminator, and the loss function of each scale domain discriminator The calculation formula is as follows: Where N S and N T Represent the number of source domain samples and the number of target domain samples, d i d j are the labels of the source domain and the target domain, respectively, and the domain label d of the source sample i = 0, the domain label d of the target sample j =1, then: The state classifier C uses cross entropy as the loss function L C , which is expressed as follows: Where ψ S , ψ T Represent the source domain samples {X S ,Y S } and target domain samples {X T ,Y T }, where Y T is the initial predicted label of the target sample, H(·) represents the cross entropy loss function; The loss function L of the maximum mean deviation MMD The calculation method is as follows: κ(·) represents the Gaussian kernel.
10. The method for intelligent diagnosis of mechanical faults based on a progressive transfer learning network according to claim 8, characterized in that: In step S4, after obtaining the optimal network parameters After that, the target domain fault samples without labels are input into the optimized model to identify the fault type, and the final output is expressed as: where Y test is the predicted label of the mechanical state, X test represents the fault sample in the target domain, and max(·) extracts the label with the highest fault probability.
Citation Information
Cited By
Underwater robot propeller type thruster cross-domain intelligent fault diagnosis method
CN120705742A