Cross-working-condition fault diagnosis method for production equipment driven by intelligent model
By sharing the feature extractor and the local maximum mean difference loss function, combined with the private feature extractor and classifier alignment loss function, the problem of misalignment of feature distribution in cross-working fault diagnosis of production equipment is solved, and higher diagnostic accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510393854.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing cross-working fault diagnosis methods for production equipment are degraded when the data distribution varies greatly under different working conditions, and traditional methods cannot fully mine the accurate information of multi-working data, resulting in low accuracy.
The shared feature extractor is used to extract the common features of the source domain and the target domain, combined with the private feature extractor and the domain-specific classifier, and the local maximum mean difference loss function and the classifier alignment loss function are achieved to achieve the consistency of feature distribution alignment and prediction labels, and improve diagnostic accuracy.
It improves the accuracy and performance of cross-working fault diagnosis of production equipment, reduces the misclassification rate of boundary areas, and enhances the model's adaptability under different working conditions.
Smart Images

Figure CN120336750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet big data and new generation information technologies, and particularly relates to a cross-condition fault diagnosis method for production equipment driven by an intelligent model. Background Art
[0002] In an actual manufacturing environment, for key components of production equipment such as gearboxes, faults often occur due to manufacturing or assembly defects, lubrication problems, harsh operating conditions, inadequate maintenance, etc. If the faults cannot be detected and diagnosed in a timely manner, it may lead to equipment shutdown, reduced production efficiency, causing serious economic losses, and even causing major casualties. Therefore, timely and accurate diagnosis of faults in key components of production equipment is of great significance for ensuring the safe and stable operation of equipment, improving production efficiency, and promoting intelligent manufacturing.
[0003] The massive data generated by the operation of mechanical equipment provides rich resources for equipment fault diagnosis. Therefore, a large number of deep learning algorithms have been widely applied to mechanical equipment fault diagnosis. However, their application depends on a large amount of labeled data, and labeled data is often scarce and costly to obtain in the industrial field. Moreover, since the operating parameters of key components of production equipment under different working conditions (such as rotational speed, load, etc.) will significantly affect the characteristic distribution of vibration signals, traditional deep learning models are usually trained under specific working conditions. When applied to different working conditions, due to the change in data distribution, the diagnostic performance of the model often drops significantly.
[0004] To solve the above problems, the fault diagnosis technology based on transfer learning has emerged. It improves the performance of tasks in the target domain by transferring knowledge from one or more domains to another, which is of great significance for the fault diagnosis of key components of production equipment in cross-condition environments. However, most current research focuses on single-source domain domain adaptation fault diagnosis. However, the key components of production equipment are often in a changing working condition environment during actual operation. The operating conditions under different working conditions lead to significant differences in data distribution, and the knowledge of a single source domain is difficult to comprehensively cover the complex features of the target domain. Although the cross-condition fault diagnosis based on multi-source domain domain adaptation shows good transfer learning ability, in order to further improve the diagnostic accuracy, some problems still need to be solved: 1) The existing methods do not adequately distinguish specific features between domains, which will result in a broad mapping relationship between the source domain and the target domain, unable to fully mine the precise information within the two domains, resulting in a decline in the performance of cross-condition fault diagnosis. 2) Traditional difference metrics mainly learn global domain transfer, that is, align the marginal distributions of the source domain and the target domain, without considering the relationship between two sub-domains in the same category of different domains, resulting in unsatisfactory transfer learning performance and low accuracy of cross-condition fault diagnosis. Therefore, how to make full use of the multi-condition labeled data generated by the operation of key components of production equipment and improve the diagnostic performance of the model in other working conditions with large differences in data distribution from these working conditions and without labeled data is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide a cross-condition fault diagnosis method for production equipment driven by an intelligent model. First, extract the common features of all source domains and the target domain through a shared feature extractor. Secondly, calculate the local maximum mean discrepancy loss function between the private features of the data samples of each pair of source domains and the target domain, so as to train the shared feature extractor and the corresponding private feature extractors to align the feature distributions of the source domain and the target domain when extracting features. Finally, after the domain-specific classifiers output the prediction labels of the source domain and the target domain, use the L1 distance to construct a classifier alignment loss function that minimizes the differences in the outputs of all domain-specific classifiers for the target domain data, thereby improving the performance and accuracy of cross-condition fault diagnosis of production equipment.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A cross-condition fault diagnosis method for production equipment driven by an intelligent model, comprising:
[0008] S1: Obtain the data sets of the target domain and N source domains; the data set of the target domain includes the data samples of the target domain, and the data sets of the source domains include the data samples of the source domains and the corresponding true fault labels;
[0009] S2: Input the data samples of the target domain into the trained fault diagnosis model, and output the corresponding fault prediction labels;
[0010] S201: Combine the target domain with N source domains respectively to form N pairs of source domains and target domains, and set corresponding sub-structure networks for each pair of source domains and target domains; The sub-structure network includes a private feature extractor, a domain-specific classifier, and a domain-specific distribution alignment module;
[0011] S202: Randomly extract the data samples of the source domain and the corresponding true fault labels, as well as the data samples of the target domain from N pairs of source domains and target domains; Input the data samples of N pairs of source domains and target domains into the shared feature extractor for feature extraction to obtain the common features of the data samples of N pairs of source domains and target domains;
[0012] S203: Input the common features of the data samples of each pair of source domains and target domains into the corresponding private feature extractor for feature extraction to obtain the private features of the data samples of this pair of source domains and target domains;
[0013] S204: Input the private features of the data samples of each pair of source domains and target domains into the corresponding domain-specific classifier for classification to obtain the source domain fault prediction labels and target domain fault prediction labels of the data samples of this pair of source domains and target domains;
[0014] S205: Use the private features of the data samples of each pair of source domains and target domains, the target domain fault prediction labels, and the true fault labels of the source domain through the corresponding domain-specific distribution alignment module to calculate the local maximum mean discrepancy loss function; The local maximum mean discrepancy loss function is used to train the shared feature extractor and the corresponding private feature extractor to achieve feature distribution alignment between the source domain and the target domain when extracting features;
[0015] S206: Calculate the domain-specific distribution alignment loss function through the local maximum mean discrepancy loss functions of N pairs of source domains and target domains;
[0016] S207: Calculate the domain-specific classifier loss function through the source domain fault prediction labels and true fault labels of all source domains;
[0017] S208: Calculate the classifier alignment loss function by minimizing the difference in the target domain fault prediction labels of the target domain data samples by all domain-specific classifiers;
[0018] S209: Calculate the overall optimization loss function through the domain-specific distribution alignment loss function, the domain-specific classifier loss function, and the classifier alignment loss function, and reversely optimize the parameters of the shared feature extractor in the fault diagnosis model and the private feature extractors and domain-specific classifiers of all sub-structure networks;
[0019] S210: Repeat steps S202 to S209 to iteratively train the fault diagnosis model until the model converges or reaches the maximum number of iterations;
[0020] Among them, input the data samples of the target domain into the trained fault diagnosis model. After passing through the shared feature extractor and the private feature extractors and domain-specific classifiers of all sub-structure networks in sequence, take the average of the outputs of the domain-specific classifiers of all sub-structure networks as the final fault prediction label;
[0021] S3: Take the fault prediction label output by the fault diagnosis model as the fault diagnosis result of the data samples of the target domain.
[0022] Preferably, in step S1, after obtaining the data sets of the target domain and N source domains, first perform data augmentation and data standardization processing on the data samples of the target domain and N source domains, and then execute the subsequent steps.
[0023] Preferably, in step S202, the shared feature extractor G(·) includes three one-dimensional convolutional modules;
[0024] The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*64, a stride of 16, and a padding of 24, a BN layer with 32 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2;
[0025] The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 64 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2;
[0026] The third one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 128 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0027] Preferably, in step S203, the private feature extractor F j (·) includes two one-dimensional convolutional modules;
[0028] The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 256 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2;
[0029] The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 512 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0030] Preferably, in step S204, the domain-specific classifier C j (·) includes two fully connected layers connected end to end;
[0031] The first fully connected layer includes neurons with 512 on both sides, a Dropout layer with a probability of 0.3, and a ReLU layer;
[0032] The second fully connected layer includes neurons with 512 and the number equal to the number of fault categories on both sides.
[0033] Preferably, in step S205, the formula for the domain-specific distribution alignment module to calculate the local maximum mean discrepancy loss function using the private features of each pair of source domain and target domain data samples, the target domain fault prediction label, and the true fault label of the source domain is expressed as:
[0034]
[0035] In the formula: represents an unbiased estimator of calculating the local maximum mean discrepancy, that is, the local maximum mean discrepancy loss function, P s , P t respectively represent the data sample feature distributions of the source domain and the target domain, N s , N t respectively represent the number of samples in the source domain and the target domain, K represents the number of fault categories, H represents the reproducing kernel Hilbert space, F(·) represents the private features of the source domain data samples or the target domain data samples, and respectively represent the data samples of the source domain and the target domain and belong to the weight of the kth type of fault; for the data samples in the source domain, is the kth value of the one-hot encoded vector corresponding to the true fault label of the data sample x m , for the unlabeled data samples in the target domain, is the kth value of the target domain fault prediction label m obtained by the data sample x through the domain-specific classifier; represents the sum of the probabilities that all samples in the domain D belong to the category k; k(·,·) represents the Gaussian kernel function, and σ represents the kernel width.
[0036] Preferably, in step S206, the domain-specific distribution alignment loss function L lmmd is calculated by the following formula:
[0037]
[0038] In the formula: represents the ith data sample of the jth source domain, represents the common feature of the data sample , Represents the private features of the data sample ; Represents the i-th data sample of the target domain, Represents the data sample 's common features, Represents the data sample 's private features; Represents the calculation of and The local maximum mean difference loss function between them.
[0039] Preferably, in step S207, the domain-specific classifier loss function L is calculated by the following formula C :
[0040]
[0041] In the formula: Represents the true fault label corresponding to the data sample , Represents the predicted label of the j-th domain-specific classifier for the source domain data sample , J(·,·) represents the calculation of the cross-entropy loss function, Represents the mathematical expectation of the source domain data sample.
[0042] Preferably, in step S208, the classifier alignment loss L is calculated by the following formula disc :
[0043]
[0044] In the formula: C i (F i (G(x k )))、C j (F j (G(x k ))) respectively represent the predicted probability vectors of the target domain data sample x k output by the i-th and j-th domain-specific classifiers, Represents the mathematical expectation of the target domain data sample.
[0045] Preferably, in step S209, the overall optimization loss function is calculated by the following formula:
[0046] L total = L C + αL lmmd + βL disc ;
[0047] In the formula: L C Represents the domain-specific classifier loss function, Llmmd Denote the domain - specific distribution alignment loss function, \(L\). disc Denote the classifier alignment loss, where \(\alpha\) and \(\beta\) are set trade - off parameters;
[0048] Among them, the formula for reverse - optimizing the parameters of the shared feature extractor, all private feature extractors, and domain - specific classifiers in the fault diagnosis model is expressed as:
[0049]
[0050] In the formula: \(\theta\) G ,\(\theta\) F ,\(\theta\) C are the parameters to be trained for the shared feature extractor, private feature extractors, and domain - specific classifiers respectively, is the partial - derivative algorithm, and \(\eta\) is the learning rate.
[0051] Compared with the prior art, the intelligent - model - driven cross - operating - condition fault diagnosis method for production equipment in the present invention has the following beneficial effects:
[0052] In the present invention, the shared feature extractor extracts common feature representations from the data samples of all source domains and the target domain, maps the vibration signals after data processing from the original feature space to the common feature space, captures the common laws between multiple source domains and the target domain. Compared with separately designing feature extraction networks, the shared feature extractor is simple in structure and can save computing resources at the same time. Moreover, the first - layer convolution of the shared feature extractor is designed with a wide convolution kernel with a kernel size of 64 to obtain a larger receptive field and good anti - noise performance.
[0053] After the shared feature extractor, the network of the present invention adopts a multi-branch sub-structure network design, and a private feature extractor, a domain-specific classifier, and a domain-specific distribution alignment module are designed on each sub-structure network. First, the common features of each pair of source domain and target domain are mapped into a specific feature space through the private feature extractor, so as to focus on capturing the data distribution characteristics of each pair of source domain and target domain, fully mine the accurate information within the two domains, and improve the cross-condition fault diagnosis performance. Secondly, in the domain-specific distribution alignment module, the locally maximum mean discrepancy loss function between the private features of the data samples of each pair of source domain and target domain is calculated by using the pseudo-label (i.e., the target domain fault prediction label)-driven locally maximum mean discrepancy algorithm. This loss function is used to train the shared feature extractor and the corresponding private feature extractor to achieve feature distribution alignment between the source domain and the target domain when extracting features. The traditional maximum mean discrepancy algorithm only aligns the marginal distributions of the two domains and does not consider the relationship between the two sub-domains in the same category of different domains. However, the locally maximum mean discrepancy adopted by the present invention makes full use of the target domain fault prediction label and introduces the class weight, calculates the mean discrepancy between the local distributions at the sub-domain level, enables the samples of the same category to be aligned in different domains, fully aligns the conditional distributions in the unsupervised scenario, finely captures the differences between the source domain and the target domain, and improves the accuracy of cross-condition fault diagnosis of production equipment. Finally, after the domain-specific classifier outputs the prediction labels of the source domain and the target domain, the classifier alignment loss function that minimizes the differences between the outputs of all domain-specific classifiers for the target domain data is constructed by using the L1 distance, which solves the problem that the domain-specific classifier is trained based on the source domain samples with different data distributions, and the prediction results of the target domain samples (especially the target domain samples near the decision boundary) are prone to divergence, resulting in an increase in the misclassification rate in the boundary region. And the classifier alignment loss function makes the prediction of the data samples of the target domain by each domain-specific classifier as consistent as possible, forcing different classifiers to form a consensus on the samples in the decision boundary region, thereby further improving the performance and accuracy of cross-condition fault diagnosis of production equipment. Description of the Drawings
[0054] In order to make the objectives, technical solutions, and advantages of the invention more clear, the present invention will be further described in detail below with reference to the drawings, where:
[0055] Figure 1 and Figure 2 are the framework diagram and diagnostic process of the intelligent model-driven cross-condition fault diagnosis method for production equipment.
[0056] Figure 3 is the grouped bar chart of the test accuracy of the PHM2009 dataset.
[0057] Figure 4 is the accuracy confusion matrix of the PHM2009 dataset.
[0058] Figure 5 It is the feature visualization result of the PHM2009 dataset.
[0059] Figure 6 It is the grouped bar chart of the test accuracy of the WT-planetary gearbox dataset.
[0060] Figure 7 It is the accuracy confusion matrix of the WT-planetary gearbox.
[0061] Figure 8 It is the feature visualization result of the WT-planetary gearbox. Specific implementation manner
[0062] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present invention provided herein is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0063] It should be noted that like reference numerals and letters refer to like items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the product of the invention is commonly placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0064] The following is a more detailed description through specific embodiments:
[0065] Embodiment:
[0066] This embodiment discloses a cross-condition fault diagnosis method for intelligent model-driven production equipment.
[0067] As Figure 1 and Figure 2 shown, the cross-condition fault diagnosis method for intelligent model-driven production equipment includes:
[0068] S1: Obtain the data sets of the target domain and N source domains; the data set of the target domain includes the data samples of the target domain, and the data sets of the source domains include the data samples of the source domains and the corresponding true fault labels;
[0069] S2: Input the data samples of the target domain into the trained fault diagnosis model and output the corresponding fault prediction labels;
[0070] S201: Combine the target domain with N source domains respectively to form N pairs of source domains and target domains, and set corresponding sub-structure networks for each pair of source domains and target domains; the sub-structure network includes a private feature extractor, a domain-specific classifier, and a domain-specific distribution alignment module;
[0071] S202: Randomly extract data samples of the source domain and corresponding true fault labels, as well as data samples of the target domain from N pairs of source domains and target domains; input the data samples of N pairs of source domains and target domains into the shared feature extractor for feature extraction to obtain the common features of the data samples of N pairs of source domains and target domains;
[0072] S203: Input the common features of the data samples of each pair of source domains and target domains into the corresponding private feature extractor for feature extraction to obtain the private features of the data samples of this pair of source domains and target domains;
[0073] S204: Input the private features of the data samples of each pair of source domains and target domains into the corresponding domain-specific classifier for classification to obtain the source domain fault prediction labels and target domain fault prediction labels of the data samples of this pair of source domains and target domains;
[0074] Among them, the target domain fault prediction label needs to be applied when calculating the local maximum mean discrepancy loss function and the classifier alignment loss function.
[0075] S205: Use the private features of the data samples of each pair of source domains and target domains, the target domain fault prediction label, and the true fault label of the source domain through the corresponding domain-specific distribution alignment module to calculate the local maximum mean discrepancy loss function; the local maximum mean discrepancy loss function is used to train the shared feature extractor and the corresponding private feature extractor to achieve feature distribution alignment between the source domain and the target domain when extracting features;
[0076] S206: Calculate the domain-specific distribution alignment loss function through the local maximum mean discrepancy loss functions of N pairs of source domains and target domains;
[0077] S207: Calculate the domain-specific classifier loss function through the source domain fault prediction labels and true fault labels of N pairs of source domains;
[0078] S208: Calculate the classifier alignment loss function by minimizing the difference in the target domain fault prediction labels of all domain-specific classifiers for the target domain data samples;
[0079] S209: Calculate the overall optimization loss function through the domain-specific distribution alignment loss function, the domain-specific classifier loss function, and the classifier alignment loss function, and reverse-optimize the parameters of the shared feature extractor in the fault diagnosis model and the private feature extractors and domain-specific classifiers of all sub-structure networks;
[0080] S210: Repeat steps S202 to S209 to iteratively train the fault diagnosis model until the model converges or reaches the maximum number of iterations;
[0081] Among them, input the data samples in the target domain into the trained fault diagnosis model. After passing through the shared feature extractor, the private feature extractors of all sub-structure networks, and the domain-specific classifiers in sequence, take the average of the outputs of the domain-specific classifiers of all sub-structure networks as the final fault prediction label;
[0082] S3: Take the fault prediction label output by the fault diagnosis model as the fault diagnosis result of the data samples in the target domain.
[0083] The algorithm code of the cross-condition fault diagnosis method for production equipment driven by an intelligent model is shown in Table 1.
[0084] Table 1 Algorithm Pseudocode
[0085]
[0086]
[0087] The workflow of this method includes three stages: data preparation, model training, and model testing. First, collect the vibration signals of the equipment under different working conditions and perform label annotation according to the health status of the equipment to obtain a labeled data set; in the model training stage, input multiple labeled source domain data and unlabeled target domain training data into the model for training; in the model testing stage, use the unlabeled target domain test data to input the trained model, and compare the predicted label output by the model with the actual label to calculate the accuracy. It should be noted that in the data preparation stage, although the true labels of the target domain samples will be labeled, these label information do not participate in the model training process and are only used to finally evaluate the classification accuracy of the model.
[0088] The present invention extracts common feature representations from the data samples of all source domains and target domains through a shared feature extractor, maps the vibration signals after data processing from the original feature space to the common feature space, and captures the common laws between multiple source domains and target domains. Compared with separately designing feature extraction networks, the shared feature extractor is simple in structure and can save computing resources at the same time. Moreover, the first layer convolution of the shared feature extractor adopts a wide convolution kernel with a kernel size of 64 to obtain a larger receptive field and good anti-noise performance.
[0089] After the shared feature extractor, the network of the present invention adopts a multi-branch sub-structure network design, and a private feature extractor, a domain-specific classifier, and a domain-specific distribution alignment module are designed on each sub-structure network. First, the common features of each pair of source domain and target domain are mapped into a specific feature space through the private feature extractor, so as to focus on capturing the data distribution characteristics of each pair of source domain and target domain, fully mine the accurate information within the two domains, and improve the cross-condition fault diagnosis performance. Secondly, in the domain-specific distribution alignment module, the local maximum mean discrepancy loss function between the private features of the data samples of each pair of source domain and target domain is calculated by using the pseudo-label (i.e., the target domain fault prediction label) driven local maximum mean discrepancy algorithm. This loss function is used to train the shared feature extractor and the corresponding private feature extractor to achieve the feature distribution alignment between the source domain and the target domain when extracting features. The traditional maximum mean discrepancy algorithm only aligns the marginal distributions of the two domains and does not consider the relationship between the two sub-domains in the same category of different domains. However, the local maximum mean discrepancy adopted by the present invention makes full use of the target domain fault prediction label and introduces the class weight, calculates the mean discrepancy between the local distributions at the sub-domain level, enables the samples of the same category to be aligned in different domains, fully aligns the conditional distributions in the unsupervised scenario, finely captures the differences between the source domain and the target domain, and improves the accuracy of cross-condition fault diagnosis of production equipment. Finally, after the domain-specific classifier outputs the prediction labels of the source domain and the target domain, the classifier alignment loss function that minimizes the differences between the outputs of all domain-specific classifiers for the target domain data is constructed by using the L1 distance, which solves the problem that the domain-specific classifier is trained based on the source domain samples with different data distributions, and the prediction results for the target domain samples (especially the target domain samples near the decision boundary) are prone to divergence, resulting in an increase in the misclassification rate in the boundary region. And the classifier alignment loss function makes the prediction of the data samples of the target domain by each domain-specific classifier as consistent as possible, forcing different classifiers to form a consensus on the samples in the decision boundary region, thereby further improving the performance and accuracy of cross-condition fault diagnosis of production equipment.
[0090] To better introduce the technical solution of the present invention, this embodiment is described through the following several parts.
[0091] I. Definition of Source Domain and Target Domain
[0092] Definition is defined as the labeled source domain, and each source domain j follows the data distribution, N represents the number of source domains, where represents the sample data of the j-th source domain, N j is the number of samples, represents the true label of the sample corresponding to the source domain, and k is the number of categories of health states.
[0093] Definition is the target domain without labels and follows the data distribution of P t (x, y), where N t is the number of samples in the target domain, and the labels Y of the target domain t are unknown.
[0094] The source domain and the target domain satisfy: 1) All source domains and the target domain should have the same label space. 2) There are differences between the data distribution of each source domain and the data distribution of the target domain, that is, P sj (x, y) ≠ P t (x, y). The present invention uses multi-source labeled data and unlabeled data in the target domain to train a model so that the prediction performance of the model on the target domain is optimal.
[0095] The present invention collects the original vibration signals of multiple labeled source domains and unlabeled target domains, performs data augmentation and standardization on all signals, and realizes the cross-reuse of sample sequences to prevent the model from being sensitive to the scale of input data, avoid the model relying on features with large numerical values, and improve the migration and generalization ability of the model.
[0096] II. Data Augmentation and Data Standardization
[0097] In this embodiment, after obtaining the data sets of the target domain and N source domains, first perform data augmentation and data standardization processing on the data samples of the target domain and N source domains, and then perform subsequent steps;
[0098] Data augmentation refers to setting an offset to perform overlapping sampling on the input signal to realize the cross-reuse of data sample sequences;
[0099] Data standardization means that the data sample (original vibration signal data) x i ′ is processed by Z-Score standardization;
[0100] The formula is expressed as:
[0101]
[0102] In the formula: x i is the data sample after standardization processing, μ(x i ′) is the mean value, and σ(x i ′) is the standard deviation.
[0103] III. Shared Feature Extractor
[0104] Establish a shared feature extractor G(·) to extract common feature representations from all source domains and the target domain, map the vibration signals after data processing from the original feature space to the common feature space, and capture the common laws between multiple source domains and the target domain.
[0105] The shared feature extractor G(·) includes three one-dimensional convolutional modules (connected end to end in sequence).
[0106] The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*64, a stride of 16, and a padding of 24 (connected end to end in sequence), a BN layer with 32 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0107] The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1 (connected end to end in sequence), a BN layer with 64 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0108] The third one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1 (connected end to end in sequence), a BN layer with 128 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0109] IV. Private Feature Extractor
[0110] The private feature extractor is essentially a 1DCNN feature extraction module that maps each pair of source domain and target domain to a specific feature space to focus on capturing the data distribution characteristics of each pair of source domain and target domain.
[0111] Private Feature Extractor F j (·) includes two one-dimensional convolutional modules (connected end to end in sequence).
[0112] The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1 (connected end to end in sequence), a BN layer with 256 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0113] The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1 (connected end to end in sequence), a BN layer with 512 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
[0114] V. Domain-Specific Classifier
[0115] The domain-specific classifier receives the feature representation from the domain-specific distribution alignment module, outputs the prediction probabilities of each category through a fully connected layer and a softmax activation function, and feeds back the target domain fault prediction label predicted for the target domain to the domain-specific distribution alignment module.
[0116] Domain-Specific Classifier C j (·) includes two fully connected layers connected end to end in sequence.
[0117] The first fully connected layer includes neurons with 512 on each side connected end to end in sequence, a Dropout layer with a probability of 0.3, and a ReLU layer;
[0118] The second fully connected layer includes neurons with 512 on one side and the number equal to the number of fault categories on the other side.
[0119] VI. Domain-Specific Distribution Alignment Module
[0120] Essentially a 1DCNN feature extraction module, which maps each pair of source domain and target domain into a specific feature space to focus on capturing the data distribution characteristics of each pair of source domain and target domain.
[0121] In this embodiment, the formula for the domain-specific distribution alignment module to calculate the Local Maximum Mean Discrepancy (LMMD) loss function using the private feature of the data sample of each pair of source domain and target domain, the fault prediction label of the target domain, and the true fault label of the source domain is expressed as:
[0122]
[0123] In the formula: represents the unbiased estimator of calculating the local maximum mean difference, that is, the local maximum mean difference loss function, P s , P t respectively represent the data sample feature distributions of the source domain and the target domain, N s , N t respectively represent the number of samples in the source domain and the target domain, K represents the number of fault categories, H represents the Reproducing Kernel Hilbert Space, F(·) represents the private feature of the source domain data sample or the target domain data sample, and respectively represent the data samples of the source domain and the target domain and belong to the weight of the k-th type of fault; for the data sample of the source domain, is the k-th value of the one-hot encoded vector corresponding to the true fault label of the data sample x m , for the unlabeled data sample of the target domain, is the k-th value of the target domain fault prediction label obtained by the data sample x m after passing through the domain-specific classifier, where the target domain fault prediction label is a pseudo label; is a pseudo label; represents the sum of the probabilities that all samples in the domain D belong to the category k; k(·,·) represents the Gaussian kernel function, and σ represents the kernel width.
[0124] When training the shared feature extractor and the corresponding private feature extractor, the feature distribution alignment between the source domain and the target domain is achieved when extracting features by minimizing the local maximum mean discrepancy loss function for the shared feature extractor and the corresponding private feature extractor.
[0125] VII. Loss Function
[0126] 1. Domain-Specific Distribution Alignment Loss Function
[0127] Specifically, the domain-specific distribution alignment loss function L is calculated by the following formula lmmd :
[0128]
[0129] In the formula: represents the i-th data sample in the j-th source domain, represents the data sample 's common feature, represents the data sample 's private feature; represents the i-th data sample in the target domain, represents the data sample 's common feature, represents the data sample 's private feature; represents calculating the local maximum mean discrepancy loss function between and .
[0130] 2. Domain-Specific Classifier Loss Function
[0131] The cross-entropy loss function is used to minimize the difference between the prediction result and the true label.
[0132] Specifically, the domain-specific classifier loss function L is calculated by the following formula C :
[0133]
[0134] In the formula: represents the true fault label corresponding to the data sample , represents the prediction label of the j-th domain-specific classifier for the source domain data sample , J(·,·) represents calculating the cross-entropy loss function, represents the mathematical expectation of the source domain data sample.
[0135] 3. Classifier Alignment Loss
[0136] Since the classifier is trained based on source domain samples with different data distributions, its prediction results for target domain samples (especially target domain samples near the decision boundary) are prone to divergence, leading to an increase in the misclassification rate in the boundary region. Although domain-specific distribution alignment achieves sufficient alignment of feature distributions, it fails to solve the problem of inconsistent predictions of the classifier for target domain samples. Therefore, the differences between all classifiers are minimized to force different classifiers to form a consensus on samples in the decision boundary region. That is, for the same target sample, the probability distribution differences output by all classifiers are calculated, and the difference loss function is constructed using the L1 distance.
[0137] Specifically, the classifier alignment loss L is calculated through the following formula disc :
[0138]
[0139] In the formula: C i (F i (G(x k ))), C j (F j (G(x k )) represent the prediction probability vectors of the target domain data sample x k output by the i-th and j-th domain-specific classifiers respectively, represents the mathematical expectation of the target domain data sample.
[0140] 4. Overall optimization loss function
[0141] In this embodiment, after extracting the common feature representations of all original vibration signals, each pair of source domain and target domain is mapped into a specific feature space to achieve distribution alignment. The sub-domain alignment algorithm and the classifier alignment algorithm are adopted to reduce the distribution offset between multiple source domains and the target domain, and realize cross-condition fault diagnosis. The three objectives during training the fault diagnosis model are: 1) minimizing the domain-specific classifier loss; 2) minimizing the difference metric loss; 3) minimizing the classifier alignment loss.
[0142] Specifically, the overall optimization loss function is calculated through the following formula:
[0143] L total = L C + αL lmmd + βL disc ;
[0144] In the formula: L C represents the domain-specific classifier loss function, L lmmd represents the domain-specific distribution alignment loss function, L disc represents the classifier alignment loss, and α, β are set trade-off parameters;
[0145] Among them, the random gradient descent algorithm (SGD) is used to inversely optimize the parameters of the shared feature extractor, all private feature extractors, and domain-specific classifiers in the fault diagnosis model;
[0146] The formula is expressed as:
[0147]
[0148] In the formula: θ G 、θ F 、θ C are the parameters to be trained for the shared feature extractor, private feature extractors, and domain-specific classifiers respectively, is the partial derivative algorithm, and η is the learning rate.
[0149] VIII. Experimental Description
[0150] To better illustrate the advantages of the technical solution of the present invention, the following experiments are disclosed in this embodiment.
[0151] In this experiment, two data sets were collected from two different gearbox experimental platforms for extensive experiments, and a variety of methods were used for comparative analysis.
[0152] 1. Data Description
[0153] 1.1 Industrial Gearbox Data Set PHM2009
[0154] The gearbox of this data set consists of two pairs of meshing gears, three shafts (input shaft IS, intermediate shaft ID, output shaft OS), and six bearings (three on the input side: IS, three on the output side: OS). An accelerometer is installed near each of the input shaft and the output shaft for vibration signal acquisition, and the sampling frequency is 66.67 kHz. The data set covers 8 compound fault modes, including typical failure forms such as gear fracture, eccentricity, and bearing damage, as shown in Table 2. PHM2009 contains five shaft speeds and two loads. In this study, only the input-end data of the last four shaft speeds under high load are considered, and the corresponding four working conditions are shown in Table 3. Three source domain scenarios are designed, including a total of 4 transfer learning tasks: A + B + C → D, A + B + D → C, A + C + D → B, B + C + D → A.
[0155] 1.2 WT - Planetary Gearbox Data Set
[0156] The data acquisition platform of this dataset consists of a motor, a planetary gearbox, a fixed-axis gearbox, and a load application device. The planetary gearbox is composed of a central sun gear and four planetary gears. Accelerometers are used to collect vibration signals in the x and y directions, and an encoder is used to collect data on the input shaft of the planetary gearbox. The sampling frequency is set to 48 kHz for both. The dataset covers five typical fault modes of the sun gear and eight different shaft speed conditions. The corresponding labels for the fault states and the condition numbers corresponding to the speeds are shown in Tables 4 and 5 respectively. In this study, the vibration signals in the x direction are used to design two-source domain scenarios, which include a total of five transfer learning tasks: E+F→H, F+G→I, G+H→J, H+I→K, I+J→L.
[0157] Table 2 Fault categories of the PHM2009 dataset
[0158]
[0159] G: Healthy; C: Cracked; E: Eccentric; Br: Damaged; B: Ball; I: Inner ring; O: Outer ring; Im: Unbalanced; Ks: Keyway damaged.
[0160] Table 3 Condition numbers of the PHM2009 dataset
[0161]
[0162] Table 4 Fault categories of the WT - planetary gearbox dataset
[0163]
[0164] B: Gear damaged; H: Healthy gear; M: Missing tooth; R: Tooth root fracture; W: Tooth surface wear
[0165] Table 5 Condition numbers of the WT - planetary gearbox dataset
[0166]
[0167] 2. Comparison methods
[0168] 1) Single-source optimal: Select a single source domain that is closest to the target domain distribution (with the highest diagnostic accuracy) from multiple source domains, and only use the data of this source domain for transfer learning, ignoring the information of other source domains. One-dimensional convolutional CNN: Directly apply the model trained with source domain data to the test of the target domain without using domain adaptation methods. Deep domain adaptation network DAN: Apply the difference metric algorithm MK-MMD to the convolutional network to achieve domain adaptation between the source domain and the target domain. Domain adversarial network DANN: Construct a domain classifier to distinguish the domain source of the input features, and use the gradient reversal layer to make the parameter update direction of the feature extractor opposite to the optimization goal of the domain classifier, forcing the feature extractor to generate features that confuse the domain classifier, thereby achieving alignment of the inter-domain distribution. Maximize classifier difference MCD: By maximizing the difference in the prediction of the target samples by two task classifiers, detect the samples in the target domain that are far from the source domain support region, and then let the generator minimize this difference, forcing the target features to be close to the source domain support region. This method effectively utilizes the task-specific decision boundary to improve cross-domain performance.
[0169] 2) Source combination: Combine the data of multiple source domains into a mixed source domain, and then use the single-source domain adaptation method for transfer. CNN com , DAN com , DANN com , MCD com are the corresponding source combination methods.
[0170] 3) Multiple sources: Construct the relationships between multiple source domains and the target domain, and make full use of the information of each source domain. ADACL: The network consists of a feature extractor, multiple classifiers, and a domain discriminator. The domain discriminator distinguishes the domain sources of all source domains and the target domain. The feature extractor generates domain-invariant features, and uses the domain adversarial and classifier alignment strategies to align multiple feature spaces. MSACL MK-MMD : Under the structure of MSACL, replace the LMMD loss of the domain-specific distribution alignment module with the multi-kernel maximum mean difference (MK-MMD) loss.
[0171] 3. Training settings
[0172] To ensure the fairness of comparison, all the above methods adopt a network structure and model parameters similar to the fault diagnosis model (MSACL) of the present invention. For the data under all working conditions of the PHM2009 dataset and the WT-planetary gearbox dataset, 300 samples are intercepted using a sliding window of size 2048, and the offset is set to 680. Therefore, 2400 samples are collected for each working condition of the PHM2009 dataset, and 1500 samples are collected for each working condition of the WT-planetary gearbox dataset. In the experiment, all source domain samples are used as the training set, 70% of the target domain samples are used as the training set, and 30% of the target domain is used as the test set. To suppress the noise signal in the early stage of training, a progressive training method is adopted, and the trade-off parameters of the model are set as:
[0173]
[0174] where θ = 10, p is a parameter that changes from 0 to 1 during the training process, p = i - 1 / I - 1, i is the current training epoch, and I is the total number of training epochs. The learning rate η of the network is 0.01, the momentum of SGD is 0.9, the total number of training epochs at one time is 30, and the batch size is 64.
[0175] 4. Analysis of Results of PHM2009 Dataset
[0176] 4.1 Analysis of Classification Accuracy
[0177] The test accuracy of the PHM2009 dataset and the corresponding grouped bar chart are shown in Table 6 and Figure 3 respectively. Under the single-source optimal and source combination methods, the diagnostic accuracy of the domain adaptation method is higher than that of CNN under each working condition, indicating that the domain adaptation algorithm can significantly alleviate the domain shift problem. Secondly, under each working condition, the diagnostic accuracy of the source combination method is higher than that of the corresponding single-source optimal method, indicating that multi-source domain information can alleviate the problem of insufficient single-source domain information. Then, the multi-source method has a higher diagnostic accuracy than the source combination method as a whole, indicating the importance of constructing the relationship between multiple source domains and the target domain and fully exploring the information of each source domain.
[0178] The network model proposed by the present invention achieves the highest average diagnostic accuracy of 98.37% in the transfer tasks under four working conditions, which fully shows that MSACL can effectively achieve cross-working condition fault diagnosis of multi-source domains. The average diagnostic accuracy of MSACL and MSACL MK-MMD is increased by 1.53% and 0.76% compared with the method ADACL that uses domain adversarial confusion for all source domains and target domains, proving the effectiveness of mapping the source domain and the target domain to specific feature spaces for alignment respectively. Under the working conditions of A + B + C → D, A + B + D → C, and A + C + D → B, MSACL is better than MSACL MK-MMDThe diagnostic accuracies were increased by 1.11%, 0.84%, and 1.25% respectively, indicating that the sub-domain alignment algorithm LMMD can capture the differences between the source domain and the target domain more granularly than the global domain alignment algorithm MK-MMD, effectively reducing the negative transfer of samples. Under B+C+D→A, MSACL has a 0.14% lower diagnostic accuracy than MSACL MK-MMD which indicates that this working condition may be more suitable for global domain alignment, but this value is relatively low and within an acceptable range.
[0179] Table 6 Test accuracies of the PHM2009 dataset
[0180]
[0181] 4.2 Confusion matrix analysis
[0182] The classification results of the target domain test data under four working conditions are presented in the confusion matrix, where the abscissa represents the predicted sample labels and the ordinate represents the true sample labels, as Figure 4 shown. There are a total of 2,880 diagnostic tasks for target domain compound fault test samples in the four tasks, and only 47 samples are misidentified. Particularly, under the A+C+D→B working condition, only 6 samples out of 720 compound fault test samples are misidentified, achieving a diagnostic accuracy of 99.17%. Among all the compound fault categories, category 2 is misdiagnosed as category 1 the most times, reaching a total of 17 times under the four tasks. This may be because the difference between the compound faults of category 2 and category 1 is only the difference between healthy and broken on the 32T gear, which is not as obvious as the differences between other categories.
[0183] 4.3 Feature visualization
[0184] The output features of the domain-specific classifier are reduced to two-dimensional space using t-distributed Stochastic Neighbor Embedding (t-SNE) for feature visualization. The feature visualization results of CNN com and MSACL on the PHM2009 dataset are as Figure 5 shown. As Figure 5 (a) CNN without domain adaptation comThe features extracted are messy, there is an obvious mixing phenomenon among the samples of each fault category, and there are certain differences between the source domain and target domain samples. For example, for the target domain samples of category 6, they are concentrated in the upper part, while the source domain samples are concentrated in the lower part; for the target domain samples of category 7, they are concentrated on the left side, while the source domain samples are concentrated in the upper part. Under the four working conditions studied, the source domain samples and target domain samples of MSACL form obvious eight clustering clusters, and the source domain samples and target domain samples corresponding to the same fault category can be well integrated together, with good inter-class discrimination and inter-domain fusion. Under the working conditions of A+B+C→D, A+C+D→B, and B+C+D→A, there is a small overlap between the samples of fault categories 1 and 2 in MSACL, indicating that there may be a phenomenon of difficult discrimination between category 1 and category 2, which is consistent with the analysis of the above confusion matrix.
[0185] 5. Results Analysis of WT - Planetary Gearbox Dataset
[0186] 5.1 Classification Accuracy Analysis
[0187] The test accuracy of the WT - planetary gearbox dataset and the corresponding grouped bar chart are shown in Table 7 and Figure 6 as follows, which is similar to the diagnostic results of the PHM2009 dataset. It should be noted that com CNN com and MCD
[0188] are 0.18% and 0.71% lower than the average diagnostic accuracy of CNN and MCD respectively. This may be because the source domain contains source domains with large distribution differences from the target domain, and simply combining the source domains will increase the distribution difference between the source domain and the target domain, resulting in a negative transfer phenomenon and causing the diagnostic accuracy of the model to decline. MK-MMD The MSACL proposed in the present invention achieves the highest average diagnostic accuracy of 98.97% in the fault diagnosis tasks under five working conditions, and reaches high accuracies of 100%, 99.78%, 99.11%, and 99.78% respectively under the working conditions of E+F→H, F+G→I, H+I→K, and I+J→L, further demonstrating the superiority of MSACL. The average diagnostic accuracy of MSACL is 1.46% and 0.49% higher than that of
[0189] Table 7 Test Accuracy of WT - Planetary Gearbox Dataset
[0190]
[0191]
[0192] 5.2 Confusion Matrix Analysis
[0193] The accuracy confusion matrix of the WT-planetary gearbox target domain test data is as Figure 7 shown. For the fault diagnosis tasks with 450 target domain test samples under each working condition, all samples are correctly identified under the E+F→H working condition, only 1 sample is mis-identified in both F+G→I and I+J→L, and the numbers of mis-identified samples in the other two working conditions are 17 and 4 respectively. Among them, it is difficult to distinguish between the gear damage samples of class 0 and the tooth surface wear samples of class 4. Class 0 is mis-diagnosed as class 4 7 times, and class 4 is mis-diagnosed as class 0 10 times.
[0194] 5.3 Feature Visualization
[0195] The feature visualization results of CNN com and MSACL on the WT-planetary gearbox dataset are as Figure 8 shown. As Figure 8 (a) In the CNN without domain adaptation com , there is a certain overlap between the target domain samples of the missing tooth fault of class 2 and the samples of the other four fault states, and the target domain samples of each fault state are gathered inside, while the source domain samples are gathered outside, showing an obvious deviation. MSACL shows strong inter-domain fusion and inter-class discrimination in the transfer learning tasks under five working conditions. The target domain samples of each fault class are closely fitted with the source domain samples and are well clustered into a cluster, and there is also less overlap between the samples of different fault classes. In the F+G→I working condition, the missing tooth samples of class 2 are gathered into two clusters, but the two clusters are clearly distinguished from the other clusters, and the samples of the two domains have good fusion, so they are easy to identify.
[0196] 6. Model Ablation Experiment
[0197] To prove the effectiveness of the domain-specific distribution alignment and classifier alignment loss modules, ablation experiments of the model were carried out, and the experimental results under two datasets are shown in Tables 8 and 9. Among them, M1: Under the structure of MSACL, the classifier alignment loss is not considered; M2: Under the structure of MSACL, the LMMD loss is not considered. In all working conditions studied on the PHM2009 dataset and the WT-planetary gearbox dataset, MSACL using LMMD sub-domain alignment and classifier alignment loss achieved the highest diagnostic accuracy, fully proving the effectiveness of the domain alignment strategy and the classifier alignment strategy can effectively reduce the probability of mis-classifying the samples on the decision boundary.
[0198] Table 8 Ablation Experiment of PHM2009 Dataset
[0199]
[0200]
[0201] Table 9 Ablation experiments on the WT-planetary gearbox dataset
[0202]
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution shall be covered by the scope of the claims of the present invention.
Claims
1. An intelligent model-driven cross-condition fault diagnosis method for production equipment, characterized in that Including: S1: Obtain the datasets of the target domain and N source domains; the dataset of the target domain includes data samples of the target domain, and the datasets of the source domains include data samples of the source domains and corresponding true fault labels; S2: Input the data samples of the target domain into the trained fault diagnosis model to output corresponding fault prediction labels; S201: Combine the target domain with N source domains respectively to form N pairs of source domains and the target domain, and set corresponding sub-structure networks for each pair of source domains and the target domain; the sub-structure network includes a private feature extractor, a domain-specific classifier, and a domain-specific distribution alignment module; S202: Randomly extract data samples of the source domain and corresponding true fault labels, as well as data samples of the target domain from N pairs of source domains and the target domain respectively; input the data samples of N pairs of source domains and the target domain into the shared feature extractor for feature extraction to obtain the common features of the data samples of N pairs of source domains and the target domain; S203: Input the common features of the data samples of each pair of source domains and the target domain into the corresponding private feature extractor for feature extraction to obtain the private features of the data samples of this pair of source domains and the target domain; S204: Input the private features of the data samples of each pair of source domains and the target domain into the corresponding domain-specific classifier for classification to obtain the source domain fault prediction labels and target domain fault prediction labels of the data samples of this pair of source domains and the target domain; S205: Use the private features of the data samples of each pair of source domains and the target domain, the target domain fault prediction labels, and the true fault labels of the source domain through the corresponding domain-specific distribution alignment module to calculate the local maximum mean discrepancy loss function; the local maximum mean discrepancy loss function is used to train the shared feature extractor and the corresponding private feature extractor to achieve feature distribution alignment between the source domain and the target domain when extracting features; S206: Calculate the domain-specific distribution alignment loss function through the local maximum mean discrepancy loss functions of N pairs of source domains and the target domain; S207: Calculate the domain-specific classifier loss function through the source domain fault prediction labels and true fault labels of all source domains; S208: Calculate the classifier alignment loss function by minimizing the differences of the target domain fault prediction labels of all domain-specific classifiers for the data samples of the target domain; S209: Calculate the overall optimization loss function through the domain-specific distribution alignment loss function, the domain-specific classifier loss function, and the classifier alignment loss function, and reverse-optimize the parameters of the shared feature extractor in the fault diagnosis model and the private feature extractors and domain-specific classifiers of all sub-structure networks; S210: Repeat steps S202 to S209 to iteratively train the fault diagnosis model until the model converges or reaches the maximum number of iterations; Among them, when inputting the data samples of the target domain into the trained fault diagnosis model, after passing through the shared feature extractor and the private feature extractors and domain-specific classifiers of all sub-structure networks in sequence, take the average of the outputs of the domain-specific classifiers of all sub-structure networks as the final fault prediction label; S3: Use the fault prediction labels output by the fault diagnosis model as the fault diagnosis results of the data samples of the target domain.
2. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 1, wherein: In step S1, after obtaining the datasets of the target domain and N source domains, first perform data augmentation and data normalization on the data samples of the target domain and N source domains, and then execute the subsequent steps.
3. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 1, wherein: In step S202, the shared feature extractor G(·) includes three one-dimensional convolutional modules; The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*64, a stride of 16, and a padding of 24, a BN layer with 32 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2; The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 64 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2; The third one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 128 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
4. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 1, characterized in that: In step S203, the private feature extractor F j (·) includes two layers of one-dimensional convolutional modules; The first one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 256 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2; The second one-dimensional convolutional module includes a one-dimensional convolution with a kernel size of 1*3, a stride of 1, and a padding of 1, a BN layer with 512 channels, a RELU layer, and a max pooling layer with a kernel size of 1*2.
5. The intelligent model-driven cross-condition fault diagnosis method for production equipment according to claim 1, wherein: In step S204, the domain-specific classifier C j (·) includes two fully connected layers connected end to end; The first fully connected layer includes neurons with 512 on both sides, a Dropout layer with a probability of 0.3, and a RELU layer; The second fully connected layer includes neurons with 512 on one side and the number equal to the number of fault categories on the other side.
6. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 1, wherein: In step S205, the formula for calculating the local maximum mean discrepancy loss function by the domain-specific distribution alignment module using the private features of each pair of source domain and target domain data samples, the target domain fault prediction labels, and the true fault labels of the source domain is: In the formula: represents an unbiased estimator for calculating the local maximum mean difference, that is, the local maximum mean difference loss function, P s , P t respectively represent the data sample feature distributions of the source domain and the target domain, N s , N t respectively represent the number of samples in the source domain and the target domain, K represents the number of fault categories, H represents the reproducing kernel Hilbert space, F(·) represents the private features of the data samples in the source domain or the target domain, and respectively represent the data samples in the source domain and the target domain and belong to the weights of the k-th type of fault; for the data samples in the source domain, is the k-th value of the one-hot encoded vector corresponding to the true fault label of the data sample x m , for the unlabeled data samples in the target domain, is the k-th value of the target domain fault prediction label m obtained by the domain-specific classifier for the data sample x ; represents the total probability that all samples in the domain D belong to the category k; k(·,·) represents the Gaussian kernel function, and σ represents the kernel width.
7. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 1, characterized in that: In step S206, the domain-specific distribution alignment loss function L is calculated by the following formula lmmd : In the formula: represents the i-th data sample in the j-th source domain, represents the common feature of the data sample ; represents the private feature of the data sample ; represents the i-th data sample in the target domain, represents the common feature of the data sample ; represents the private feature of the data sample ; represents the calculation of and the local maximum mean discrepancy loss function between them.
8. The intelligent model-driven cross-condition fault diagnosis method for production equipment according to claim 7, wherein: In step S207, the domain-specific classifier loss function L is calculated by the following formula C : Wherein: represents a data sample corresponding true fault label represents the predicted label of the source domain data sample by the j-th domain-specific classifier , J(·,·) represents the calculation of the cross-entropy loss function represents the mathematical expectation of the source domain data sample 9. The cross-condition fault diagnosis method for intelligent model-driven production equipment according to claim 8, wherein: In step S208, the classifier alignment loss L is calculated by the following formula disc : Where: C i (F i (G(x k ))), C j (F j (G(x k ))) represent the predicted probability vectors of the target domain data sample x k for the i-th and j-th domain-specific classifiers respectively, represents the mathematical expectation of the target domain data sample.
10. The intelligent model-driven cross-condition fault diagnosis method for production equipment according to claim 9, characterized in that: In step S209, calculate the overall optimization loss function through the following formula: L total = L C + αL lmmd + βL disc ; Where: L C represents the domain-specific classifier loss function, L lmmd represents the domain-specific distribution alignment loss function, L disc represents the classifier alignment loss, and α, β are set trade-off parameters; Among them, the formula for back-optimizing the parameters of the shared feature extractor and all private feature extractors and domain-specific classifiers in the fault diagnosis model is: where: θ G , θ F , θ C are the parameters to be trained for the shared feature extractor, the private feature extractor, and the domain-specific classifier, respectively, is the partial derivative algorithm, and η is the learning rate.
Citation Information
Cited By
Underwater robot propeller type thruster cross-domain intelligent fault diagnosis method
CN120705742A