Rolling Bearing Fault Diagnosis Method, System, Device, Medium and Program Product

The dual-mode adversarial deep migration learning network addresses the alignment issue in traditional methods by using sparse wavelet convolution and multi-linear mapping to enhance fault diagnosis performance and knowledge transfer in rolling bearing fault recognition.

CN120086803BActive Publication Date: 2025-07-15SHANDONG UNIV

Patent Information

Application Number
CN202510549457.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-15
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Traditional domain adversarial methods ignore category influences in rolling bearing fault diagnosis, resulting in a decrease in the similarity of the characteristic of the source domain and the target domain, and fail to effectively identify faults under load changes or speed changes.

Method used

A dual-mode adversarial deep transfer learning network is adopted, and a sparse wavelet convolution module is used to extract specific fault frequency characteristics, combine multilinear mapping and dual classifier to retain category information, and align the source domain and target domain feature distribution through maximum mean difference.

Benefits of technology

It improves fault diagnosis performance and knowledge transfer capabilities, enhances the sensitivity and interpretability of the model to transient shocks, effectively deals with the uncertainty of the target domain, and improves the accuracy of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086803B_ABST
    Figure CN120086803B_ABST
Patent Text Reader

Abstract

The present invention discloses a rolling bearing fault diagnosis method, system, device, medium and program product, relating to the technical field of rotating machinery fault diagnosis, including training a dual-mode adversarial deep transfer learning network based on source domain features and target domain features; wherein, training a dual-classifier based on source domain features and target domain features to obtain a source domain classification error and a dual-classifier certainty error, training a domain discriminator based on a multilinear mapping composed of source domain features, target domain features and the prediction results of the dual-classifier to obtain a domain classification error, and combining the maximum mean discrepancy error of source domain features and target domain features as a loss function; classifying target domain samples using the dual-mode adversarial deep transfer learning network to obtain a fault diagnosis result. Designing a dual-mode adversarial deep transfer learning network, using multilinear mapping and the certainty difference of the dual-classifier to retain category information, introducing the maximum mean discrepancy, and improving the fault diagnosis performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rotating machinery fault diagnosis, and particularly to a rolling bearing fault diagnosis method, system, device, medium and program product. Background Art

[0002] In rolling bearing fault diagnosis, through adversarial transfer learning, a model trained with a large number of labeled bearing vibration data under a certain working condition can be transferred to a working condition with load change or speed change to achieve accurate identification of bearing faults under the new condition.

[0003] However, the traditional adversarial transfer method has at least the following problems:

[0004] (1) Traditional domain adversarial methods mostly use marginal distributions to align source domain and target domain samples, ignoring the influence of categories.

[0005] (2) Traditional domain adversarial methods do not directly align the features of the source domain and the target domain, resulting in a decrease in the similarity of the features of the source domain and the target domain. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a rolling bearing fault diagnosis method, system, device, medium and program product, designs a dual-mode adversarial deep transfer learning network, extracts features containing specific fault frequencies by using the constructed sparse wavelet convolution module, retains category information by using multi-linear mapping and the deterministic difference of the dual-classifier, and introduces the maximum mean discrepancy to better align the feature distributions of the source domain and the target domain and improve the fault diagnosis performance.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a rolling bearing fault diagnosis method, including:

[0009] Taking the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, extracting features from the target domain sample and the source domain sample to obtain source domain features and target domain features;

[0010] Training the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, training the dual-classifier based on the source domain features and the target domain features to obtain the source domain classification error and the dual-classifier deterministic error, training the domain discriminator based on the multi-linear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combining the maximum mean discrepancy error of the source domain features and the target domain features as the loss function of the training process;

[0011] The trained dual-mode adversarial deep transfer learning network is used to classify the target domain samples to obtain the fault diagnosis results.

[0012] As an alternative implementation, for the target domain samples and the source domain samples, a sparse wavelet convolution module is used to extract features respectively; specifically including: the sparse wavelet convolution module includes a wavelet convolution layer and a sparse constraint module composed of two fully connected layers and two activation functions, and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module; the wavelet convolution layer is initialized with a wavelet convolution kernel with translation parameters and scale parameters, and the initialized wavelet convolution layer and sparse constraint module are used to extract features from the target domain samples and the source domain samples respectively; specifically:

[0013] ; ;

[0014] Among them, u and s are the translation parameter and the scale parameter respectively, t is the index of the convolution kernel, is the selected wavelet basis function, is the input sample, h is the output of the wavelet convolution layer, is the convolution operation;

[0015] In the sparse constraint module, a random vector is transformed into half of the original length through a fully connected layer, then transformed by the ReLU activation function, and then restored to the original length through another fully connected layer, and finally normalized to the range of (0, 1) by the sigmoid activation function to obtain the final sparse vector V , multiply h by the sparse vector V to obtain the output features.

[0016] As an alternative implementation, the spectral kurtosis constraint term is:

[0017] ; ;

[0018] The sparse constraint term is: ;

[0019] Among them, I represents the number of features in represents the i th feature in L represents the length of denote the l th element, denote h the envelope; is the kernel Hilbert space; is the L1 norm.

[0020] As an alternative implementation, based on the source domain classification error , the double-classifier certainty error and the domain classification error the loss function of the training process is obtained as:

[0021] ;

[0022] wherein, are the network parameters of the first classifier, the second classifier and the domain discriminator respectively; are the weights of the double-classifier certainty error and the domain classification error respectively; is the number of source domain samples; is the number of target domain samples; is the cross-entropy loss; is the i th prediction result of the j th source domain sample under the th classifier; i is the actual label of the th source domain sample; is the operator for calculating the double-classifier certainty error; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; is the mathematical expectation calculated from the source domain samples satisfying the source domain distribution ; is the mathematical expectation calculated from the target domain samples satisfying the target domain distribution ; are the i th source domain sample and the j th target domain sample respectively; is the i th feature g and prediction result p of the th source domain sample; j is the g th feature p and prediction result of the

[0023] As an alternative implementation, the input of the domain discriminator is:

[0024] ;

[0025] where g is the source domain feature or target domain feature extracted; p is the prediction result of the dual classifier, obtained by taking the average of the predictions of the two classifiers respectively; and are both multilinear mappings; is the tensor product; ⊙ is the element-wise product, is g and p the dimensions of.

[0026] As an alternative implementation, the maximum mean discrepancy error is used to align the extracted features;

[0027] ;

[0028] where are the network parameters of the feature extractor; are respectively the dual classifier deterministic error , the domain classification error and the maximum mean discrepancy error weights; is the number of source domain samples; is the number of target domain samples; is the operator for calculating the dual classifier deterministic error; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; are respectively the i th source domain sample and the j th target domain sample; is the i th source domain sample's feature g and prediction result p joint variable; is the j th target domain sample's feature g and prediction result p joint variable; is a multilinear mapping; is the domain discriminator;H is a reproducing kernel Hilbert space, and is a feature mapping function mapped to the reproducing kernel Hilbert space; is the feature extracted from the i th source domain sample; is the feature extracted from the i th target domain sample.

[0029] In a second aspect, the present invention provides a rolling bearing fault diagnosis system, including:

[0030] A feature extraction module configured to take the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, and perform feature extraction on the target domain sample and the source domain sample to obtain source domain features and target domain features;

[0031] A training module configured to train the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, the dual-classifier is trained based on the source domain features and the target domain features to obtain a source domain classification error and a dual-classifier certainty error, and the domain discriminator is trained based on a multilinear mapping composed of the source domain features, the target domain features, and the prediction results of the dual-classifier to obtain a domain classification error, and then combined with the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process;

[0032] A classification module configured to classify the target domain sample by using the trained dual-mode adversarial deep transfer learning network to obtain a fault diagnosis result.

[0033] In a third aspect, the present invention provides an electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0034] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in the first aspect is completed.

[0035] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, the method described in the first aspect is implemented.

[0036] Compared with the prior art, the beneficial effects of the present invention are:

[0037] The present invention proposes a rolling bearing fault diagnosis method based on a dual-mode adversarial deep transfer learning network, which uses the constructed sparse wavelet convolution module to extract features containing specific fault frequencies, uses multi-linear mapping and the deterministic difference of dual classifiers to retain class information, and introduces the maximum mean difference to better align the feature distributions of the source domain and the target domain, having good fault diagnosis performance and knowledge transfer ability.

[0038] The sparse wavelet convolution module constructed in the present invention initializes the convolution using a wavelet convolution kernel with trainable translation and scale parameters, and uses spectral kurtosis constraint and sparse constraint to restrict the extracted features, so as to effectively extract the features containing fault features, improve the sensitivity of the model to transient impact extraction, allow the selection of the most suitable wavelet kernel for fault diagnosis, improve the fault diagnosis performance and enhance the interpretability.

[0039] The present invention designs a dual-mode adversarial deep transfer learning network, which introduces the dual-classifier prediction variable as conditional information into the adversarial learning, can better align the feature distributions of the source domain and the target domain, and retain class information at the same time; uses the multi-linear mapping of features and class predictions as the discriminator input to capture the complex interaction relationship between the two.

[0040] The present invention generates a more discriminative feature representation by maximizing the deterministic difference between the two classifiers, enhances the classification and generalization ability, and effectively deals with the uncertainty of the target domain; at the same time, it introduces the maximum mean difference to better help the feature extractor extract domain-invariant features.

[0041] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Brief Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0043] Figure 1 It is the flow chart of the rolling bearing fault diagnosis method provided in Embodiment 1 of the present invention;

[0044] Figure 2 It is the schematic diagram of the rolling bearing fault diagnosis method provided in Embodiment 1 of the present invention. Detailed Embodiment

[0045] The following further describes the present invention in conjunction with the drawings and embodiments.

[0046] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0049] Embodiment 1

[0050] As Figure 1 - Figure 2 shown, this embodiment provides a rolling bearing fault diagnosis method, which specifically includes the following steps:

[0051] Taking the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, feature extraction is performed on the target domain sample and the source domain sample to obtain the source domain feature and the target domain feature;

[0052] Based on the source domain feature and the target domain feature, the constructed dual-mode adversarial deep transfer learning network is trained; among them, based on the source domain feature and the target domain feature, the dual-classifier is trained to obtain the source domain classification error and the dual-classifier certainty error. Based on the multilinear mapping composed of the source domain feature, the target domain feature and the prediction result of the dual-classifier, the domain discriminator is trained to obtain the domain classification error, and then combined with the maximum mean difference error of the source domain feature and the target domain feature as the loss function of the training process;

[0053] The trained dual-mode adversarial deep transfer learning network is used to classify the target domain sample to obtain the fault diagnosis result.

[0054] In this embodiment, the vibration signal of the rolling bearing with unknown fault category obtained is used as the target domain sample, and the historical vibration signal of the rolling bearing in each fault state and its corresponding fault category label are used as the source domain sample. The target domain sample and the source domain sample are input into the dual-mode adversarial deep transfer learning network, and through training this network, the fault diagnosis of the rolling bearing is realized.

[0055] Among them, vibration signals of rolling bearings can be collected by sensors, or vibration signals of rolling bearings uploaded by acquisition terminals can be obtained.

[0056] In this embodiment, a sparse wavelet convolution module is used to extract features from target domain samples and source domain samples, obtaining source domain features and target domain features; the sparse wavelet convolution module consists of two parts, including a wavelet convolution layer containing wavelet convolution kernels, and a sparse constraint module composed of two fully connected layers and two activation functions; and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module. The wavelet convolution layer is initialized with wavelet convolution kernels having trainable translation parameters and scale parameters, and the extracted features are restricted by spectral kurtosis constraints and sparse constraints. Thus, after inputting the target domain samples and source domain samples into the sparse wavelet convolution module, the extracted source domain features and target domain features are obtained. The design of the sparse wavelet convolution module improves the sensitivity of the network to transient impact extraction, and allows selection of the most suitable wavelet kernel for fault diagnosis, improving the model performance and enhancing interpretability.

[0057] Specifically:

[0058] The kernel of the wavelet convolution layer in the sparse wavelet convolution module is initialized with wavelet convolution kernels having trainable translation and scale parameters. The initialization formula of the wavelet convolution kernel is:

[0059] ;

[0060] Among them, u and s respectively represent the translation parameter and the scale parameter, t represents the index of the specific value of the convolution kernel, represents the selected wavelet basis function.

[0061] The operation of the wavelet convolution layer is expressed as:

[0062] ;

[0063] Among them, is the input time domain samples of the source domain and the target domain, h is the output of the wavelet convolution layer, is the convolution operation.

[0064] In this embodiment, three different wavelet convolution kernels, namely Laplace wavelet, Mexh wavelet and Morlet wavelet, are adopted, and the scale parameters on each channel of the convolution kernel are uniformly distributed in the ranges of (0.1, 2), (0.1, 3) and (0.1, 4.5) respectively.

[0065] In this embodiment, spectral kurtosis is used as a typical sparsity metric to better select the frequency band for fault feature extraction. The spectral kurtosis constraint term is expressed as:

[0066] ; ;

[0067] In the formula, I represents the number of features in, represents the i th feature in, L represents the length of, represents the l th element of, represents h the envelope of; is the kernel Hilbert space.

[0068] Although initializing with wavelet convolution kernels can generate more interpretable and representative features, the diversity in the selection of translation parameters, scale parameters, and wavelet basis functions may hinder the improvement of fault diagnosis accuracy. Therefore, a sparse constraint module is introduced to screen the extracted features. First, a random vector is transformed to half of its original length through a fully connected layer, then transformed with the ReLU activation function, followed by restoration to the original length through another fully connected layer, and finally, these values are normalized to the range (0, 1) using the sigmoid activation function to obtain the final sparse vector V , and then h is multiplied by the sparse vector V to obtain the output feature vector .

[0069] To remove the channels that have little impact on the results and retain the relatively critical channels containing fault features, the L1-norm sparse constraint term is introduced: ; is the L1 norm.

[0070] In this embodiment, the sparse wavelet convolution module enriches the physical meaning of the model using sparse wavelet convolution kernels and spectral kurtosis constraint terms, improves the interpretability of the model, and effectively extracts the features containing fault features, thereby improving the performance of fault diagnosis.

[0071] In this embodiment, the dual-mode adversarial deep transfer learning network constructed based on source domain features and target domain features is trained, specifically including:

[0072] First, the dual classifier is basically trained using source domain features to ensure that the first classifier and the second classifier can correctly classify source domain samples. Based on the source domain prediction results , the source domain classification error is determined , specifically:

[0073] ;

[0074] wherein, are the network parameters of the feature extractor, the first classifier, and the second classifier respectively; is the number of source domain samples; is the cross-entropy loss; is the i th prediction result of the j rd classifier for the th source domain sample; i is the actual label of the

[0075] Then, on the premise of ensuring the correct classification results of source domain samples, the dual classifier is trained using target domain features to learn the discriminative decision boundary on target domain samples. Based on the target domain prediction results , the dual classifier certainty error is determined ; meanwhile, the domain discriminator is trained to enhance the ability of the domain discriminator to distinguish whether the extracted features come from source domain samples or target domain samples.

[0076] In summary, the training process of the dual-mode adversarial deep transfer learning network is expressed as:

[0077] ;

[0078] ;

[0079] ;

[0080] wherein, is the domain classification error; are the weights of the dual classifier certainty error and the domain classification error respectively; is the domain discriminator; is the network parameter of the domain discriminator; is the number of target domain samples; is the operator for calculating the dual classifier certainty error; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; are the iThe \(i\)-th source domain sample and the j \(j\)-th target domain sample; \(\mu_s\) is the mathematical expectation calculated for the source domain samples that satisfy the source domain distribution ; \(\mu_t\) is the mathematical expectation calculated for the target domain samples that satisfy the target domain distribution ; \(X_{is}\) is the joint variable of the features i and prediction results g of the \(i\)-th source domain sample p ; \(X_{jt}\) is the joint variable of the features j and prediction results g of the \(j\)-th target domain sample p ;

[0081] \(M\) is a multilinear mapping that takes the extracted source domain features and target domain features, as well as the binary classifier prediction results as input, and uses the domain discriminator \(D\) to output the domain classification prediction results of the multilinear mapping. Based on the domain classification prediction results, the domain classification error is determined. Among them, the binary classifier prediction results are obtained by taking the average of the results predicted by two classifiers for each sample; the domain classification error refers to the error between the prediction result of whether a sample is determined by the domain discriminator to belong to the source domain or the target domain and the true result of whether the sample actually belongs to the source domain or the target domain.

[0082] A \(P\) is a square matrix of size K \(C\times C\) K . \(P_{ij}\) A is the element in the m \(i\)-th row and the n \(j\)-th column of m \(P\), representing the product of the probabilities that the first classifier classifies the sample into the n \(i\)-th category and the second classifier classifies it into the A \(j\)-th category, and

[0083] can effectively evaluate the prediction correlation of the binary classifier on different categories. A To minimize the classifier differences in terms of prediction correlation, the two predictions need to be consistent and relevant, that is, to maximize the A diagonal elements of

[0084] In addition, T \(M\) is a multilinear mapping used to overcome the problem that the multimodal information conveyed in the classifier prediction cannot be fully utilized to match the multimodal distribution of complex domains. The multilinear mapping It can fully capture the multimodal structure behind complex data distributions and is thus used for the o calculation of joint variables, denoted as:

[0085] ;

[0086] where g is the source domain feature or target domain feature extracted; is the tensor product, p is the classifier prediction result.

[0087] At this time , to avoid dimensional explosion, when g and p the dimensions of satisfy the randomization method is applied to the calculation of the joint variable o , denoted as:

[0088] ;

[0089] where ⊙ is the element-wise product; and are random matrices sampled only once and fixed during the training process, and each element follows a symmetric distribution with a single variance, i.e., , , such as Gaussian distribution and uniform distribution.

[0090] Then the joint variable o , i.e., the input of the domain discriminator (multilinear mapping), is denoted as:

[0091] ;

[0092] where, and are both multilinear mappings.

[0093] In this embodiment, the feature extractor G is trained based on the classification results of the classifier and the classification results of the domain discriminator, so that the features it extracts can make the effects of the two classifiers as similar as possible, and at the same time can make the domain discriminator as confused as possible.

[0094] In addition, the above process does not directly align the features extracted by the feature extractor, so the maximum mean difference between the source domain features and the target domain features is introduced to G further align the features extracted by the feature extractor;

[0095] The process is as follows:

[0096] ;

[0097] Among them, H is a reproducing kernel Hilbert space, is a feature mapping function mapped to the reproducing kernel Hilbert space; is the weight of the maximum mean discrepancy; is the i feature extracted from the th source domain sample; i is the feature extracted from the

[0098] By performing adversarial training on the feature extractor G、 domain discriminator D、 and the binary classifier in the dual-mode adversarial deep transfer learning network, domain-invariant features for knowledge transfer are extracted through the feature extractor G ; finally, after training the network until convergence, the trained feature extractor and binary classifier are used to classify the target domain samples to achieve rolling bearing fault diagnosis under different working conditions.

[0099] In this embodiment, experiments are carried out on the bearing fault diagnosis dataset, and the vibration signals of the bearing are collected by an accelerometer installed at the end of the bearing. The dataset includes three types of fault samples - inner race fault, outer race fault, and rolling element fault. One set of vibration signals is collected from each of the bearing end and the fan end. Each type of sample has four fault diameters (0.007 inches, 0.014 inches, 0.021 inches, and 0.028 inches), and is tested under four different loads (0, 1 horsepower, 2 horsepower, and 3 horsepower) and corresponding speeds.

[0100] The proposed sparse wavelet convolution module is applied to the fault feature extraction of vibration signals. Three different wavelet convolution kernels, namely Laplace wavelet, Mexh wavelet, and Morlet wavelet, are used, and the scale parameters on each channel of the convolution kernel are uniformly distributed in the ranges of (0.1, 2), (0.1, 3), and (0.1, 4.5) respectively. For the hyperparameters in the dual-mode adversarial deep transfer learning network, they are determined by the grid search method to be , , .

[0101] The experimental results show that the dual-mode adversarial deep transfer learning network proposed in this embodiment can well extract fault features, utilize multi-linear mapping and the deterministic difference of the binary classifier to retain class information, and introduce the maximum mean discrepancy to help better align the source domain and target domain feature distributions.

[0102] Under the qualitative analysis of feature visualization, it can be seen that the features extracted from samples of the same fault type are clustered, while there are obvious boundaries between the features extracted from samples of different fault types, which proves that the model has good fault diagnosis ability and provides an effective solution for the fault diagnosis of rolling bearings.

[0103] Of course, in other embodiments, the selection of the wavelet convolution kernel and the setting of the gating factor value can be changed according to specific situations and requirements, and are not limited to the above values.

[0104] Embodiment 2

[0105] This embodiment provides a rolling bearing fault diagnosis system, including:

[0106] A feature extraction module, configured to use the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, and perform feature extraction on the target domain sample and the source domain sample to obtain source domain features and target domain features;

[0107] A training module, configured to train the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, the dual-classifier is trained based on the source domain features and the target domain features to obtain the source domain classification error and the dual-classifier certainty error, and the domain discriminator is trained based on the multi-linear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combined with the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process;

[0108] A classification module, configured to classify the target domain samples by using the trained dual-mode adversarial deep transfer learning network to obtain the fault diagnosis result.

[0109] It should be noted here that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules can be executed in a computer system such as a set of computer executable instructions as part of the system.

[0110] In more embodiments, there is also provided:

[0111] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0112] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0113] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0114] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0115] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or completed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0116] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.

[0117] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.

[0118] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0119] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier, so that the device, apparatus, or processor can execute the various processes and operations described above. Examples of carriers include signals, computer-readable media, and so on. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0120] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0121] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A rolling bearing fault diagnosis method, characterized in that Including: Taking the vibration signal to be measured of the rolling bearing obtained as the target domain samples and the historical vibration signals as the source domain samples, feature extraction is performed on the target domain samples and the source domain samples to obtain source domain features and target domain features; Based on the source domain features and the target domain features, the constructed dual-mode adversarial deep transfer learning network is trained; among them, based on the source domain features and the target domain features, the dual-classifier is trained to obtain the source domain classification error and the dual-classifier certainty error, and the domain discriminator is trained based on the multilinear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combined with the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process; The trained dual-mode adversarial deep transfer learning network is used to classify the target domain samples to obtain the fault diagnosis result; For the target domain samples and the source domain samples, a sparse wavelet convolution module is used to perform feature extraction respectively; specifically including: the sparse wavelet convolution module includes a wavelet convolution layer and a sparse constraint module composed of two fully connected layers and two activation functions, and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module; the wavelet convolution layer is initialized with a wavelet convolution kernel with translation parameters and scale parameters, and the initialized wavelet convolution layer and the sparse constraint module are used to perform feature extraction on the target domain samples and the source domain samples respectively; specifically: ; ; Among them, u and s are the translation parameter and the scale parameter respectively, t is the index of the convolution kernel, is the selected wavelet basis function, is the input sample, h is the output of the wavelet convolution layer, is the convolution operation; In the sparse constraint module, a random vector is transformed to half of its original length through a fully connected layer, then transformed by the ReLU activation function, and subsequently restored to the original length through another fully connected layer. Finally, it is normalized to the range (0, 1) by the sigmoid activation function to obtain the final sparse vector. V , multiply h by the sparse vector V to obtain the output feature.

2. The rolling bearing fault diagnosis method according to claim 1, characterized in that, Spectral kurtosis constraint term is as follows: ; ; Sparse constraint term is as follows: ; Among them, I represents the number of features in represents the i th feature in L represents the length of represents the l th element of represents h the envelope of is the kernel Hilbert space; is the L1 norm.

3. A rolling bearing fault diagnosis method according to claim 1, characterized in that, Based on the source domain classification error , the double-classifier certainty error and the domain classification error The loss function of the training process is obtained as follows: ; Among them, are the network parameters of the first classifier, the second classifier, and the domain discriminator respectively ; are the weights of the binary classifier's certainty error and the domain classification error respectively; is the number of source domain samples; is the number of target domain samples; is the cross-entropy loss; is the i th j source domain sample's prediction result under the th i classifier; is the operator for calculating the binary classifier's certainty error; is the multilinear mapping; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; are the i th source domain sample and the j th target domain sample respectively; is the i th source domain sample's feature g and the joint variable of the prediction result p ; is the j th target domain sample's feature g and the joint variable of the prediction result p .

4. The method for diagnosing faults of a rolling bearing according to claim 1, wherein, The input of the domain discriminator is: ; where, g is the source domain feature or target domain feature extracted; p is the prediction result of the double classifier, obtained by taking the average after the two classifiers predict respectively; and are both multi-linear mappings; is the tensor product; ⊙ is the element-wise product, is g and p the dimensions of.

5. The method for diagnosing faults of a rolling bearing according to claim 1, wherein, The maximum mean difference error is used to align the extracted features; ; Among them, are the network parameters of the feature extractor; are the weights of the deterministic error of the binary classifier , domain classification error and maximum mean discrepancy error respectively; is the number of source domain samples; is the number of target domain samples; is the operator for calculating the deterministic error of the binary classifier; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; are the i th source domain sample and the j th target domain sample respectively; is the i th source domain sample's feature g and prediction result p joint variable; is the j th target domain sample's feature g and prediction result p joint variable; is the multilinear mapping; is the domain discriminator; H is the reproducing kernel Hilbert space, is the feature mapping function mapped to the reproducing kernel Hilbert space; is the feature extracted from the i th source domain sample; is the feature extracted from the i th target domain sample.

6. A rolling bearing fault diagnosis system, characterized in that, Including: A feature extraction module configured to take the vibration signal to be measured of the rolling bearing obtained as the target domain samples and the historical vibration signals as the source domain samples, and perform feature extraction on the target domain samples and the source domain samples to obtain source domain features and target domain features; For the target domain samples and the source domain samples, a sparse wavelet convolution module is used to perform feature extraction respectively; specifically including: the sparse wavelet convolution module includes a wavelet convolution layer and a sparse constraint module composed of two fully connected layers and two activation functions, and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module; the wavelet convolution layer is initialized with a wavelet convolution kernel with translation parameters and scale parameters, and the initialized wavelet convolution layer and the sparse constraint module are used to perform feature extraction on the target domain samples and the source domain samples respectively; specifically: ; ; Among them, u and s are the translation parameter and the scale parameter respectively, t is the index of the convolution kernel, is the selected wavelet basis function, is the input sample, h is the output of the wavelet convolution layer, is the convolution operation; In the sparse constraint module, a random vector is transformed to half of its original length through a fully connected layer, then transformed using the ReLU activation function, subsequently restored to its original length through another fully connected layer, and finally normalized to the range (0, 1) using the sigmoid activation function to obtain the final sparse vector. V , multiply h by the sparse vector V to obtain the output feature; A training module configured to train the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; among them, based on the source domain features and the target domain features, the dual-classifier is trained to obtain the source domain classification error and the dual-classifier certainty error, and the domain discriminator is trained based on the multilinear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combined with the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process; A classification module configured to use the trained dual-mode adversarial deep transfer learning network to classify the target domain samples to obtain the fault diagnosis result.

7. An electronic device, characterized in that, Comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the method according to any one of claims 1-5 is completed.

8. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method according to any one of claims 1-5 is completed.

9. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by the processor, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Equipment fault diagnosis method based on semi-supervised small sample

    CN116451150A

  • Bearing fault diagnosis method based on iCORAL-MMD and adversarial transfer learning

    CN116793682A

Cited By

  • High-speed train bearing fault identification method based on maximum mean difference migration learning

    CN122777988A