Rolling bearing fault diagnosis method, system, equipment, medium and program product
By designing a dual-mode convergence deep transfer learning network in rolling bearing fault diagnosis, using technical means such as sparse wavelet convolution module and multilinear mapping, the problem of insufficient feature alignment in traditional methods is solved, and better fault diagnosis performance and knowledge transfer capabilities are achieved.
Patent Information
- Application Number
- CN202510549457.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-29
AI Technical Summary
In rolling bearing fault diagnosis, traditional adversarial migration methods have problems such as domain adversarial methods ignore category influence and insufficient feature alignment, resulting in poor fault diagnosis performance.
A dual-mode adversarial deep transfer learning network is designed, and the sparse wavelet convolution module is used to extract the characteristics of specific fault frequencies, combined with multilinear mapping and dual classifier deterministic differences, and introduced the maximum mean difference to align the feature distribution of the source domain and the target domain.
The performance and knowledge transfer capabilities of rolling bearing fault diagnosis are improved, and the sensitivity of the model to transient shocks and interpretability of fault diagnosis are enhanced.
Smart Images

Figure CN120086803A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rotary machinery fault diagnosis, and particularly to a rolling bearing fault diagnosis method, system, device, medium and program product. Background Art
[0002] In the fault diagnosis of rolling bearings, through adversarial transfer learning, the model trained by a large number of labeled bearing vibration data under a certain working condition can be transferred to the working conditions with load changes or speed changes to achieve accurate identification of bearing faults under the new conditions.
[0003] However, the traditional adversarial transfer method has at least the following problems: (1) The traditional domain adversarial method mostly uses marginal distributions to align source domain and target domain samples, ignoring the influence of categories.
[0004] (2) The traditional domain adversarial method does not directly align the features of the source domain and the target domain, resulting in a decrease in the similarity of the features of the source domain and the target domain. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a rolling bearing fault diagnosis method, system, device, medium and program product, designs a dual-mode adversarial deep transfer learning network, extracts features containing specific fault frequencies by using the constructed sparse wavelet convolution module, retains category information by using multi-linear mapping and the deterministic difference of the dual-classifier, and introduces the maximum mean discrepancy to better align the feature distributions of the source domain and the target domain and improve the fault diagnosis performance.
[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a rolling bearing fault diagnosis method, including: Taking the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, extracting features from the target domain sample and the source domain sample to obtain source domain features and target domain features; Training the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, training the dual-classifier based on the source domain features and the target domain features to obtain the source domain classification error and the dual-classifier deterministic error, training the domain discriminator based on the multi-linear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combining the maximum mean discrepancy error of the source domain features and the target domain features as the loss function of the training process; Classifying the target domain sample by using the trained dual-mode adversarial deep transfer learning network to obtain the fault diagnosis result.
[0007] As an alternative implementation, for the target domain samples and the source domain samples, a sparse wavelet convolution module is adopted for feature extraction respectively; specifically including: the sparse wavelet convolution module includes a wavelet convolution layer and a sparse constraint module composed of two fully connected layers and two activation functions, and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module; the wavelet convolution layer is initialized with a wavelet convolution kernel having a translation parameter and a scale parameter, and the initialized wavelet convolution layer and the sparse constraint module are used to extract features from the target domain samples and the source domain samples respectively; specifically: ; ; Among them, u and s are the translation parameter and the scale parameter respectively, t is the index of the convolution kernel, is the selected wavelet basis function, is the input sample, h is the output of the wavelet convolution layer, is the convolution operation; In the sparse constraint module, a random vector is transformed into half of the original length through a fully connected layer, then transformed by the ReLU activation function, and then restored to the original length through another fully connected layer, and finally normalized to the range (0, 1) by the sigmoid activation function to obtain the final sparse vector V , multiply h by the sparse vector V to obtain the output feature.
[0008] As an alternative implementation, the spectral kurtosis constraint term is: ; ; The sparse constraint term is: ; Among them, I represents the number of features in represents the i th feature in L represents the length of represents the l th element of represents h the envelope of is the kernel Hilbert space; is the L1 norm.
[0009] As an alternative implementation, based on the source domain classification error , the double classifier certainty error and the domain classification error , the loss function of the training process is obtained as follows: ; wherein, are the network parameters of the first classifier, the second classifier, and the domain discriminator respectively; are the weights of the double classifier certainty error and the domain classification error respectively; is the number of source domain samples; is the number of target domain samples; is the cross-entropy loss; is the prediction result of the i th source domain sample under the j th classifier; is the actual label of the i th source domain sample; is the operator for calculating the double classifier certainty error; is the multilinear mapping; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; are the i th source domain sample and the j th target domain sample respectively; is the i th source domain sample's feature g and the joint variable of the prediction result p ; is the j th target domain sample's feature g and the joint variable of the prediction result p ;
[0010] As an alternative implementation, the input of the domain discriminator is: ; wherein, g is the extracted source domain feature or target domain feature; p is the double classifier prediction result, which is obtained by taking the average after the two classifiers predict respectively; and are both multilinear mappings; is the tensor product; ⊙ is the element-wise product, is g and p the dimensions of.
[0011] As an alternative implementation, the maximum mean discrepancy error is used to align the extracted features; ; where are the network parameters of the feature extractor; are respectively the binary classifier certainty error , the domain classification error and the maximum mean discrepancy error weights; is the number of source domain samples; is the number of target domain samples; is the operator for calculating the binary classifier certainty error; is the prediction result of the first classifier for the i -th target domain sample; is the prediction result of the second classifier for the i -th target domain sample; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; are respectively the i -th source domain sample and the j -th target domain sample; is the i -th source domain sample's feature g and prediction result p joint variable; is the j -th target domain sample's feature g and prediction result p joint variable; is the multilinear mapping; is the domain discriminator; H is the reproducing kernel Hilbert space, is the feature mapping function that maps to the reproducing kernel Hilbert space; is the feature extracted from the i -th source domain sample; is the feature extracted from the i -th target domain sample.
[0012] In a second aspect, the present invention provides a rolling bearing fault diagnosis system, including: A feature extraction module, configured to use the vibration signal to be measured of the rolling bearing obtained as the target domain sample, and the historical vibration signal as the source domain sample, extract features from the target domain sample and the source domain sample, and obtain source domain features and target domain features; A training module, configured to train the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, training the dual-classifier based on the source domain features and the target domain features to obtain a source domain classification error and a dual-classifier certainty error, training the domain discriminator based on a multilinear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain a domain classification error, and then combining the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process; A classification module, configured to classify the target domain sample by using the trained dual-mode adversarial deep transfer learning network to obtain a fault diagnosis result.
[0013] In a third aspect, the present invention provides an electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0015] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a rolling bearing fault diagnosis method based on a dual-mode adversarial deep transfer learning network, uses the constructed sparse wavelet convolution module to extract features containing specific fault frequencies, uses multilinear mapping and dual-classifier certainty differences to retain class information, and introduces the maximum mean difference to better align the feature distributions of the source domain and the target domain, and has good fault diagnosis performance and knowledge transfer ability.
[0017] The sparse wavelet convolution module constructed by the present invention initializes the convolution with a wavelet convolution kernel having trainable translation and scale parameters, and uses spectral kurtosis constraints and sparse constraints to limit the extracted features, so as to effectively extract features containing fault features, improve the sensitivity of the model to transient shock extraction, allow selection of the most suitable wavelet kernel for fault diagnosis, improve fault diagnosis performance and enhance interpretability.
[0018] The present invention designs a dual-mode adversarial deep transfer learning network, which introduces the prediction variables of the dual-classifiers as conditional information into adversarial learning, can better align the feature distributions of the source domain and the target domain, and retain class information at the same time; uses the multi-linear mapping of feature and class prediction as the discriminator input to capture the complex interaction relationship between the two.
[0019] The present invention generates more discriminative feature representations by maximizing the deterministic differences between the two classifiers, enhances the classification and generalization capabilities, and effectively deals with the uncertainty of the target domain; at the same time, introduces the maximum mean difference to better help the feature extractor extract domain-invariant features.
[0020] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative work.
[0022] Figure 1 It is a flowchart of the rolling bearing fault diagnosis method provided in Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the rolling bearing fault diagnosis method provided in Embodiment 1 of the present invention. Detailed Embodiments
[0023] The following will further illustrate the present invention in conjunction with the drawings and embodiments.
[0024] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0025] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] Without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0027] Embodiment 1 As Figure 1 - Figure 2 shown, this embodiment provides a rolling bearing fault diagnosis method, which specifically includes the following steps: Taking the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample, feature extraction is performed on the target domain sample and the source domain sample to obtain source domain features and target domain features; Based on the source domain features and the target domain features, the constructed dual-mode adversarial deep transfer learning network is trained; among them, based on the source domain features and the target domain features, the dual-classifier is trained to obtain the source domain classification error and the dual-classifier certainty error. Based on the multi-linear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier, the domain discriminator is trained to obtain the domain classification error, and then combined with the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process; The trained dual-mode adversarial deep transfer learning network is used to classify the target domain samples to obtain the fault diagnosis result.
[0028] In this embodiment, the vibration signal of the rolling bearing with unknown fault category obtained is used as the target domain sample, and the historical vibration signal of the rolling bearing in each fault state and its corresponding fault category label are used as the source domain sample. The target domain sample and the source domain sample are input into the dual-mode adversarial deep transfer learning network, and through training of the network, the fault diagnosis of the rolling bearing is realized.
[0029] Among them, the vibration signal of the rolling bearing can be collected by a sensor, or the vibration signal of the rolling bearing uploaded by the collection terminal can be obtained.
[0030] In this embodiment, a sparse wavelet convolution module is used to perform feature extraction on the target domain sample and the source domain sample to obtain source domain features and target domain features; the sparse wavelet convolution module consists of two parts, including a wavelet convolution layer containing a wavelet convolution kernel, and a sparse constraint module composed of two fully connected layers and two activation functions; and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module. The wavelet convolution layer is initialized with a wavelet convolution kernel with trainable translation parameters and scale parameters, and the extracted features are restricted by spectral kurtosis constraint and sparse constraint. Thus, after the target domain sample and the source domain sample are input into the sparse wavelet convolution module, the extracted source domain features and target domain features are obtained. The design of the sparse wavelet convolution module improves the sensitivity of the network to transient shock extraction and allows the selection of the most suitable wavelet kernel for fault diagnosis, improving the model performance and enhancing the interpretability.
[0031] Specifically: The kernel of the wavelet convolutional layer in the sparse wavelet convolutional module is initialized by a wavelet convolutional kernel with trainable translation and scale parameters. The initialization formula of the wavelet convolutional kernel is: ; where, u and s represent the translation parameter and the scale parameter respectively, t represents the index of the specific value of the convolutional kernel, represents the selected wavelet basis function.
[0032] The operation of the wavelet convolutional layer is expressed as: ; where, is the time-domain sample of the input source domain and target domain, h is the output of the wavelet convolutional layer, is the convolutional operation.
[0033] In this embodiment, three different wavelet convolutional kernels of Laplace wavelet, Mexh wavelet and Morlet wavelet are adopted, and the scale parameters on each channel of the convolutional kernel are uniformly distributed in the ranges of (0.1, 2), (0.1, 3) and (0.1, 4.5) respectively.
[0034] In this embodiment, spectral kurtosis is used as a typical sparsity measure to better select the frequency band for fault feature extraction. The spectral kurtosis constraint term is expressed as: ; ; In the formula, I represents the number of features in, represents the i th feature in, L represents the length of, represents the l th element of, represents h the envelope of; is the kernel Hilbert space.
[0035] Although initializing with wavelet convolution kernels can generate more interpretable and representative features, the diversification of translation parameters, scale parameters, and the selection of wavelet basis functions may hinder the improvement of fault diagnosis accuracy. Therefore, a sparse constraint module is introduced to screen the extracted features. First, a random vector is transformed into half of its original length through a fully connected layer, then transformed using the ReLU activation function, and then restored to the original length through another fully connected layer. Finally, these values are normalized to the range (0, 1) using the sigmoid activation function to obtain the final sparse vector V , and then h is multiplied by the sparse vector V to obtain the output feature vector .
[0036] To remove the channels that have little impact on the results and retain the relatively critical channels containing fault features, an L1-norm sparse constraint term is introduced: ; where
[0037] is the L1 norm. In this embodiment, the sparse wavelet convolution module enriches the physical meaning of the model using sparse wavelet convolution kernels and spectral kurtosis constraint terms, improves the interpretability of the model, and effectively extracts the features containing fault features, thereby improving the performance of fault diagnosis.
[0038] In this embodiment, the dual-mode adversarial deep transfer learning network constructed based on source domain features and target domain features is trained, specifically including: First, the dual-classifier is basically trained using source domain features to ensure that the first classifier and the second classifier can correctly classify source domain samples. Based on the source domain prediction results , the source domain classification error is determined, specifically as: ; where are the network parameters of the feature extractor, the first classifier, and the second classifier respectively; is the number of source domain samples; is the cross-entropy loss; is the i th source domain sample's prediction result under the j th classifier; is the i th source domain sample's actual label.
[0039] Then, on the premise of ensuring the correct classification results of source domain samples, the dual-classifier is trained using target domain features to learn the discriminative decision boundary on target domain samples. Based on the target domain prediction results , determine the certainty error of the dual-classifier ; Meanwhile, train the domain discriminator to enhance its ability to distinguish whether the extracted features come from source domain samples or target domain samples.
[0040] In summary, the training process of the dual-mode adversarial deep transfer learning network is expressed as: ; ; ; Among them, is the domain classification error; are the weights of the certainty error of the dual-classifier and the domain classification error respectively; is the domain discriminator; are the network parameters of the domain discriminator; is the number of target domain samples; is the operator for calculating the certainty error of the dual-classifier; is the prediction result of the first classifier for the i th target domain sample; is the prediction result of the second classifier for the i th target domain sample; are the i th source domain sample and the j th target domain sample respectively; is the mathematical expectation calculated from the source domain samples that satisfy the source domain distribution ; is the mathematical expectation calculated from the target domain samples that satisfy the target domain distribution ; is the i th source domain sample's feature g and prediction result p 's joint variable; is the j th target domain sample's feature g and prediction result p 's joint variable.
[0041] is a multilinear mapping that takes the extracted source domain features and target domain features, as well as the prediction results of the dual-classifier as inputs, and uses the domain discriminator D to output the domain classification prediction results of the multilinear mapping , and thus determine the domain classification error based on the domain classification prediction results. Among them, the prediction result of the dual-classifier is obtained by taking the average of the results predicted by the two classifiers for each sample; the domain classification error refers to the error between the prediction result of whether a certain sample is determined by the domain discriminator to belong to the source domain or the target domain, and the true result of whether this sample actually belongs to the source domain or the target domain.
[0042] A is a square matrix of size K × K , is A the element in the m -th row and the n -th column of , representing the product of the probabilities that the first classifier classifies the sample into the m -th category and the second classifier classifies it into the n -th category, A which can effectively evaluate the prediction correlation of the binary classifier on different categories.
[0043] To minimize the classifier differences in terms of prediction correlation, the two predictions need to be consistent and relevant, that is, to maximize the A diagonal elements, which also enables the prediction to determine the predicted category with high confidence. At the same time, A the non-diagonal elements of can be regarded as the confusion information of the two classifiers.
[0044] In addition, T is a multilinear mapping, which is used to overcome the problem that the multimodal information conveyed in the classifier prediction cannot be fully utilized to match the multimodal distribution in complex domains. The multilinear mapping can fully capture the multimodal structure behind the complex data distribution, so it is used for the o calculation of the joint variable, denoted as: ; where g is the source domain feature or target domain feature extracted; is the tensor product, p is the classifier prediction result.
[0045] At this time , to avoid the dimension explosion, when g and p the dimensions of satisfy , the randomization method is applied to the calculation of the joint variable o , denoted as: ; where ⊙ is the element-wise product; and are random matrices that are sampled only once and fixed during the training process, and each element follows a symmetric distribution with a single variance, that is , , such as the Gaussian distribution and the uniform distribution.
[0046] Thus, the joint variable o, that is, the input of the domain discriminator (multilinear mapping) is denoted as: ; where and are both multilinear mappings.
[0047] In this embodiment, the feature extractor G is trained based on the classification results of the classifier and the domain discriminator, so that the features it extracts can make the effects of the two classifiers as similar as possible, and at the same time can make the domain discriminator as confused as possible.
[0048] In addition, the above process does not directly align the features extracted by the feature extractor. Therefore, the maximum mean difference between the source domain features and the target domain features is introduced G to further align the features extracted by the feature extractor The process is as follows: ; where H is the reproducing kernel Hilbert space, is the feature mapping function that maps to the reproducing kernel Hilbert space; is the weight of the maximum mean difference; is the feature extracted from the i th source domain sample; is the feature extracted from the i th target domain sample.
[0049] Through the adversarial training of the feature extractor G、 domain discriminator D、 and the dual classifier in the dual-mode adversarial deep transfer learning network, domain-invariant features for knowledge transfer are extracted by the feature extractor G Finally, after the network is trained until convergence, the trained feature extractor and dual classifier are used to classify the target domain samples to achieve rolling bearing fault diagnosis under different working conditions.
[0050] In this embodiment, experiments are carried out on the bearing fault diagnosis dataset, and the vibration signals of the bearing are collected by accelerometers installed at the end of the bearing. This dataset includes three types of fault samples - inner race fault, outer race fault, and rolling element fault. A set of vibration signals are collected from both the bearing end and the fan end respectively. Each type of sample has four fault diameters (0.007 inches, 0.014 inches, 0.021 inches, and 0.028 inches), and is tested under four different loads (0, 1 horsepower, 2 horsepower, and 3 horsepower) and corresponding speeds.
[0051] Apply the proposed sparse wavelet convolution module to the fault feature extraction of vibration signals. Three different wavelet convolution kernels, namely Laplace wavelet, Mexh wavelet and Morlet wavelet, are used. The scale parameters on each channel of the convolution kernel are uniformly distributed in the ranges of (0.1, 2), (0.1, 3) and (0.1, 4.5) respectively. For the hyperparameters in the dual-mode adversarial deep transfer learning network, they are determined by the grid search method as , , .
[0052] The experimental results show that the dual-mode adversarial deep transfer learning network proposed in this embodiment can extract fault features well, retain class information by using multi-linear mapping and the deterministic difference of the dual-classifier, and introduce the maximum mean difference to help better align the feature distributions of the source domain and the target domain.
[0053] It can be seen from the qualitative analysis of feature visualization that the features extracted from samples of the same fault type are clustered, while there are obvious boundaries between the features extracted from samples of different fault types, which can prove that the model has good fault diagnosis ability and provides an effective solution for the fault diagnosis of rolling bearings. Of course, in other embodiments, the selection of the wavelet convolution kernel and the setting of the gating factor value can be changed according to specific situations and requirements, and are not limited to the above values.
[0054] Embodiment 2 This embodiment provides a rolling bearing fault diagnosis system, including: A feature extraction module, configured to use the measured vibration signal of the rolling bearing obtained as the target domain sample and the historical vibration signal as the source domain sample to extract features from the target domain sample and the source domain sample, and obtain the source domain features and the target domain features; A training module, configured to train the constructed dual-mode adversarial deep transfer learning network based on the source domain features and the target domain features; wherein, train the dual-classifier based on the source domain features and the target domain features to obtain the source domain classification error and the dual-classifier deterministic error, train the domain discriminator based on the multi-linear mapping composed of the source domain features, the target domain features and the prediction results of the dual-classifier to obtain the domain classification error, and then combine the maximum mean difference error of the source domain features and the target domain features as the loss function of the training process; A classification module, configured to classify the target domain samples by using the trained dual-mode adversarial deep transfer learning network to obtain the fault diagnosis result.
[0055] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0056] In more embodiments, there is also provided: An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0057] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, field-programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0058] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0059] A computer-readable storage medium for storing computer instructions, which when executed by the processor, completes the method described in Embodiment 1.
[0060] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0061] A computer program product, including a computer program, which when executed by the processor, implements the method described in Embodiment 1.
[0062] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided among program modules as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.
[0063] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0064] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier such that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0065] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0066] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A rolling bearing fault diagnosis method, characterized in that: include: The vibration signal of the rolling bearing to be measured is used as the target domain sample, and the historical vibration signal is used as the source domain sample, and features are extracted from the target domain sample and the source domain sample to obtain source domain features and target domain features; The constructed dual-mode adversarial deep transfer learning network is trained based on the source domain features and the target domain features; wherein, the dual classifier is trained based on the source domain features and the target domain features to obtain the source domain classification error and the dual classifier deterministic error, and the domain discriminator is trained based on the multilinear mapping composed of the source domain features, the target domain features and the prediction results of the dual classifier to obtain the domain classification error, and then the maximum mean difference error of the source domain features and the target domain features is combined as the loss function of the training process; The trained dual-mode adversarial deep transfer learning network is used to classify the target domain samples and obtain the fault diagnosis results.
2. A rolling bearing fault diagnosis method according to claim 1, characterized in that: For target domain samples and source domain samples, sparse wavelet convolution modules are used to extract features respectively; specifically, the sparse wavelet convolution module includes a wavelet convolution layer and a sparse constraint module composed of two fully connected layers and two activation functions, and a spectral kurtosis constraint term is introduced in the wavelet convolution layer, and a sparse constraint term is introduced in the sparse constraint module; the wavelet convolution layer is initialized with a wavelet convolution kernel with a translation parameter and a scale parameter, and the initialized wavelet convolution layer and sparse constraint module are used to extract features for target domain samples and source domain samples respectively; specifically, ; ; in, u and s are the translation parameter and scale parameter respectively, t is the index of the convolution kernel, is the selected wavelet basis function, is the input sample, h is the output of the wavelet convolution layer, is the convolution operation; In the sparse constraint module, a random vector is transformed to half of its original length through a fully connected layer, then converted using the ReLU activation function, and then restored to its original length through another fully connected layer. Finally, the sigmoid activation function is used to normalize it to the range of (0, 1) to obtain the final sparse vector. V ,Will h Multiplying a sparse vector V , and get the output features.
3. A rolling bearing fault diagnosis method as claimed in claim 2, characterized in that: Spectral Kurtosis Constraint for: ; ; Sparse Constraints for: ; in, I express The number of features in express The i Features, L express Length, express No. l elements, express h The envelope of is the nuclear Hilbert space; is the L1 norm.
4. A rolling bearing fault diagnosis method as claimed in claim 1, characterized in that: Based on the source domain classification error , dual classifier deterministic error and domain classification error The loss function of the training process is obtained as: ; in, They are the first classifier, the second classifier and the domain discriminator respectively. Network parameters; are the weights of the dual classifier deterministic error and domain classification error, respectively; is the number of source domain samples; is the number of samples in the target domain; is the cross entropy loss; For the i The source domain samples are j The prediction results under the classifier; For the i The actual labels of source domain samples; is the operator for calculating the deterministic error of the dual classifier; is a multilinear mapping; The first classifier is i The prediction results of target domain samples; The second classifier is i The prediction results of target domain samples; To meet the source domain distribution The mathematical expectation calculated from the source domain samples; To meet the target domain distribution The mathematical expectation calculated from the target domain samples; Respectively i source domain samples and j target domain samples; For the i Features of source domain samples g and prediction results p The joint variable of For the j Features of target domain samples g and prediction results p The combined variable.
5. A rolling bearing fault diagnosis method as claimed in claim 1, characterized in that: The input of the domain discriminator is: ; Among them, g is the extracted source domain feature or target domain feature; p is the prediction result of the dual classifier, which is obtained by taking the average value after the two classifiers predict separately; and They are all multilinear mappings; is the tensor product; ⊙ is the element-wise product, for g and p Dimension.
6. A rolling bearing fault diagnosis method according to claim 1, characterized in that: The maximum mean difference error is used to align the extracted features; ; in, is the network parameter of the feature extractor; are the deterministic errors of the dual classifiers , domain classification error and the maximum mean difference error The weight of is the number of source domain samples; is the number of samples in the target domain; is the operator for calculating the deterministic error of the dual classifier; The first classifier is i The prediction results of target domain samples; The second classifier is i The prediction results of target domain samples; To meet the source domain distribution The mathematical expectation calculated from the source domain samples; To meet the target domain distribution The mathematical expectation calculated from the target domain samples; Respectively i source domain samples and j target domain samples; For the i Features of source domain samples g and prediction results p The joint variable of For the j Features of target domain samples g and prediction results p The joint variable of is a multilinear mapping; is the domain discriminator; H is the reproducing kernel Hilbert space, is the characteristic mapping function mapped to the reproducing kernel Hilbert space; For the i Features extracted from source domain samples; For the i The features extracted from the target domain samples.
7. A rolling bearing fault diagnosis system, characterized in that: include: The feature extraction module is configured to use the vibration signal of the rolling bearing to be measured as the target domain sample and the historical vibration signal as the source domain sample, perform feature extraction on the target domain sample and the source domain sample, and obtain the source domain feature and the target domain feature; A training module is configured to train a dual-mode adversarial deep transfer learning network constructed based on source domain features and target domain features; wherein the dual classifier is trained based on the source domain features and the target domain features to obtain a source domain classification error and a dual classifier deterministic error, and the domain discriminator is trained based on a multilinear mapping composed of the source domain features, the target domain features and the prediction results of the dual classifier to obtain a domain classification error, and then the maximum mean difference error of the source domain features and the target domain features is combined as a loss function of the training process; The classification module is configured to classify the target domain samples using the trained dual-mode adversarial deep transfer learning network to obtain fault diagnosis results.
8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.
9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Equipment fault diagnosis method based on semi-supervised small sample
CN116451150A
Bearing fault diagnosis method based on iCORAL-MMD and adversarial transfer learning
CN116793682A
Gear case fault diagnosis method based on deep transfer learning
CN116894187A
Equipment fault diagnosis method, machine readable storage medium and processor
CN118227957A
Vacuum dry pump bearing fault diagnosis method based on domain confrontation and attention transfer learning
CN119322967A
Cited By
Bearing fault diagnosis method based on cross-domain adaptive weighting
CN120524302A
Rotating machine fault diagnosis method and system based on Gaussian boundary constraint network
CN120890673A
Rotating machinery fault diagnosis method and system based on gaussian boundary constraint network
CN120890673B