Rolling bearing depth universal domain self-adaptive cross-working-condition fault diagnosis method and device and medium
By constructing a deep universal domain adaptive cross-condition fault diagnosis model in rolling bearing fault diagnosis, the problem of difficulty in accurately diagnosing rolling bearing faults and identifying private fault types in the prior art is solved, and higher fault diagnosis accuracy is achieved.
Patent Information
- Application Number
- CN202510183202.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-27
Smart Images

Figure CN120217138A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent fault diagnosis of rolling bearings, and particularly to a method, device and medium for deep general domain adaptive cross-condition fault diagnosis of rolling bearings. Background Art
[0002] In the related art, during the service life of rotating machinery, as a key load-bearing component, once a rolling bearing fails, it will lead to production stagnation, increased equipment damage, and even cause safety accidents, resulting in immeasurable economic losses and casualties. Existing intelligent fault diagnosis methods for rotating machinery are difficult to accurately extract valuable fault information from complex mixed signals such as vibration and noise for common non-stationary signals in rotating machinery. Moreover, existing domain adaptive methods for rolling bearing fault diagnosis often require the source task and the target task to have a consistent label space (the same fault type). However, the occurrence of rolling bearing faults is random, and the fault type of the target task is unknown, and there may be private fault types (for example, the known fault types of the source task are inner ring faults and outer ring faults, while the actual fault types of the target task are inner ring faults and rolling element faults). A large number of rolling bearing fault diagnosis methods based on the premise assumption that the source task and the target task have the same fault type are difficult to effectively identify the private classes in the target task, thereby resulting in low accuracy of the fault diagnosis results.
[0003] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a method, device and medium for deep general domain adaptive cross-condition fault diagnosis of rolling bearings, which can improve the accuracy of fault diagnosis results.
[0005] To achieve the above object, on the one hand, an embodiment of the present application proposes a method for deep general domain adaptive cross-condition fault diagnosis of rolling bearings, the method comprising the following steps:
[0006] Obtain the original fault data of the rolling bearing under different operating conditions;
[0007] Divide the original fault data into a source domain data set and a target domain data set according to the operating conditions, the source domain data set consists of several original fault data with fault type labels, and the target domain data set consists of several original fault data without fault type labels;
[0008] Divide the source domain data set into a source domain training set according to a preset sample length and the number of fault samples per class, and divide the target domain data set into a target domain training set and a target domain test set;
[0009] Construct a fault diagnosis model, where the fault diagnosis model includes a feature extractor, a binary classification network, and an adversarial domain discriminator. The binary classification network includes a number of binary classifiers, and the number of binary classifiers is equal to the absolute value of the number of fault types in the source domain training set;
[0010] Train the fault diagnosis model using the source domain training set or the target domain training set, and calculate the total loss function during the training process;
[0011] Adjust the model parameters of the fault diagnosis model according to the total loss function and a preset model parameter optimization algorithm to achieve the alignment of shared class feature distributions between the source domain training set and the target domain training set, and identify the private fault types of the target domain data set according to a confidence threshold;
[0012] When the number of training iterations of the fault diagnosis model is greater than or equal to a preset maximum number of iterations, test the trained fault diagnosis model using the target domain test set, and perform rolling bearing depth general domain adaptive cross-condition fault diagnosis based on the tested fault diagnosis model.
[0013] In some embodiments, the step of training the fault diagnosis model using the source domain training set or the target domain training set and calculating the total loss function during the training process includes:
[0014] Train the feature extractor and the random binary classification network in the fault diagnosis model according to the source domain training set;
[0015] Calculate the one-versus-many loss function in the total loss function during the training process. The calculation formula of the one-versus-many loss function is as follows:
[0016]
[0017] In the formula, represents the one-versus-many loss function; represents the positive class probability that any input sample x in the source domain training set s belongs to the positive class of fault class c; represents the positive class probability that any input sample x in the source domain training set s belongs to the positive class of fault class j.
[0018] In some embodiments, the step of training the fault diagnosis model using the source domain training set or the target domain training set and calculating the total loss function during the training process further includes:
[0019] Calculate the entropy loss function of the binary classifier in the fault diagnosis model. The calculation formula of the entropy loss function is as follows:
[0020]
[0021] In the formula, represents the entropy loss function; represents any input sample x in the target domain training set t the positive class probability belonging to fault class j; |C s | represents the absolute value of the number of fault types in the source domain training set.
[0022] In some embodiments, training the fault diagnosis model with the source domain training set or the target domain training set, and calculating the total loss function during the training process further includes:
[0023] Calculating the discrimination loss function corresponding to the adversarial domain discriminator in the fault diagnosis model, and the calculation formula of the discrimination loss function is as follows:
[0024]
[0025] In the formula, represents the discrimination loss function; and represent the shared weights of the samples in the source domain dataset and the samples in the target domain training set belonging to the shared class; represents the i-th sample in the source domain dataset; represents the j-th sample in the target domain training set; D f represents the calculation process of the feature extractor; D d represents the calculation process of the adversarial domain discriminator; represents calculating the sample of the target domain training set expectation; represents calculating the sample of the source domain training set expectation.
[0026] In some embodiments, training the fault diagnosis model with the source domain training set or the target domain training set, and calculating the total loss function during the training process further includes:
[0027] Calculating the class-level feature alignment loss function in the total loss function, and the calculation formula of the class-level feature alignment loss function is as follows:
[0028]
[0029] In the formula, represents the class-level feature alignment loss function; C s represents the label space of the source domain training set; n s represents the total number of samples in the source domain training set; nt represents the total number of samples in the target domain training set; and respectively represent the weights of the samples in the source domain training set and the samples in the target domain training set that belong to class c; and respectively represent the feature vectors corresponding to the samples in the source domain training set and the samples in the target domain training set <.,.> represents the inner product of vectors; φ(g) represents the feature mapping that maps samples to the kernel Hilbert space;
[0030] Calculate the entropy maximization loss function for the positive class prediction in the source domain training set in the binary classification network. The calculation formula of the entropy maximization loss function is as follows:
[0031]
[0032] In the formula, represents the entropy maximization loss function, y s represents the sample label of the source domain training set; D s represents the source domain training set; represents the sample of the source domain training set, represents the expectation of the sample in the source domain training set;
[0033] Calculate the entropy minimization loss function for the positive class prediction in the target domain training set in the binary classification network. The calculation formula of the entropy minimization loss function is as follows:
[0034]
[0035] In the formula, represents the entropy minimization loss function, y t represents the sample label of the target domain training set; D t represents the target domain training set; represents the sample of the target domain training set, represents calculating the expectation of the sample in the target domain training set.
[0036] In some embodiments, training the fault diagnosis model with the source domain training set or the target domain training set, and calculating the total loss function during the training process further includes:
[0037] Calculate the regularization loss function. The calculation formula of the regularization loss function is as follows:
[0038]
[0039] In the formula, represents the regularization loss function; n t represents the total number of samples in the target domain training set; represents the prediction output by the binary classifier for the sample; represents the average of the predictions of all binary classifiers in the random binary network; represents if takes the value of 1, otherwise 0; represents calculating and the cross-entropy between them and obtaining the average value; λ represents the confidence threshold; represents the average of the positive class predictions of multiple binary classifiers.
[0040] In some embodiments, the calculation formula of the total loss function is as follows:
[0041]
[0042] In the formula, represents the total loss function; (μ1, μ2, μ3, μ4, μ5, μ6) = (0.5, 0.1, 0.01, 0.1, 0.1, 1.0).
[0043] To achieve the above object, on the other hand, an embodiment of the present application proposes a rolling bearing depth general domain adaptive cross-condition fault diagnosis device, and the device includes:
[0044] A first module, configured to obtain the original fault data of the rolling bearing under different operating conditions;
[0045] A second module, configured to divide the original fault data into a source domain data set and a target domain data set according to the operating conditions, the source domain data set is composed of several original fault data with fault type labels, and the target domain data set is composed of several original fault data without fault type labels;
[0046] A third module, configured to divide the source domain data set into a source domain training set according to a preset sample length and the number of samples of each type of fault, and divide the target domain data set into a target domain training set and a target domain test set;
[0047] A fourth module, configured to construct a fault diagnosis model, the fault diagnosis model includes a feature extractor, a binary classification network and an adversarial domain discriminator, the binary classification network includes several binary classifiers, and the number of binary classifiers is equal to the absolute value of the number of fault types in the source domain training set;
[0048] The fifth module is used to train the fault diagnosis model with the source domain training set or the target domain training set, and calculate the total loss function during the training process;
[0049] The sixth module is used to adjust the model parameters of the fault diagnosis model according to the total loss function and a preset model parameter optimization algorithm, so as to achieve the alignment of the shared class feature distributions between the source domain training set and the target domain training set, and identify the private fault types of the target domain data set according to a confidence threshold;
[0050] The seventh module is used to, when the number of training iterations of the fault diagnosis model is greater than or equal to a preset maximum number of iterations, test the trained fault diagnosis model with the target domain test set, and perform rolling bearing depth general domain adaptive cross-condition fault diagnosis based on the tested fault diagnosis model.
[0051] To achieve the above object, another aspect of the embodiments of the present application proposes a computer device, including:
[0052] At least one processor;
[0053] At least one memory for storing at least one program;
[0054] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0055] To achieve the above object, another aspect of the embodiments of the present application proposes a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0056] The embodiments of the present application at least include the following beneficial effects: The present application provides a rolling bearing depth general domain adaptive cross-condition fault diagnosis method, device and medium. This solution trains a fault diagnosis model including a feature extractor, a binary classification network and an adversarial domain discriminator with source domain data containing fault type labels and target domain data without fault type labels, and calculates the loss function during the training process. Then, according to the total loss function and a preset model parameter optimization algorithm, the model parameters of the fault diagnosis model are adjusted to achieve the alignment of the shared class feature distributions between the source domain training set and the target domain training set, and the private fault types of the target domain data set are identified according to a confidence threshold; when the number of training iterations of the fault diagnosis model is greater than or equal to a preset maximum number of iterations, the trained fault diagnosis model is tested with the target domain test set, so that when the trained and tested fault diagnosis model performs rolling bearing fault diagnosis, the accuracy of the fault diagnosis result can be effectively improved. Description of the Drawings
[0057] Figure 1 is the flowchart of the rolling bearing depth general domain adaptive cross - working - condition fault diagnosis method provided by the embodiments of the present application;
[0058] Figure 2 is the schematic diagram of training and optimizing the fault diagnosis model using the source domain and the target domain provided by the embodiments of the present application;
[0059] Figure 3 is the structural schematic diagram of the fault diagnosis model provided by the embodiments of the present application;
[0060] Figure 4 is the schematic diagram of the test results of the rolling bearing depth general domain adaptive cross - working - condition fault diagnosis method provided by the embodiments of the present application;
[0061] Figure 5 is the structural schematic diagram of the rolling bearing depth general domain adaptive cross - working - condition fault diagnosis device provided by the embodiments of the present application;
[0062] Figure 6 is the hardware structural schematic diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application.
[0064] It can be understood that the terms "first", "second", etc. used in the present application can be used to describe various concepts herein, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be called the second information, and similarly, the second information can also be called the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "at the time of...", or "in response to determining".
[0065] The terms "at least one", "multiple", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0067] In the related art, during the service of rotating machinery, the rolling bearing is a key load-bearing component. Once a failure occurs, it will lead to production stagnation, increased equipment damage, and even safety accidents, causing immeasurable economic losses and casualties. Currently, the research hotspots in the field of intelligent fault diagnosis of rotating machinery focus on signal processing technology, fault feature enhancement, and the development of intelligent fault diagnosis models. Signal processing technology is an important basis for fault diagnosis. Traditional methods such as Fourier transform have certain advantages in processing stationary signals, but for the common non-stationary signals in rotating machinery, it is often difficult to accurately extract valuable fault information from complex mixed signals such as vibration and noise.
[0068] Deep learning technology has attracted much attention from scholars for its ability to end-to-end learn the mapping relationship between fault features and fault types and quickly and accurately diagnose and classify rotating machinery faults. It often assumes that the feature distributions of training samples and test samples are the same. However, affected by factors including but not limited to high speed, heavy load, frequent start-stop, and harsh working environments, and the operating state of the rolling bearing will gradually degrade over time, resulting in differences in feature distributions for the same type of fault. Due to the above cross-condition diagnosis problem, deep learning methods will experience significant performance degradation when performing fault diagnosis in the target task. Transfer learning provides a theoretical prior for alleviating the feature distribution differences caused by cross-conditions. Some scholars have integrated transfer learning theory and deep learning models and used domain adaptation methods such as adversarial training and distance metrics to alleviate the problem of cross-domain fault diagnosis of rolling bearings. However, existing domain adaptation methods for rolling bearing fault diagnosis often require the source task and the target task to have a consistent label space (the same fault types). However, the occurrence of rolling bearing faults is random, and the fault types of the target task are unknown, and there may be private fault types (for example, the known fault types in the source task are inner race faults and outer race faults, while the actual fault types in the target task are inner race faults and rolling element faults). A large number of rolling bearing fault diagnosis methods based on the premise assumption that the source task and the target task have the same fault types are difficult to effectively identify private classes in the target task, thereby resulting in low accuracy of fault diagnosis results.
[0069] In view of this, embodiments of this application provide a rolling bearing deep general domain adaptation cross-condition fault diagnosis method, device, and medium, which can improve the accuracy of fault diagnosis results.
[0070] The rolling bearing depth general domain adaptive cross - working condition fault diagnosis method provided by the embodiments of the present application relates to the technical field of intelligent fault diagnosis of rolling bearings. The rolling bearing depth general domain adaptive cross - working condition fault diagnosis method provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the rolling bearing depth general domain adaptive cross - working condition fault diagnosis method, etc., but is not limited to the above forms.
[0071] The present application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet - type devices, multi - processor systems, microprocessor - based systems, set - top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer - executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0072] The following specifically elaborates on the embodiments of the present application with reference to the accompanying drawings:
[0073] Figure 1 is an optional flowchart of the rolling bearing depth general domain adaptive cross - working condition fault diagnosis method provided by the embodiments of the present application, Figure 1 The method in can include but is not limited to steps S110 to S170:
[0074] Step S110, obtain the original fault data of the rolling bearing under different operating conditions;
[0075] Step S120: Divide the original fault data into a source domain dataset and a target domain dataset according to the operating conditions. Among them, the source domain dataset consists of several original fault data with fault type labels, and the target domain dataset consists of several original fault data without fault type labels;
[0076] Step S130: Divide the source domain dataset into a source domain training set according to the preset sample length and the number of samples of each type of fault, and divide the target domain dataset into a target domain training set and a target domain test set;
[0077] Step S140: Build a fault diagnosis model. Among them, the fault diagnosis model includes a feature extractor, a binary classification network, and an adversarial domain discriminator. The binary classification network includes several binary classifiers, and the number of binary classifiers is equal to the absolute value of the number of fault types in the source domain training set;
[0078] Step S150: Train the fault diagnosis model through the source domain training set or the target domain training set, and calculate the total loss function during the training process;
[0079] Step S160: Adjust the model parameters of the fault diagnosis model according to the total loss function and the preset model parameter optimization algorithm to achieve the alignment of the shared class feature distributions between the source domain training set and the target domain training set, and identify the private fault types of the target domain dataset according to the confidence threshold;
[0080] Step S170: When the number of training iterations of the fault diagnosis model is greater than or equal to the preset maximum number of iterations, test the trained fault diagnosis model through the target domain test set, and perform rolling bearing depth general domain adaptive cross-condition fault diagnosis based on the tested fault diagnosis model.
[0081] It can be understood that, as Figure 2 shown, in this embodiment, different operating conditions can be set according to the operating state (load or speed) of the rolling bearing. Then, the original vibration signals under different operating conditions are obtained as the training samples of the model. According to different service conditions, the source domain and the target domain are divided. At the same time, according to the set sample length and step size, the original vibration signals are intercepted to obtain several fault samples. Among them, the fault types of the rolling bearings in the source domain are known, denoted as consisting of n s samples with fault type labels, and the label space is C s ; the fault types in the target domain are unknown, denoted as consisting of n t samples without fault type labels, and the label space is C t . The same fault types (shared classes) in the source domain and the target domain are C = C s I C t, there may also be mutually independent private fault types (private classes). The source domain private class is The target domain private class is Next, several migration fault diagnosis tasks are divided according to different source domains and target domains, and the source domain training set D of the model is set strain , the target domain training set D ttrain and the target domain test sample set D ttest , which are used for model training and testing respectively.
[0082] In the embodiment of the present application, as Figure 3 shown, the fault diagnosis model can be a fine-grained depth cross-domain intelligent fault diagnosis model for the general scenario of rolling bearings. This model includes a feature extractor G f (·), a random binary classification network, and an adversarial domain discriminator G d (·). Among them, the random binary classification network consists of |C s | binary classifiers. Each binary classifier outputs a two-dimensional vector respectively, representing the probabilities that the sample belongs to the positive class (a certain fault class) and the negative class (other fault classes). For each source domain sample, the class corresponding to the given label is called the positive class, and all other classes are negative classes. The sample training process sets mini-batch iterative training, and at the same time, an adaptive learning rate is set during the model training process.
[0083] Based on the above settings, this embodiment constructs a model optimization total loss function using prediction accuracy, weighted adversarial training, shared class local maximum mean discrepancy measurement, and soft consistency regularization Using the source domain training set D strain and the target domain training set D ttrain Train the model, and optimize the model through the Adam algorithm to achieve the alignment of the shared class feature distributions in the source domain and the target domain, and at the same time, achieve the discrimination of the target domain private class by virtue of the confidence threshold.
[0084] Specifically, in order to reduce the computational overhead in this embodiment, the hard negative classifier sampling is adopted in the training process of the random binary classification network. For each source domain fault sample, only two classifiers belonging to the positive class and the hard negative class are trained. By training the corresponding classifiers, the decision boundaries between different classes can be effectively learned. Then calculate and minimize the one-vs-rest loss function to achieve model optimization. Among them, the calculation formula of the one-vs-rest loss function is as follows:
[0085]
[0086] In the formula, represents the one-vs-rest loss function; represents the positive class probability that the arbitrary input sample x s in the source domain training set belongs to the fault class c; represents the arbitrary input sample x in the source domain training sets Positive class probability of fault class j
[0087] In the embodiments of the present application, in order to enhance the low-density separation of the fault sample features in the target domain, ensure the shared class alignment between the target samples and the source samples, and at the same time keep the private class samples in the target domain isolated. During the model training process, the advantages of the binary classifier containing the concept of the "unknown" category and the ability of the multi-class classifier to force the alignment of the target samples with one of the fault classes in the source domain through entropy minimization are respectively utilized. For the target domain sample x t , using open-set entropy minimization, calculate the entropy of all binary classifiers in the random binary network, then calculate the average value and minimize it, and calculate the entropy loss function of the binary classifier to optimize the model. Among them, the calculation formula of the entropy loss function is as follows:
[0088]
[0089] In the formula, represents the entropy loss function; represents any input sample x in the target domain training set t belonging to the positive class probability of fault class j; |C s | represents the absolute value of the number of fault types in the source domain training set.
[0090] In the embodiments of the present application, based on the consistency between the networks and the fact that the classifier will give a higher prediction probability score to samples of the same class, the fault diagnosis model uses the m-sampling random binary network to estimate the consistency confidence score of the target samples and the source samples, and obtains the confidence score of the shared class. For a given sample x, the steps for calculating its confidence score are as follows: First, draw m random binary networks from the multivariate Gaussian distribution; Second, pass the sample through each random binary network to generate m independent |C s |×2-dimensional outputs; Then, calculate the average value of the m independent |C s |×2-dimensional outputs to generate the final output; Finally, take the maximum positive class prediction probability of the output to calculate the confidence score, where the maximum positive class prediction probability of the source domain is The maximum positive class prediction probability of the target domain is Among them, the sample with a higher confidence score is regarded as the positive class. Therefore, the calculation formula for the fault class prediction of the target domain is as follows:
[0091]
[0092] In the formula, "unk" represents the private class in the target domain, represents the sample belonging to the positive class probability of category k, and λ represents the confidence threshold.
[0093] It can be understood that the adversarial domain discriminator G in the fault diagnosis model of this embodiment d (·) aims to adversarially match the feature distributions of source domain samples and target domain samples in the shared class label set C. Since there may be private classes in the target domain and the source domain, the existing adversarial training process will cause the private classes to be misaligned to the shared classes. Therefore, the construction process of the total loss function can adopt sample-level shared weighted adversarial to achieve the alignment of the shared class feature distribution, and then through sample-level shared class weighting, limit the adversarial domain discriminator G d (·) to distinguish source data and target data in the common label set C. Through adversarial training, the feature extractor G f (·) tries to confuse the adversarial domain discriminator G d (·) to obtain cross-domain invariant features of the shared classes in the source domain and the target domain. Specifically, the training process of this embodiment needs to calculate the discriminant loss function corresponding to the adversarial domain discriminator. Among them, the calculation formula of the discriminant loss function is as follows:
[0094]
[0095] In the formula, represents the discriminant loss function; and represent the shared weights of the samples in the source domain dataset and the samples in the target domain training dataset belonging to the shared classes; represents the i-th sample in the source domain dataset; represents the j-th sample in the target domain training dataset; D f represents the calculation process of the feature extractor; D d represents the calculation process of the adversarial domain discriminator; represents calculating the expectation of the sample in the target domain training set; represents calculating the expectation of the sample in the source domain training set.
[0096] It can be understood that if only the adversarial training for global alignment of the source domain and the target domain is considered during the model training process, the category information will be ignored, and there is a risk of misclassifying the target domain samples. To mitigate the risk of global alignment, this embodiment further realizes class-level feature alignment by using category information during the model training process. Among them, the calculation formula of the class-level feature alignment loss function is as follows:
[0097]
[0098] In the formula, and represent the weights of the samples and belonging to the category c, Represents the reproducing kernel Hilbert space mapped by the Gaussian kernel function; φ(x) represents mapping x to the kernel Hilbert space.
[0099] Different from the method of directly using source domain labels and classifier predictions to construct and the local maximum mean discrepancy method, in the method of this embodiment, and the construction process takes into account the influence of private classes in the source domain and the target domain. Among them, the construction formula is as follows:
[0100]
[0101] In the formula, y ic represents the value at the i-th position of the prediction probability vector of the sample belonging to class c, and ω i represents the sample-level shared weight.
[0102] For samples in the source domain, true labels are used and one-hot encoding is performed to calculate the weight of each sample belonging to class c. For the target domain without labels, the maximum positive class prediction probability is used as the probability of assigning target domain samples to each class c, and then the weight of each sample in the target domain belonging to class c is calculated. The larger ω i is, the more likely the sample belongs to the shared class. On the contrary, in the ideal state (ω i = 0), the smaller ω i is, the more likely the sample belongs to the private class. In the ideal state, when the sample belongs to the private class the shared class local maximum mean discrepancy metric method proposed in this embodiment can effectively alleviate the private class risk.
[0103] In order to calculate and optimize the diagnostic model in this embodiment, the above class-level feature alignment loss function is transformed to obtain the following formula:
[0104]
[0105] In the formula, represents the class-level feature alignment loss function; C s represents the label space of the source domain training set; n s represents the total number of samples in the source domain training set; n t represents the total number of samples in the target domain training set; and respectively represent the samples in the source domain training set and the samples in the target domain training set belonging to class c; and respectively represent the output of the feature extractor for the samples in the source domain training set and the feature vectors corresponding to the samples in the target domain training set ; <.,.> represents the inner product of vectors; φ(g) represents the feature mapping that maps the samples to the kernel Hilbert space. <.,.> represents the inner product of vectors; φ(g) represents the feature mapping that maps the samples to the kernel Hilbert space.
[0106] In this embodiment, by minimizing the class-level feature alignment loss function during the training process of the fault diagnosis model, fine-grained information can be utilized to make the distributions of relevant sub-domains within the same category close to each other.
[0107] In addition, in order to force the binary classification network to have a uniform probability distribution on the source domain private classes, this embodiment further proposes to maximize the entropy value of the source domain positive class prediction of the binary classification network, and calculates the entropy value maximization loss function through the following formula:
[0108]
[0109] In the formula, represents the entropy value maximization loss function, y s represents the sample label of the source domain training set; D s represents the source domain training set; represents the sample of the source domain training set, represents the sample in the source domain training set expectation.
[0110] Meanwhile, by minimizing the entropy of the target domain positive class prediction, the prediction of the fault diagnosis model for the target samples belonging to the shared class becomes more reliable. Specifically, this embodiment calculates the entropy value minimization loss function of the positive class prediction in the target domain training set of the binary classification network through the following formula:
[0111]
[0112] In the formula, represents the entropy value minimization loss function, y t represents the sample label of the target domain training set; D t represents the target domain training set; represents the sample of the target domain training set, represents calculating the expectation of the sample in the target domain training set ;
[0113] It can be understood that due to the existence of private classes, relying solely on the prediction of a single classifier may lead to suboptimal diagnostic results; in addition, lacking prior information about the target distribution, it is difficult to enhance the balance of clustering using the hard labels output by the classifier. And consistency regularization can effectively learn a compact representation in semi-supervised learning. Therefore, soft consistency regularization uses a confidence threshold to select target domain samples with high positive class prediction confidence (more likely to be shared class samples) to mitigate the negative impact of private class samples in the target domain. At the same time, a soft consistency regularization framework is introduced. Based on the consistency between networks and the fact that classifiers will give higher prediction probability scores for samples of the same class, the prediction average (soft label) of all binary classifiers in the random binary network is used to assist in the assignment. By considering the overall feature structure of the target domain, structural regularization is enforced, thereby achieving implicit discrimination of the target domain. Specifically, the calculation formula of the regularization loss function involved in this embodiment is as follows:
[0114]
[0115] In the formula, represents the regularization loss function; n t represents the total number of samples in the target domain training set; represents the prediction output by the binary classifier for the sample; represents the prediction average of all binary classifiers in the random binary network; represents if takes the value of 1, otherwise 0; represents calculating and the cross-entropy between and and obtaining the average value; λ represents the confidence threshold; represents the positive class prediction average of multiple binary classifiers.
[0116] After calculating each loss function in this embodiment, the total loss function is calculated through the following formula:
[0117]
[0118] In the formula, represents the total loss function; (μ1, μ2, μ3, μ4, μ5, μ6) = (0.5, 0.1, 0.01, 0.1, 0.1, 1.0).
[0119] Specifically, during the training process of the fault diagnosis model in this embodiment, the model optimization process is achieved by minimizing the total loss function, and the maximum-minimum adversarial training is implemented using the gradient reversal layer.
[0120] After completing a single training of the fault diagnosis model in this embodiment, the target domain test sample set D is used ttestPerform model verification. Specifically, determine whether the number of model training iterations reaches the set maximum number of iterations. If not, continue to execute the above training process. If so, output the trained model and test it using the target domain test sample set, and determine whether the training and testing of the fault diagnosis model are completed with the fault diagnosis accuracy as the evaluation criterion. After completing the training and testing process of the fault diagnosis model, the trained and tested fault diagnosis model can be used for deep general domain adaptive cross-condition fault diagnosis of rolling bearings.
[0121] As can be seen from the above, the method of the embodiment of the present application trains a fault diagnosis model by using the known fault classes in the source domain and the unknown fault class samples in the target domain. At the same time, weight Gaussian distribution modeling is performed using any number of random binary classifiers in the fault diagnosis model. Through the consistency discrimination between classifiers, the same-class confidence scores of source domain samples and target domain fault samples are generated. Based on the confidence threshold, the discrimination of the private classes in the target domain is completed. Further, the total model loss function is constructed by using prediction accuracy, weighted adversarial training, shared class local maximum mean discrepancy metric, and soft consistency regularization, so as to effectively improve the private class discrimination and shared class fault diagnosis accuracy of the fault diagnosis model.
[0122] In some embodiments, the method of the present application is tested using preset publicly available drive-end bearing fault data. Among them, the preset publicly available drive-end bearing fault data has four health states, namely normal bearing (NA), inner ring fault (IF), outer ring fault (OF), and rolling element fault (BF). Each fault includes three fault sizes, namely 0.1778mm, 0.3556mm, and 0.5334mm. The specific classification and labels are shown in Table 1.
[0123] Table 1
[0124]
[0125] The sample length is set to 1024 sampling points, and different cross-condition fault diagnosis tasks are set according to different working conditions. The specific tasks are shown in Table 2:
[0126] Table 2
[0127]
[0128]
[0129] Among them, the inter-domain commonality calculation is |C s I C t | / |C s UC t|. There are 120 training samples for each class in the source domain and the target domain, and 30 test samples for each class in the target domain; the confidence threshold λ = 0.65, and the maximum number of iterations is 60 epochs. The batch size is 8. Each transfer task is trained five times and the average value is taken. The fault diagnosis results of the model are tested using the test set as Figure 4 shown. In the transfer tasks of the commonality and working condition differences between different domains, the method of the embodiment of the present application has achieved high fault diagnosis accuracy, fully demonstrating the superiority and feasibility of the method of the present application.
[0130] Referring to Figure 5 , the embodiment of the present application also provides a rolling bearing depth general domain adaptive cross-working condition fault diagnosis device, and the device includes:
[0131] The first module 510 is used to obtain the original fault data of the rolling bearing under different operating conditions;
[0132] The second module 520 is used to divide the original fault data into a source domain data set and a target domain data set according to the operating condition, wherein the source domain data set consists of several original fault data with fault type labels, and the target domain data set consists of several original fault data without fault type labels;
[0133] The third module 530 is used to divide the source domain data set into a source domain training set according to the preset sample length and the number of fault samples per class, and divide the target domain data set into a target domain training set and a target domain test set;
[0134] The fourth module 540 is used to construct a fault diagnosis model, wherein the fault diagnosis model includes a feature extractor, a binary classification network and an adversarial domain discriminator, the binary classification network includes several binary classifiers, and the number of binary classifiers is equal to the absolute value of the number of fault types in the source domain training set;
[0135] The fifth module 550 is used to train the fault diagnosis model through the source domain training set or the target domain training set, and calculate the total loss function during the training process;
[0136] The sixth module 560 is used to adjust the model parameters of the fault diagnosis model according to the total loss function and the preset model parameter optimization algorithm to achieve the alignment of the shared class feature distributions between the source domain training set and the target domain training set, and identify the private fault types of the target domain data set according to the confidence threshold;
[0137] The seventh module 570 is used to test the trained fault diagnosis model through the target domain test set when the number of training iterations of the fault diagnosis model is greater than or equal to the preset maximum number of iterations, and perform rolling bearing depth general domain adaptive cross-working condition fault diagnosis based on the tested fault diagnosis model.
[0138] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0139] The embodiments of the present application also provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0140] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0141] Please refer to Figure 6 , Figure 6 , which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:
[0142] A processor 610, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0143] A memory 620, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 620 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 620 and are called by the processor 610 to execute the above method of the embodiments of the present application;
[0144] An input / output interface 630, which is used to implement information input and output;
[0145] A communication interface 640, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0146] The bus 650 transmits information between various components of the device (such as the processor 610, the memory 620, the input / output interface 630, and the communication interface 640).
[0147] Among them, the processor 610, the memory 620, the input / output interface 630, and the communication interface 640 realize communication connections with each other inside the device through the bus 650.
[0148] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0149] It can be understood that the content in the above method embodiment is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method embodiment.
[0150] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0152] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0155] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0156] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0157] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0158] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0160] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0161] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A rolling bearing deep universal domain adaptive cross-operating condition fault diagnosis method, characterized in that: The method comprises the following steps: Obtain the original fault data of rolling bearings under different operating conditions; Dividing the original fault data into a source domain data set and a target domain data set according to the operating conditions, the source domain data set is composed of a plurality of original fault data with fault type labels, and the target domain data set is composed of a plurality of original fault data without fault type labels; Dividing the source domain data set into a source domain training set according to a preset sample length and a sample size of each type of fault, and dividing the target domain data set into a target domain training set and a target domain test set; Constructing a fault diagnosis model, wherein the fault diagnosis model includes a feature extractor, a binary classification network and an adversarial domain discriminator, wherein the binary classification network includes a plurality of binary classifiers, and the number of the binary classifiers is equal to the absolute value of the number of fault types in the source domain training set; Training the fault diagnosis model using the source domain training set or the target domain training set, and calculating a total loss function during the training process; Adjusting the model parameters of the fault diagnosis model according to the total loss function and a preset model parameter optimization algorithm to achieve alignment of shared class feature distributions between the source domain training set and the target domain training set, and identifying the private fault type of the target domain data set according to a confidence threshold; When the number of training iterations of the fault diagnosis model is greater than or equal to the preset maximum number of iterations, the trained fault diagnosis model is tested using the target domain test set, and deep universal domain adaptive cross-operating condition fault diagnosis of rolling bearings is performed based on the tested fault diagnosis model.
2. The method according to claim 1, characterized in that The method of training the fault diagnosis model by using the source domain training set or the target domain training set and calculating the total loss function during the training process includes: Training the feature extractor and the random binary classification network in the fault diagnosis model according to the source domain training set; The one-to-many loss function in the total loss function during the training process is calculated. The calculation formula of the one-to-many loss function is as follows: In the formula, represents the one-to-many loss function; Represents any input sample x in the source domain training set s The positive class probability belonging to fault class c; Represents any input sample x in the source domain training set s Positive class probability belonging to fault class j.
3. The method according to claim 2, characterized in that The method of training the fault diagnosis model by using the source domain training set or the target domain training set and calculating the total loss function during the training process further includes: The entropy loss function of the binary classifier in the fault diagnosis model is calculated, and the calculation formula of the entropy loss function is as follows: In the formula, represents the entropy loss function; Represents any input sample x in the target domain training set t The positive class probability belonging to fault class j; |C s | represents the absolute value of the number of fault types in the source domain training set.
4. The method according to claim 3, characterized in that The method of training the fault diagnosis model by using the source domain training set or the target domain training set and calculating the total loss function during the training process further includes: The identification loss function corresponding to the adversarial domain discriminator in the fault diagnosis model is calculated. The calculation formula of the identification loss function is as follows: In the formula, represents the identification loss function; and A sharing weight indicating that the samples in the source domain dataset and the samples in the target domain training set belong to a shared class; represents the i-th sample in the source domain dataset; represents the jth sample in the target domain training set; D f Denotes the calculation process of the feature extractor; D d represents the computation process of the adversarial domain discriminator; Indicates the calculation of samples of the target domain training set expectations; Indicates the calculation of samples of the source domain training set expectations.
5. The method according to claim 4, characterized in that The method of training the fault diagnosis model by using the source domain training set or the target domain training set and calculating the total loss function during the training process further includes: The class-level feature alignment loss function in the total loss function is calculated. The calculation formula of the class-level feature alignment loss function is as follows: In the formula, represents the class-level feature alignment loss function; C s represents the label space of the source domain training set; n s represents the total number of samples in the source domain training set; n t represents the total number of samples in the target domain training set; and Respectively represent the samples in the source domain training set and the samples in the target domain training set The weight belonging to category c; and Respectively represent the samples in the source domain training set output by the feature extractor and the samples in the target domain training set The corresponding eigenvector; .,.> represents the vector inner product; φ(g) represents the feature map that maps the sample to the kernel Hilbert space; The entropy maximization loss function of the positive class prediction in the source domain training set in the binary classification network is calculated. The calculation formula of the entropy maximization loss function is as follows: In the formula, represents the entropy maximization loss function, y s represents the sample label of the source domain training set; D s represents the source domain training set; represents the samples of the source domain training set, Represents the samples in the source domain training set expectations; The entropy minimization loss function of the positive class prediction in the target domain training set in the binary classification network is calculated. The calculation formula of the entropy minimization loss function is as follows: In the formula, represents the entropy minimization loss function, y t represents the sample label of the target domain training set; D t represents the target domain training set; represents the samples of the target domain training set, Indicates the calculation of the samples in the target domain training set expectations.
6. The method according to claim 5, characterized in that The method of training the fault diagnosis model by using the source domain training set or the target domain training set and calculating the total loss function during the training process further includes: The regularized loss function is calculated. The calculation formula of the regularized loss function is as follows: In the formula, represents the regularization loss function; n t represents the total number of samples in the target domain training set; Represents the prediction of the sample output by the binary classifier; represents the average prediction of all binary classifiers in a random binary network; If The value is 1, otherwise it is 0; Representation calculation and The cross entropy between and the average value is obtained; λ represents the confidence threshold; Represents the average positive class prediction of multiple binary classifiers.
7. The method according to claim 6, characterized in that The calculation formula of the total loss function is as follows: In the formula, Represents the total loss function; (μ1, μ2, μ3, μ4, μ5, μ6) = (0.5, 0.1, 0.01, 0.1, 0.1, 1.0).
8. A rolling bearing deep universal domain adaptive cross-operating condition fault diagnosis device, characterized in that: The device comprises: The first module is used to obtain the original fault data of the rolling bearing under different operating conditions; The second module is used to divide the original fault data into a source domain data set and a target domain data set according to the operating conditions, wherein the source domain data set is composed of a plurality of original fault data with fault type labels, and the target domain data set is composed of a plurality of original fault data without fault type labels; The third module is used to divide the source domain data set into a source domain training set and divide the target domain data set into a target domain training set and a target domain test set according to a preset sample length and a sample size of each type of fault; A fourth module is used to construct a fault diagnosis model, wherein the fault diagnosis model includes a feature extractor, a binary classification network and an adversarial domain discriminator, wherein the binary classification network includes a plurality of binary classifiers, and the number of the binary classifiers is equal to the absolute value of the number of fault types in the source domain training set; A fifth module is used to train the fault diagnosis model using the source domain training set or the target domain training set, and calculate a total loss function during the training process; A sixth module is used to adjust the model parameters of the fault diagnosis model according to the total loss function and the preset model parameter optimization algorithm to achieve the shared class feature distribution alignment between the source domain training set and the target domain training set, and identify the private fault type of the target domain data set according to the confidence threshold; The seventh module is used to test the trained fault diagnosis model through the target domain test set when the number of training iterations of the fault diagnosis model is greater than or equal to the preset maximum number of iterations, and to perform deep general domain adaptive cross-working condition fault diagnosis of rolling bearings based on the tested fault diagnosis model.
9. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Rotating machine fault diagnosis method and system based on Gaussian boundary constraint network
CN120890673A
Fault diagnosis method and system for multi-phase permanent magnet motor in cross-working mode
CN121350893A