A method and system for diagnosing faults in rolling element bearings

By constructing a diagnostic network and optimizing the weights of the feature extractor and classifier using dynamic simulation and adaptive spectrum modulation modules, the problems of the validity and recognition accuracy of rolling bearing simulation data were solved, and high-precision fault diagnosis was achieved.

CN120873874BActive Publication Date: 2025-11-28JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511387449.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-11-28
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing technologies struggle to construct effective rolling bearing simulation data, resulting in low fault identification accuracy. Furthermore, traditional domain adaptive methods are limited in performance when faced with unlabeled target domain data or incomplete fault categories.

Method used

A diagnostic network is constructed, which generates labeled simulation data as source domain samples through dynamic simulation and combines it with unlabeled real data from industrial sites. The weights of the feature extractor and classifier are optimized using adaptive spectral modulation and adversarial semantic alignment modules. A total loss function is constructed for training to achieve cross-domain feature extraction and fault diagnosis.

Benefits of technology

It improves the accuracy of rolling bearing fault identification, ensures high classification performance in industrial environments with scarce samples and incomplete categories, reduces the distribution difference between simulation data and measured data, and enhances knowledge transfer capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873874B_ABST
    Figure CN120873874B_ABST
Patent Text Reader

Abstract

The application relates to a rolling bearing fault diagnosis method and system, and belongs to the technical field of bearing fault diagnosis, wherein the method comprises the following steps: constructing a diagnosis network, wherein the diagnosis network comprises a feature extractor and a classifier which are connected in sequence; inputting source domain samples and target domain samples into the feature extractor in parallel to extract source domain features and target domain features; optimizing the weight of the source domain samples through the source domain features and the target domain features, so that the source domain samples with the same class as the target domain samples obtain higher weights; calculating the cross-entropy loss of the source domain based on the weight of the optimized source domain samples, combining the target domain conditional entropy loss to construct a total loss function, training the feature extractor and the classifier through the total loss function to obtain a training-converged diagnosis network; and performing fault diagnosis on unlabeled real samples to be diagnosed through the training-converged diagnosis network. The application can effectively detect the fault types of rolling bearings, and the detection precision is high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bearing fault diagnosis, in particular to a rolling bearing fault diagnosis method and system. BACKGROUND

[0002] Rolling bearings are key mechanical components, and their health conditions are crucial for the safe operation of industrial equipment, so efficient fault diagnosis technology is of great significance. However, in actual engineering, there are serious challenges in obtaining bearing fault samples: on the one hand, the long service life and low failure rate of bearings make it difficult to collect effective data; on the other hand, destructive experiments are costly, and limited samples cannot cover all working conditions and fault types, which seriously affects the generalization ability of the diagnosis model.

[0003] To cope with data scarcity, using digital twin technology for virtual simulation has become a new paradigm, which can flexibly and low-cost generate a large amount of simulated data. However, due to model simplification and environmental noise, etc., there is a "domain difference" between simulation data and real data, which makes the simulation model cannot be directly applied to real fault diagnosis. Although transfer learning is widely used to bridge this gap, traditional domain adaptation methods are limited in performance when facing the "partial domain transfer" problem of unlabeled data in the target domain (real scene) and incomplete fault categories, and may even cause negative transfer.

[0004] Moreover, due to the "domain difference" between simulation data and real data, it seriously affects the generalization ability of existing network models, and if the existing network model cannot effectively extract the potential effective features of the data, it will further reduce the fault recognition accuracy of the rolling bearing.

[0005] Therefore, how to construct effective rolling bearing simulation data and improve the fault recognition accuracy of rolling bearings is a key technical problem to be solved at present. SUMMARY

[0006] Therefore, the technical problem to be solved by the present application is to overcome the problem that it is difficult to construct effective rolling bearing simulation data in the prior art, and the fault recognition accuracy of the rolling bearing is not high.

[0007] To solve the above technical problems, the present application provides a rolling bearing fault diagnosis method, comprising:

[0008] Step S1: constructing a diagnosis network, the diagnosis network comprising a feature extractor and a classifier connected in sequence;

[0009] Step S2: Input the source domain samples and target domain samples in parallel into the feature extractor to extract source domain features and target domain features. Specifically, simulated data of rolling bearings with labels generated by dynamic simulation is used as source domain samples, and real data of rolling bearings without labels and with incomplete categories collected from industrial sites is used as target domain samples.

[0010] Step S3: Optimize the weights of source domain samples using the source domain features and target domain features, so that source domain samples of the same category as target domain samples receive higher weights;

[0011] Step S4: Calculate the cross-entropy loss of the source domain based on the weights of the optimized source domain samples, and construct the total loss function by combining it with the conditional entropy loss of the target domain. Use the total loss function to train the feature extractor and classifier to obtain the diagnostic network that has been trained and converged.

[0012] Step S5: Perform fault diagnosis on the unlabeled real samples to be diagnosed using the diagnostic network that has been trained and converged.

[0013] In one embodiment of the present invention, the feature extractor in step S2 adopts a deep architecture with convolutional modules and adaptive spectrum modulation modules interleaved, including a first convolutional module, a first adaptive spectrum modulation module, a second convolutional module, a second adaptive spectrum modulation module, a third convolutional module, and a fully connected layer connected in sequence.

[0014] In one embodiment of the present invention, step S2, which involves inputting source domain samples and target domain samples in parallel into the feature extractor to extract source domain features and target domain features, includes:

[0015] The first, second, and third convolutional modules in the feature extractor are all used to extract preliminary features from the data, and each includes a convolutional layer, a batch normalization layer, and an activation function ReLU connected in sequence.

[0016] The first and second adaptive spectral modulation modules in the feature extractor convert the initial feature representations from the convolutional module. As input, the features are transformed using Fourier transform. Mapped to the frequency domain, it can be represented as:

[0017] ;

[0018] in, These are the features after Fourier transform. For batch size, For the number of channels, For time domain length, Fourier transform;

[0019] right The amplitude spectrum is obtained by separation. and phase spectrum , denoted as:

[0020] ;

[0021] wherein, is the amplitude spectrum of by taking the modulus, denoted as: ; is the phase spectrum of by taking the phase angle, denoted as: ;

[0022] Local frequency domain feature extraction is performed on the amplitude spectrum by deep separable convolution, denoted as:

[0023] ;

[0024] wherein, is the local frequency domain feature, is the deep separable convolution;

[0025] is input into the feature extraction network, which includes two parallel paths for extracting mean statistical features and peak response features, respectively, and the mean statistical features and peak response features are spliced in the channel dimension to obtain the fusion features, denoted as:

[0026] ;

[0027] ;

[0028] ;

[0029] wherein, is the mean statistical feature, is the global average pooling, is the peak response feature, is the global maximum pooling, denotes the splicing operation in the channel dimension, is the fusion feature;

[0030] Based on the fusion feature , an adaptive frequency domain modulator is generated, denoted as:

[0031] ;

[0032] wherein, is the adaptive frequency domain modulator, is the Sigmoid activation function, which ensures that the modulation weight is normalized to the interval [0, 1]; is a​ dot-product convolutional layers;

[0033] Using the generated adaptive frequency domain modulator Amplitude spectrum The element-wise multiplication operation is represented as:

[0034] ;

[0035] in, The modulated spectrum. This indicates element-wise multiplication.

[0036] modulated spectrum Phase spectrum Recombining them into a complex spectrum, and then transforming the complex spectrum back to the time domain using a fast inverse Fourier transform, it can be represented as:

[0037] ;

[0038] in, For the reconstructed time-domain signal, This is the inverse Fourier transform;

[0039] Will and Performing a residual join is represented as:

[0040] ;

[0041] in, This refers to the characteristics of the source or target domain output by the adaptive spectrum modulation module.

[0042] In one embodiment of the present invention, the method for generating labeled simulation data of rolling bearings as source domain samples through dynamic simulation in step S2 includes:

[0043] A simulation model of the rolling bearing was established by meshing the rolling bearing using finite element analysis and treating the intersections of the meshes as nodes. Based on this model, different fault types of the rolling bearing were simulated, and the nodal acceleration signals of the rolling bearing under different fault types were calculated using explicit dynamic methods. These nodal acceleration signals were then used as source domain samples.

[0044] The formula for calculating the nodal acceleration signal of the rolling bearing under different fault types using explicit dynamics is as follows:

[0045] ;

[0046] in, For the quality matrix, The nodal acceleration is the source domain simulation data. and These represent the external load and internal force caused by internal stress in the rolling bearing, respectively.

[0047] Under the operating conditions of a rolling bearing, the nodal speed update formula is:

[0048] ;

[0049] Under the operating conditions of a rolling bearing, the formula for updating the nodal displacement is:

[0050] ;

[0051] in, For nodal displacement, For nodal displacement The nodal velocity obtained by first differentiation, It is about node speed The first derivative, This is the sequence number of the time step. For time step.

[0052] In one embodiment of the present invention, step S3, which optimizes the weights of source domain samples using the source domain features and target domain features, includes:

[0053] The source domain features and target domain features are input into an adversarial game-based semantic alignment module to optimize the weights of the simulation data in the source domain samples. The adversarial game-based semantic alignment module optimizes the weight vector set ω of the source domain samples based on Wasserstein distance adversarial game principles. The formula for the adversarial game-based semantic alignment module is as follows:

[0054] ;

[0055] Where, n s and n t These represent the total number of samples in the source and target domains, respectively. For discriminator The parameters, and The first feature extracted by the feature extractor The source domain, the first Features of a target domain sample It is the feasible region of the weight vector set ω of the source domain samples. The weight vector set ω of the source domain samples The weights corresponding to each sample;

[0056] For the semantic alignment module formula of the aforementioned resistance game, a discriminator is used under the premise of fixing the weights of the current source domain samples. the parameter update of the discriminator is trained to maximize the difference between the source domain and target domain feature outputs, denoted as:

[0057] ;

[0058] wherein, is a gradient penalty term used to satisfy the Lipschitz constraint of the discriminator ; δ is a weight coefficient of the gradient penalty;

[0059] For the semantic alignment module formula of the resistant game, the optimization solution of the weight vector set ω is solved under the premise of fixing the parameters of the discriminator , and the formula is:

[0060] ;

[0061] wherein, is the score vector of the source domain sample calculated by the discriminator ; n s is the total of the weights of all source domain samples; and ρ is an adjustment parameter.

[0062] In an embodiment of the present application, the method of constructing a total loss function to train the feature extractor and the fault classifier in the step S4 comprises:

[0063] Inputting the source domain data corresponding to the optimized weights into the diagnostic network to construct a cross-entropy loss L about the source domain, denoted as:

[0064] ;

[0065] wherein, denotes a cross-entropy function, denotes a feature extractor, C denotes a classifier, is the total number of source domain samples, is the true label of the source domain sample , is the i-th source domain sample, is the weight corresponding to each sample in the optimized source domain, and are the parameters of the feature extractor and the classifier, respectively, and ω' is the weight vector set of the samples in the optimized source domain;

[0066] Inputting the target domain data into the diagnostic network to construct a conditional entropy loss L ent about the target domain, denoted as:

[0067] ;

[0068] in, Represents the entropy function. Indicates the total number of categories. It is the classifier's analysis of the target domain samples. The predicted probability distribution; The total number of samples in the target domain. This is the j-th target domain sample;

[0069] Based on source domain cross-entropy loss and target domain conditional entropy loss Construct the total loss function The formula is:

[0070] ;

[0071] Where λ is the equilibrium hyperparameter.

[0072] In one embodiment of the present invention, step S4, training the feature extractor and fault classifier using the total loss function, further includes: adjusting the parameters of the feature extractor through backpropagation and employing a stochastic gradient descent algorithm. and classifier parameters Perform iterative optimization, adjusting the parameters of the feature extractor. For the parameters of the classifier All are recorded as Then the iterative optimization formula is:

[0073] ;

[0074] ;

[0075] in, For the first Model parameters at the next iteration For the first Model parameters at the next iteration For momentum, For learning rate, The momentum coefficient, This is the gradient of the loss parameter with respect to the parameter.

[0076] To solve the above-mentioned technical problems, the present invention provides a rolling bearing fault diagnosis system, comprising:

[0077] Network building module: used to build a diagnostic network, which includes a feature extractor and a classifier connected in sequence;

[0078] The feature extraction module is used for inputting source domain samples and target domain samples into the feature extractor in parallel to extract source domain features and target domain features, wherein the labeled simulation data of the rolling bearing generated by dynamic simulation is used as the source domain sample, and the unlabeled and incomplete-class real data of the industrial field rolling bearing is used as the target domain sample;

[0079] The weight optimization module is used for optimizing the weight of the source domain sample through the source domain features and the target domain features, so that the source domain sample with the same class as the target domain sample obtains a higher weight.

[0080] The network training module is used for calculating the cross-entropy loss of the source domain based on the weight of the optimized source domain sample, combining the target domain conditional entropy loss to construct a total loss function, training the feature extractor and the classifier through the total loss function, and obtaining a training converged diagnostic network.

[0081] The fault diagnosis module is used for performing fault diagnosis on the unlabeled real sample to be diagnosed through the training converged diagnostic network.

[0082] To solve the above technical problems, the present application provides an electronic device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to realize the steps of the above rolling bearing fault diagnosis method.

[0083] To solve the above technical problems, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the steps of the above rolling bearing fault diagnosis method.

[0084] The above technical solution of the present application has the following advantages compared with the prior art:

[0085] The present application solves the problem of lack of fault samples and missing labels in industrial field caused by high cost and safety of equipment by constructing a high-fidelity dynamic simulation model and creating a "digital twin" data source that can be effectively generated.

[0086] The feature extractor constructed by the present application embeds a feature extractor of an adaptive spectrum modulation module, which can inject physical priori knowledge of signal processing into a deep network, intelligently enhance the frequency domain features shared across domains by adaptive filtering of data, and suppress irrelevant frequency band interference, which greatly reduces the distribution difference between simulation data and measured data, and lays a solid foundation for efficient knowledge transfer from virtual model to physical entity.

[0087] To address the issue of incomplete matching between the label spaces of the source and target domains, this invention constructs an adversarial semantic alignment module. By introducing adversarial training and weight optimization mechanisms, it dynamically adjusts the weights of the source domain data to align the feature distributions of the source and target domains. This transforms the complex distribution alignment problem into a sophisticated adversarial game process, dynamically quantifying the similarity between each sample in the source domain and the target domain by alternately optimizing the discriminator and the weights of the source domain samples. This mechanism ensures that shared class samples with features similar to the target domain receive higher weights, while source domain-specific class samples with significant differences from the target domain are effectively suppressed, thus accurately mitigating the negative transfer problem. This ensures that the diagnostic model maintains high-precision classification performance even in real-world, non-ideal industrial environments with scarce samples and incomplete categories. Attached Figure Description

[0088] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0089] Figure 1 This is a flowchart of the method of the present invention;

[0090] Figure 2 This is a schematic diagram of the training of the diagnostic network in an embodiment of the present invention. Detailed Implementation

[0091] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0092] Example 1

[0093] Reference Figure 1 As shown, this invention relates to a method for diagnosing rolling bearing faults, comprising:

[0094] Step S1: Construct a diagnostic network, which includes a feature extractor and a classifier connected in sequence;

[0095] Step S2: Input the source domain samples and target domain samples in parallel into the feature extractor to extract source domain features and target domain features. The simulated data of a rolling bearing with labels generated through dynamic simulation is used as the source domain sample (i.e., Figure 1 Labeled simulation data), collecting unlabeled and incomplete real-world data of rolling bearings from industrial sites as target domain samples (i.e. Figure 1 (Unlabeled real data)

[0096] Step S3: Optimize the weights of source domain samples using the source domain features and target domain features, so that source domain samples of the same category as target domain samples receive higher weights;

[0097] Step S4: Calculate the cross-entropy loss of the source domain based on the weights of the optimized source domain samples, and combine the target domain conditional entropy loss to construct a total loss function, train the feature extractor and the classifier through the total loss function, and obtain a training converged diagnostic network;

[0098] Step S5: Perform fault diagnosis on the unlabeled real samples to be diagnosed through the training converged diagnostic network.

[0099] In combination with Figure 1 and Figure 2 , the following will introduce the embodiment in detail:

[0100] In step S1, the classifier of the embodiment adopts a softmax layer, and the feature extractor adopts a deep architecture with convolution modules and adaptive spectral modulation modules (ASMM) interleaved, which includes three convolution modules and two adaptive spectral modulation modules. Specifically, the feature extractor includes a first convolution module, a first adaptive spectral modulation module (i.e. ASMM1 in Figure 2 ), a second convolution module, a second adaptive spectral modulation module (i.e. ASMM2 in Figure 2 ), a third convolution module and a full connection layer (the full connection layer is used to integrate all the features extracted by the network into a fixed length vector as the input of the classifier) connected in sequence. Among them, the adaptive spectral modulation module aims to learn a mask to achieve intelligent signal enhancement. The module first receives the preliminary feature representation from the convolution module as input, and maps the time domain feature signal (i.e. preliminary feature ) to the frequency domain space through Fourier transform (FFT), the process is as follows:

[0101] ;

[0102] Wherein, is the feature after Fourier transform, is the batch size, is the number of channels, is the time domain length, is the Fourier transform.

[0103] Separate to obtain amplitude spectrum and phase spectrum , the process is as follows:

[0104] ;

[0105] Wherein, is the amplitude spectrum Modulo operation, get the amplitude spectrum ; For the Take the phase angle, get the phase spectrum .

[0106] Then, the amplitude spectrum is locally extracted in the frequency domain by a lightweight deep separable convolution, as follows:

[0107] ;

[0108] where, is the local frequency domain feature, is the deep separable convolution.

[0109] This convolution operation can identify domain-shared patterns and domain-specific interference components in different frequency bands, providing spectral sensing capabilities for subsequent adaptive modulation; the obtained feature is input into a parallel deep convolution-based feature extraction network, which contains at least two parallel processing paths for extracting mean statistical features and peak response features, respectively. The frequency domain features extracted by the two paths are spliced in the channel dimension to form a fused feature representation, as follows:

[0110] ;

[0111] ;

[0112] ;

[0113] where, is the mean statistical feature, is the global average pooling, is the peak response feature, is the global maximum pooling, denotes the splicing operation in the channel dimension, is the fused feature.

[0114] Then, based on the fused feature generate an adaptive frequency domain modulator, as follows:

[0115] ;

[0116] where, is the adaptive frequency domain modulator, is the Sigmoid activation function, which ensures that the modulation weight is normalized to the interval [0, 1]; is a dot product convolution layer.

[0117] ​Utilizing generated adaptive frequency domain modulator On the amplitude spectrum Element-wise multiplication is performed, as follows:

[0118] ;

[0119] wherein, is the modulated frequency spectrum, denotes an element-wise multiplication operation.

[0120] The modulated frequency spectrum is recombined with the phase spectrum into a complex frequency spectrum, and the enhanced frequency domain feature is transformed back into the time domain space by an inverse fast Fourier transform, as follows:

[0121] ;

[0122] wherein, is the reconstructed time domain signal, is an inverse Fourier transform.

[0123] Finally, a residual connection mechanism is adopted to ensure information preservation and gradient propagation, and the output feature has enhanced domain invariance, which is beneficial to subsequent cross-domain transfer learning tasks, as follows:

[0124] ;

[0125] wherein, is the feature about the source domain or the target domain output by the adaptive frequency spectrum modulation module.

[0126] Further, assuming that the size of the input data is 1x2048, the parameters of each module in the feature extractor in the embodiment are as follows:

[0127] The convolution kernel size of the convolution layer in the first convolution module is 8x8, the padding is 3x3, and the feature size output by the data after passing through the first convolution module is 32x1024.

[0128] The convolution kernel size of the convolution layer in the second convolution module is 5x5, the padding is 2x2, and the feature size output by the data after passing through the first convolution module is 64x512.

[0129] The convolution kernel size of the convolution layer in the third convolution module is 3x3, the padding is 1x1, and the feature size output by the data after passing through the first convolution module is 128x256.

[0130] In step S2, the embodiment performs grid division on the rolling bearing through finite element analysis, takes the intersection between grids as a node, thereby establishing a simulation model, simulates different fault types about the rolling bearing based on the simulation model, and calculates the node acceleration signal of the rolling bearing under different fault types through an explicit dynamics method, takes the node acceleration signal as a source domain sample, the method does not depend on solving a stiffness matrix, but directly calculates the node acceleration through the difference between a mass matrix and a load, and gradually updates the system state. Wherein, the explicit dynamics method is used to calculate the node acceleration signal of the rolling bearing under different fault types, the formula is:

[0131] ;

[0132] Wherein, is a mass matrix, is a node acceleration, i.e. a source domain sample, and respectively represent an external load of the rolling bearing and an internal stress caused internal force. The explicit dynamics method adopts a central difference method for time integral calculation, the method takes a time step as a unit, gradually advances the system response, and avoids solving a large linear equation set in a traditional implicit method. The specific integral format is as follows:

[0133] In the running state of the rolling bearing, the node velocity update (half step) is:

[0134] ;

[0135] In the running state of the rolling bearing, the node displacement update is:

[0136] ;

[0137] Wherein, is a node displacement, is a node velocity obtained by once differentiating the node displacement, is a node acceleration obtained by once differentiating the node velocity, is a time step number, representing a time point currently calculated; is a time step length, representing a time interval between adjacent two time points. It is not difficult to find that the explicit dynamics method only needs the state information of the current step to complete the next step prediction, which is convenient for parallel calculation and suitable for large-scale nonlinear structure simulation.

[0138]

[0139] ​​Further, the embodiment needs to be normalized before the simulation signal and the actual signal are input into the feature extractor, and the method for signal normalization is as follows:

[0140]

[0141] wherein, is the data normalization result, is a data point in the one-dimensional signal is the mean value of the data and is expressed as:

[0142]

[0143] is the standard deviation of the data and is expressed as:

[0144]

[0145] In step S3, the method for optimizing the weight of the source domain sample through the source domain feature and the target domain feature includes: inputting the source domain feature and the target domain feature into the semantic alignment module based on the adversarial game to optimize the weight of the simulation data in the source domain sample, wherein the core of the semantic alignment module based on the adversarial game is to construct a “min-max” adversarial game based on the Wasserstein distance, and the weight vector set ω of the source domain sample is optimized through a delicate alternating optimization process, and the semantic alignment module based on the adversarial game (i.e., the alternating optimization process) can be expressed as:

[0146]

[0147] wherein, n s and n t are the total number of source domain samples and target domain samples respectively, is the parameter of the discriminator , the discriminator in the embodiment is a multi-layer perception (MLP); and are the features of the n th source domain sample and the n th target domain sample extracted by the feature extractor, is the feasible region of the weight vector set ω of the source domain sample, is the weight corresponding to the n th sample in the weight vector set ω of the source domain sample.

[0148] Further, for the semantic alignment module based on the adversarial game, in the first stage of the alternating optimization, firstly, the parameter of the discriminator is updated under the premise of fixing the current weight of the source domain sample, and the discriminator​​​​​ is trained to maximize the difference between the source domain and the target domain feature outputs, specifically formulated as follows:

[0149] ;

[0150] wherein, is a gradient penalty term to enforce the Lipschitz constraint on the discriminator , and δ is the weight coefficient of the gradient penalty, which is linearly increased during training.

[0151] Further, for the semantic alignment module of the adversarial game, in the second phase of the alternating optimization, i.e., after fixing the parameters of the discriminator , the optimization of the weight vector set ω is solved. Specifically, the samples that are considered more different from the target domain (i.e., with a higher score) by the discriminator will be assigned a lower weight; conversely, the samples that are more similar to the target domain (i.e., with a lower score) will be assigned a higher weight to enhance their contribution in subsequent training. This weight optimization problem is constructed as a linear programming problem with a quadratic constraint, and the objective function and constraint conditions are defined as follows:

[0152] ;

[0153] wherein, is the score vector of the source domain samples calculated by the discriminator , and this problem is a linear programming with a quadratic constraint, which can be efficiently solved using the convex optimization solver CVXPY in the feasible region of the weight. The feasible region is defined by the following constraints: the weight ω i is non-negative; n s is the total of the weights of all source domain samples; and the adjustment parameter ρ controls the degree of deviation of the average weight from the average value 1, and a smaller ρ makes the weight tend to be average, and a larger ρ allows greater weight adjustment.

[0154] It should be noted that the discriminator is required to distinguish as much as possible between source domain samples and target domain samples, and only in this way can it be said that the features extracted by the feature extractor meet the requirements.

[0155] In step S4, the source domain data corresponding to the optimized weights are input into the diagnostic network to construct the cross-entropy loss about the source domain, and the core of the design is to dynamically adjust the contribution of each source domain sample to the parameter update in the back propagation using the weight vector set ω of the source domain samples optimized by the above-mentioned semantic alignment module based on the adversarial game. Specifically, the mathematical expression of the cross-entropy loss is as follows:

[0156] ;

[0157] in, Represents the cross-entropy function. C represents the feature extractor. Represents a classifier. The total number of samples in the source domain. Source domain sample The true label, For the i-th source domain sample, These are the weights corresponding to each sample in the source domain after optimization. and These are the parameters of the feature extractor and classifier, respectively, and ω' is the set of weight vectors of samples in the source domain after optimization.

[0158] In step S4, taking advantage of the unlabeled nature of the target domain data, the target domain data is input into the diagnostic network to construct the conditional entropy loss L for the target domain. ent This strategy aims to improve the predictive certainty of the model in the target domain. By minimizing the information entropy of the prediction probability distribution, it incentivizes the classifier to output higher-confidence predictions for unlabeled target domain samples, thereby enabling the learned decision boundary to better traverse high-density regions of the target domain data. The conditional entropy loss L... ent The specific mathematical expression is:

[0159] ;

[0160] in, Represents the entropy function. Indicates the total number of categories. It is the classifier's analysis of the target domain samples. The predicted probability distribution; The total number of samples in the target domain. Let j be the target domain sample.

[0161] Furthermore, in step S4, to synergistically optimize the model's classification and domain adaptation capabilities, the total loss function in this embodiment integrates the aforementioned supervised classification objective and unsupervised adaptation objective through linear weighting. Specifically, based on the source domain cross-entropy loss... and target domain conditional entropy loss Construct the total loss function Both are balanced by a hyperparameter. Adjustments are made to ensure that while learning knowledge from the source domain, the network can effectively adapt to the distribution characteristics of the target domain. The total loss function of this network model... Defined as:

[0162] ;

[0163] wherein, is a balance hyper-parameter, used to adjust the importance of conditional entropy loss in the total weight.

[0164] Further, the network model parameter optimization of the present application follows the standard end-to-end training paradigm. In each training iteration, the model first performs a forward propagation to calculate the total loss function value L( ) under the current model parameters, and then uses the back propagation algorithm and its automatic differentiation mechanism to efficiently calculate the gradient of the total loss function with respect to all trainable parameters . Finally, an efficient optimizer, specifically the stochastic gradient descent (SGD) algorithm with momentum, is used to update the network parameters according to the calculated gradient. This iteration process continues until the total loss function converges. The parameter update rule is as follows:

[0165] ;

[0166] ;

[0167] wherein the parameters of the feature extractor and the parameters of the classifier are both denoted as , is the model parameter at the th iteration, is the model parameter at the th iteration, is the momentum term, is the learning rate, is the momentum coefficient, is the gradient of the loss parameter with respect to the parameter.

[0168] In step S5, the method for fault diagnosis of the unlabeled real sample to be diagnosed by the training-converged diagnostic network comprises: extracting the features of the sample by the training-converged feature extractor, and then performing fault diagnosis on the extracted sample features by the training-converged classifier.

[0169] Embodiment Two

[0170] The present embodiment provides a rolling bearing fault diagnosis system, comprising:

[0171] a network construction module: for constructing a diagnostic network, the diagnostic network comprising a feature extractor and a classifier connected in sequence;

[0172] The feature extraction module is configured to input source domain samples and target domain samples into the feature extractor in parallel to extract source domain features and target domain features, wherein the labeled simulation data of the rolling bearing generated by the dynamic simulation is used as the source domain samples, and the unlabeled and incomplete-class real data of the rolling bearing collected in the industrial field is used as the target domain samples.

[0173] The weight optimization module is configured to optimize the weights of the source domain samples based on the source domain features and the target domain features, so that the source domain samples with the same class as the target domain samples obtain higher weights.

[0174] The network training module is configured to calculate the cross-entropy loss of the source domain based on the weights of the optimized source domain samples, and construct a total loss function by combining the target domain conditional entropy loss, so as to train the feature extractor and the classifier through the total loss function to obtain a trained diagnostic network.

[0175] The fault diagnosis module is configured to perform fault diagnosis on the unlabeled real samples to be diagnosed through the trained diagnostic network.

[0176] Embodiment three

[0177] The embodiment provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the rolling bearing fault diagnosis method in the embodiment one when executing the computer program.

[0178] Embodiment four

[0179] The embodiment provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the steps of the rolling bearing fault diagnosis method in the embodiment one.

[0180] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, such as object-oriented programming languages Java and interpreted scripting language JavaScript.

[0181] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0183] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks.

[0184] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those of skill in the art once they have the benefit of the present disclosure. Therefore, the appended claims are intended to cover all such variations and modifications as falling within the scope of the application.

[0185] Obviously, the embodiments described above are only examples for clarity and are not intended to limit the scope of the application. Based on the above description, other different forms can also be made by those of ordinary skill in the art. Here, it is not necessary or possible to exhaust all the embodiments. The obvious changes or modifications derived therefrom are still within the protection scope of the present application.

Claims

1. A method of fault diagnosis of a rolling bearing, characterized in that: The method comprises the steps of: Step S1: constructing a diagnostic network comprising a feature extractor and a classifier connected in sequence; Step S2: inputting source domain samples and target domain samples into the feature extractor in parallel to extract source domain features and target domain features, wherein labeled simulation data of rolling bearings generated by dynamic simulation is used as the source domain samples, and unlabeled and incomplete-class real data of industrial rolling bearings is used as the target domain samples; The feature extractor in the step S2 adopts a deep architecture with convolution modules and adaptive frequency modulation modules interleaved, and comprises a first convolution module, a first adaptive frequency modulation module, a second convolution module, a second adaptive frequency modulation module, a third convolution module and a fully connected layer connected in sequence; The method of inputting the source domain samples and the target domain samples into the feature extractor in parallel to extract the source domain features and the target domain features in the step S2 comprises: The first convolution module, the second convolution module and the third convolution module in the feature extractor are all used for extracting preliminary features of data, and all comprise a convolution layer, a batch normalization layer and an activation function ReLU connected in sequence; The first and second adaptive spectral modulation modules in the feature extractor map the preliminary feature representation from the convolutional module As input, the features are mapped to the frequency domain space by a Fourier transform, denoted as: ​ ; wherein, is the Fourier transformed feature, is the batch size, is the number of channels, is the time domain length, is the Fourier transform; To perform the separation, the magnitude spectrum and the phase spectrum are represented as: ; wherein, is the complex spectrum of the signal modulus, obtaining its amplitude spectrum ; is the complex spectrum of the signal phase angle, obtaining its phase spectrum ; Amplitude spectrum through depthwise separable convolution Local frequency domain feature extraction is performed, represented as: ; wherein, is a local frequency domain feature, is a depthwise separable convolution; Will The input is fed into a feature extraction network, which includes two parallel paths for extracting mean statistical features and peak response features, respectively. The mean statistical features and peak response features are then concatenated along the channel dimension to obtain a fused feature, represented as follows: ; ; ; wherein, is a mean statistical feature, is a global average pooling, is a peak response feature, is a global max pooling, denotes a concatenation operation in the channel dimension, is a fusion feature; Fusion-based feature An adaptive frequency domain modulator is generated, denoted as: ; wherein, is an adaptive frequency domain modulator, is a sigmoid activation function that ensures the modulation weights are normalized to the interval [0, 1]; is a is a dot product convolutional layer; Utilizing generated adaptive frequency domain modulator On amplitude spectrum Element-wise multiplication operations, denoted as: ; wherein is the modulated frequency spectrum, denotes an element-wise multiplication operation; modulated spectrum Phase spectrum Recombining them into a complex spectrum, and then transforming the complex spectrum back to the time domain using a fast inverse Fourier transform, it can be represented as: ; wherein is the reconstructed time domain signal, is the inverse Fourier transform; The residual connection is performed with and is expressed as: ; wherein, are features output by the adaptive spectral modulation module with respect to the source domain or the target domain; Step S3: optimizing the weight of the source domain samples based on the source domain features and the target domain features, so that the source domain samples of the same class as the target domain samples obtain higher weights; The method of optimizing the weight of the source domain samples based on the source domain features and the target domain features in the step S3 comprises: inputting the source domain features and the target domain features into a semantic alignment module based on an antagonistic game to optimize the weight of the simulation data in the source domain samples, wherein the semantic alignment module based on the antagonistic game optimizes a weight vector set ω of the source domain samples based on an antagonistic game of a Wasserstein distance, and a formula of the semantic alignment module based on the antagonistic game is: ; wherein n s and n t are the total number of source domain and target domain samples respectively, is the parameter of the discriminator , and are the features of the n th source domain sample and the n th target domain sample extracted by the feature extractor, is the feasible region of the weight vector set ω of the source domain samples, is the weight corresponding to the n th sample in the weight vector set ω of the source domain samples; For the semantic alignment module formula of the resistance game, under the premise of fixing the weight of the current source domain sample, the parameter update of the discriminator is carried out, and the discriminator is trained to maximize the difference between the source domain and the target domain feature output, represented as: ; wherein, is a gradient penalty term used to penalize the discriminator satisfies the Lipschitz constraint; δ is a weight coefficient of the gradient penalty. For the semantic alignment module formula of the resistant game, under the premise of fixing the parameters of the discriminator , the optimization solution of the weight vector set ω is as follows: ; wherein, is a discriminator a computed source domain sample score vector; n s is the sum of all source domain sample weights; p is an adjustment parameter; Step S4: calculating a cross-entropy loss of the source domain based on the weight of the optimized source domain samples, combining a conditional entropy loss of the target domain to construct a total loss function, training the feature extractor and the classifier based on the total loss function, and obtaining a trained diagnostic network; Step S5: performing fault diagnosis on unlabeled real samples to be diagnosed based on the trained diagnostic network.

2. The rolling bearing fault diagnosis method according to claim 1, characterized in that: The method of generating labeled simulation data of rolling bearings by dynamic simulation as the source domain samples in the step S2 comprises: grid division is performed on the rolling bearing by finite element analysis, the intersection points between grids are taken as nodes, thereby establishing a simulation model, different fault types of the rolling bearing are simulated based on the simulation model, and node acceleration signals of the rolling bearing under different fault types are calculated by an explicit dynamic method, and the node acceleration signals are taken as the source domain samples, wherein the formula for calculating the node acceleration signals of the rolling bearing under different fault types by the explicit dynamic method is: ; wherein, is a mass matrix, is a nodal acceleration, i.e. the source domain simulation data, and denote the internal forces caused by the external load and the internal stress of the rolling bearing, respectively; in the running state of the rolling bearing, the node velocity update formula is: ; in the running state of the rolling bearing, the node displacement update formula is: ; wherein, is the node displacement, is the derivative of the node displacement is the node velocity, is the derivative of the node velocity , is the time step number, is the time step length.

3. The rolling bearing fault diagnosis method according to claim 1, characterized in that: The method of constructing the total loss function to train the feature extractor and the fault classifier in the step S4 comprises: Input the source domain data corresponding to the optimized weight into the diagnosis network to build a cross-entropy loss about the source domain is expressed as: ; wherein, represents a cross-entropy function, represents a feature extractor, C represents a classifier, is the total number of source domain samples, is the real label of the source domain sample , is the i-th source domain sample, is the weight corresponding to each sample in the optimized source domain, and are the feature extractor and classifier parameters, respectively, and ω' is the weight vector set of the samples in the optimized source domain; The target domain data is input into the diagnosis network, and a conditional entropy loss L about the target domain is constructed ent , which is represented as: ; wherein, denotes an entropy function, denotes the total number of classes, is the predicted probability distribution of the target domain sample by the classifier; is the total number of target domain samples, is the jth target domain sample; According to the source domain cross-entropy loss and the target domain conditional entropy loss Construct a total loss function , the formula is: ; wherein λ is a balance hyperparameter.

4. The rolling bearing fault diagnosis method according to claim 1, characterized in that: Step S4, which trains the feature extractor and fault classifier using the total loss function, further includes: adjusting the parameters of the feature extractor through backpropagation and employing a stochastic gradient descent algorithm. and classifier parameters Perform iterative optimization, adjusting the parameters of the feature extractor. For the parameters of the classifier All are recorded as Then the iterative optimization formula is: ; ; wherein, is the model parameter at the th iteration, is the model parameter at the th iteration, is the momentum term, is the learning rate, is the momentum coefficient, is the gradient of the loss parameter with respect to the parameter.

5. A rolling bearing fault diagnosis system for implementing the rolling bearing fault diagnosis method according to any one of claims 1 to 4, characterized in that: The method comprises the steps of: a network construction module configured to construct a diagnostic network comprising a feature extractor and a classifier connected in sequence; a feature extraction module configured to input source domain samples and target domain samples to the feature extractor in parallel to extract source domain features and target domain features, wherein labeled simulation data of rolling bearings generated by dynamic simulation is used as the source domain samples, and unlabeled real data of rolling bearings collected in an industrial field is used as the target domain samples; a weight optimization module configured to optimize the weights of the source domain samples by using the source domain features and the target domain features, so that source domain samples of the same class as the target domain samples obtain higher weights; a network training module configured to calculate a cross-entropy loss of the source domain based on the weights of the optimized source domain samples, and construct a total loss function by combining a target domain conditional entropy loss, so as to train the feature extractor and the classifier by using the total loss function to obtain a trained diagnostic network; a fault diagnosis module configured to perform fault diagnosis on unlabeled real samples to be diagnosed by using the trained diagnostic network.

6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The processor executes the computer program to implement the steps of the rolling bearing fault diagnosis method according to any one of claims 1 to 4.

7. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to implement the steps of the rolling bearing fault diagnosis method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Rolling bearing fault diagnosis method based on dynamic index antagonism self-adaption

    CN114429152A

  • Rolling bearing fault diagnosis method based on autoregression data generation

    CN118936887A