Physical knowledge driven transfer learning method for bearing fault classification across operating conditions

By employing a physics-driven transfer learning approach, combined with PDE multi-scale convolution and Transformer, we achieved accurate identification and stable classification of bearing faults in complex industrial environments. This solved the problem of insufficient feature extraction in traditional methods across different operating conditions, and improved the accuracy and robustness of fault identification.

CN121598176BActive Publication Date: 2026-04-10NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING TECH UNIV
Filing Date
2026-01-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional bearing fault classification methods struggle to maintain stability and generalization ability in complex industrial environments, fail to effectively extract key feature information across operating conditions, and lack robust anti-interference capabilities in high-noise environments.

Method used

Employing a physics-driven transfer learning approach, this method introduces physical quantities such as Gaussian smoothing, first-order gradient, and second-order curvature to design a PDE multi-scale convolution and a physics-driven Transformer. Combined with a physics-adaptive MMD domain alignment transfer mechanism, it adaptively captures the real physical structure in vibration signals and achieves physical consistency alignment between the source and target domains under various operating conditions.

Benefits of technology

It significantly improves the accuracy and robustness of fault identification in complex real-world scenarios, enhances the accuracy and stability of fault classification, and strengthens the ability to adapt to different operating conditions and resist noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598176B_ABST
    Figure CN121598176B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of physical knowledge driven migration learning cross-working condition bearing fault classification method, including by gathering bearing drive end vibration acceleration signal and constructing source domain and target domain data set;Cross-working condition bearing fault classification model is constructed, based on Gaussian kernel, Laplace kernel, four different directions Sobel kernel is carried out depth separable convolution, extract physical quantity and construct physical enhancement multi-head self-attention mechanism, traditional feedforward network is optimized in conjunction with gate mechanism, and the classifier based on multiple expert path and physical gate is designed;Combined with Gaussian kernel MMD, migration learning is carried out;The bearing fault classification result is outputted to the measured bearing drive end vibration acceleration signal using the trained working condition bearing fault classification model, and the bearing fault classification under different working conditions is realized.The present application effectively inhibits the difference of working condition and noise disturbance, and significantly improves bearing fault identification precision and robustness in complex actual scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bearing fault classification, and particularly relates to a physical knowledge driven transfer learning cross-condition bearing fault classification method. BACKGROUND

[0002] Under the background of the rapid development of current industrialization and intelligent manufacturing, bearings, as key components in rotating machinery, are directly related to the stability and safety of equipment operation. Bearings are widely used in transportation, aerospace, industrial manufacturing and automation production fields. If the bearing fault cannot be classified in time, it will lead to equipment downtime and performance degradation. Therefore, it is of great significance to realize efficient and reliable bearing fault classification for ensuring the safe operation of industrial equipment. However, in the actual industrial environment, the working conditions change frequently, and the fluctuations of load, speed and external interference lead to strong non-stationarity and significant cross-condition differences of vibration signals, which makes it difficult for traditional methods to maintain classification stability and generalization ability. In order to solve this problem, it is necessary to build an intelligent classification model that can migrate across conditions and has physical interpretation. The cross-condition bearing fault transfer learning method based on physical knowledge driven proposed in the present application is designed to meet the reliable classification requirements under complex working condition conditions. By introducing the physical mechanism of bearing vibration, the migration ability and classification accuracy of the model are improved.

[0003] Traditional bearing fault classification methods mainly include signal processing and machine learning. Signal processing methods rely on high-quality vibration signals, and in complex industrial environments such as noise interference and multiple fault coupling, key features are easily hidden, and classification performance is significantly limited. Machine learning methods classify through artificial features, but shallow models are difficult to extract deep fault information, and have insufficient adaptability to non-stationary and cross-condition changes. Deep learning can automatically extract features from raw signals and improve fault recognition ability, but it has poor anti-interference ability in strong noise environment, and the model is sensitive to changes in different loads, speeds and other working conditions, and the feature distribution is obviously shifted, so the cross-condition classification stability is poor. SUMMARY

[0004] In view of the deficiencies of the prior art, the physical knowledge driven transfer learning cross-condition bearing fault classification method is provided, which solves the problems that the traditional method is difficult to maintain fault classification stability and generalization ability, cannot extract complete key feature information in a complex industrial environment, has insufficient cross-condition adaptability, and has limited noise resistance, etc.The present application introduces Gaussian smoothing, first-order gradient and second-order curvature and other physical quantities in the feature modeling process, designs PDE multi-scale convolution and physically driven Transformer, and adaptively captures the real physical structure in the vibration signal.Meanwhile, combined with the physical adaptive MMD domain alignment transfer mechanism, the physical consistency alignment of the source domain and the target domain under the condition of multiple working conditions is realized, the difference between the working conditions and the noise disturbance are effectively suppressed, and the fault recognition precision and robustness in complex actual scenes are significantly improved.

[0005] To achieve the above technical purpose, the present application provides the following technical scheme: a physical knowledge driven transfer learning cross-condition bearing fault classification method, comprising the following steps:

[0006] The acceleration vibration sensors arranged at multiple points are used to collect the driving end vibration acceleration signals of the bearing under different working condition conditions, and the original vibration sequences under the multiple working condition conditions are constructed in time sequence combined with the time index;

[0007] From the constructed original vibration sequences under the multiple working condition conditions, an original vibration sequence under an arbitrary working condition condition is selected as a source domain signal, and an original vibration sequence under a working condition condition is selected from the remaining original vibration sequences under the multiple working condition conditions which are not selected as a target domain signal, to construct a source domain training set, a target domain validation set and a target domain test set in an unsupervised form; the source domain training set contains fault category labels, and the target domain validation set and the target domain test set do not contain fault category labels;

[0008] A cross-condition bearing fault classification model is constructed, and the cross-condition bearing fault classification model comprises:

[0009] A backbone network converts the input signal into a learnable tensor;

[0010] A PDE multi-scale convolution module performs depth separable convolution on the learnable tensor based on Gaussian kernel, Laplacian kernel and four different direction Sobel kernels, constructs a PDE regularization term for total loss calculation, and generates physical gate fusion features combined with a physical gating mechanism;

[0011] A physically driven Transformer extracts physical quantities from the physical gate fusion features and designs a physical enhancement multi-head self-attention mechanism and a physical gate feedforward network to obtain final output features;

[0012] A classifier converts the final output features into a predicted fault category.

[0013] The total loss is constructed to train the cross-condition bearing fault classification model, and the trained cross-condition bearing fault classification model is taken as an optimal cross-condition bearing fault classification model;

[0014] The bearing driving end vibration acceleration signals under different working conditions are collected by using the multi-point arranged acceleration vibration sensors, and the original vibration sequence under the multi-condition is constructed in time sequence in combination with the time index, including:

[0015] Optionally, the driving end vibration acceleration signals of the bearing under different working conditions are collected by using the multi-point arranged acceleration vibration sensors, and the original vibration sequence under the multi-condition is constructed in time sequence in combination with the time index, including:

[0016] A plurality of measuring points are selected on the bearing seat surface, and the acceleration vibration sensors are fixedly installed;

[0017] The bearing is operated under different working conditions, the acceleration vibration sensors collect the driving end vibration acceleration signals of the normal state and each fault state under each working condition respectively, the sampling frequency is set as a fixed value, the collected driving end vibration acceleration signals are taken as the original vibration data, and the original vibration sequence is constructed in time sequence.

[0018] Optionally, the source domain training set, the target domain verification set and the target domain test set are constructed in an unsupervised form, including:

[0019] In a sliding window clipping manner, each driving end vibration acceleration signal in the original vibration sequence from the source domain signal and the target domain signal is clipped into a plurality of signal sample units with a uniform length, to construct a source domain original sample library and a target domain original sample library;

[0020] According to the time index in the original vibration sequence, the source domain samples and the target domain samples are sampled from the source domain original sample library and the target domain original sample library respectively;

[0021] According to the fault type, the source domain samples are assigned with fault category labels, and one-hot encoding is performed to obtain the labeled source domain samples, to construct a source domain training set;

[0022] The mean and the standard deviation are calculated based on the source domain samples, and the source domain samples and the target domain samples are uniformly standardized according to the mean and the standard deviation;

[0023] The indexes of the standardized target domain samples are randomly shuffled, and are divided into a target domain verification set and a target domain test set at a ratio of 1:1.

[0024] Optionally, the sampling source domain samples and target domain samples from the source domain original sample library and the target domain original sample library according to the time indexes in the original vibration sequence comprises:

[0025] Setting a time section threshold For the source domain original sample library and the target domain original sample library constructed based on the original vibration sequence with a total length of : randomly selecting signal sample units in a part of the time indexes in the target domain original sample library as the source domain samples, wherein is a sample number control parameter; and randomly selecting signal sample units in a part of the time indexes in the target domain original sample library as the target domain samples.

[0026] Optionally, the PDE multi-scale convolution module performs depth separable convolution on the learnable tensor based on a Gaussian kernel, a Laplacian kernel and four different direction Sobel kernels respectively, to generate Gaussian features, Laplacian features and Sobel features, and simultaneously calculate a PDE regularization term.

[0027] The four different direction Sobel kernels are Sobel kernels in four gradient directions of 0°, 45°, 90° and 135°.

[0028] The physical gating mechanism calculates energy values of the Gaussian features, the Laplacian features and the Sobel features respectively to convert into respective physical fusion weights, performs weighted fusion, and then performs element-by-element multiplication operation on the learnable tensor to generate physical gating fusion features.

[0029] The physical enhancement multi-head self-attention mechanism comprises: performing linear mapping and dimension adjustment on the physical quantity, and stacking the physical quantity into the physical gating fusion features to obtain physically enhanced query, key and value for calculating multi-head self-attention; and performing residual connection on the physical gating fusion features and the physical attention features and performing layer normalization to obtain physical attention features.

[0030] The physical gating feedforward network comprises: inputting the physical attention features into a traditional feedforward network layer to obtain feedforward features; mapping the physical quantity into a physical embedding vector with the same dimension as the physical gating fusion features, and obtaining gating weights through nonlinear activation; performing weighted fusion on the physical attention features and the feedforward features based on the gating weights to obtain attention feedforward features; performing residual connection on the attention feedforward features and the physical attention features and performing layer normalization to obtain final output features.

[0031] ​​Optionally, the Gaussian kernel, Laplacian kernel, four different direction Sobel kernel, respectively, to the learnable tensor depth separable convolution, generate Gaussian feature, Laplacian feature, Sobel feature, while calculating PDE regularization term, including:

[0032] The Gaussian kernel, Laplacian kernel, four different direction Sobel kernel, are copied in the channel dimension, so that they are aligned with the feature channel number of the learnable tensor to constitute the basis kernel, further introducing the learnable kernel parameter, forming Gaussian depth separable convolution kernel, Laplacian depth separable convolution kernel, four different direction Sobel depth separable convolution kernel;

[0033] Introducing the learnable diffusion tensor parameter to constitute the learnable expansion tensor matrix :

[0034] ;

[0035] Among them, , Respectively, the learnable diffusion tensor parameter controls the diffusion intensity of two main directions, The learnable diffusion tensor parameter controls the coupling between the two main directions;

[0036] Using the learnable diffusion tensor parameter to process each depth separable convolution kernel As follows:

[0037] ;

[0038] Generate learnable PDE convolution kernel , the learnable PDE convolution kernel Including learnable Gaussian PDE convolution kernel, learnable Laplacian PDE convolution kernel, four different direction learnable Sobel-PDE convolution kernel, respectively, to constitute Gaussian convolution layer, Laplacian convolution layer, four different direction Sobel convolution layer, and further respectively to the learnable tensor Gaussian depth separable convolution, Laplacian depth separable convolution, four different direction Sobel depth separable convolution, generate Gaussian feature, Laplacian feature, four different direction initial Sobel feature;

[0039] Linear scaling and global average pooling are performed on the four different direction initial Sobel features, and then splicing and inputting the fully connected layer are performed, converting into four direction corresponding weights, based on the weight, the four direction initial Sobel features are weighted and fused to obtain Sobel feature;

[0040] The sum of the square of the L2 norm of each depth separable convolution kernel corresponding to the basis kernel and each learnable PDE convolution kernel is calculated as the PDE regularization term.

[0041] Optionally, the energy values of the high-frequency features, Laplacian features, and Sobel features are calculated respectively to convert into respective physical fusion weights, including:

[0042] The absolute values of the feature values of the high-frequency features, Laplacian features, and Sobel features at each spatial position are averaged respectively to obtain respective energy values, which are sent to a multilayer perception to convert into physical fusion weights corresponding to the high-frequency features, Laplacian features, and Sobel features through a softmax operation.

[0043] Optionally, the physical gating fusion features are converted into sequence low-frequency energy, sequence gradient energy, and second-order curvature energy, including:

[0044] The physical gating fusion features are reshaped into an input feature sequence, the sequence length is , and the th element in the sequence is denoted as .

[0045] The sequence low-frequency energy is calculated based on the average amplitude :

[0046] .

[0047] wherein is an absolute value symbol;

[0048] The sequence gradient energy is constructed based on the change speed :

[0049] .

[0050] .

[0051] wherein represents a change amount of the th input feature sequence to the th input feature sequence;

[0052] The second-order curvature energy is constructed based on the mutation degree :

[0053] .

[0054] .

[0055] Optionally, the total loss is constructed to train the cross-working-condition bearing fault classification model, including:

[0056] The final output features of the source domain training set extracted by the cross-condition bearing fault classification model are first dimensionally reduced by a projection head, then normalized, so that all features are located on a unit sphere, and then a similarity matrix is constructed between each other to calculate the InfoNCE contrast loss as the source domain contrast learning loss;

[0057] The sequence low-frequency energy difference, sequence gradient energy difference, and second-order curvature energy difference of the source domain and the target domain are calculated by the cross-condition bearing fault classification model for the source domain training set and the target domain verification set respectively, and the three differences are combined to construct a physical difference quantity, and the physical quantity difference is input into a multilayer perception to be converted into a physical weight;

[0058] The final output features from the source domain training set and the target domain verification set are normalized and a Gaussian kernel is selected, and then the Gaussian kernel MMD of the two is calculated, and the discrete estimate of the square of the Gaussian kernel MMD is taken as the physical adaptive MMD domain alignment loss;

[0059] The physical weight is applied to the physical adaptive MMD domain alignment loss, and combined with the source domain contrast learning loss as the total transfer loss;

[0060] The cross-entropy loss of the predicted fault class generated by the cross-condition bearing fault classification model for the source domain training set and the corresponding fault class label is calculated as the classification loss;

[0061] The weighted sum of the classification loss, the total transfer loss, and the PDE regularization term is taken as the total loss.

[0062] Optionally, the classifier comprises: a multilayer perception based Gaussian expert path, Laplacian expert path, and Sobel expert path are constructed respectively, the final output features from the physically driven Transformer are converted into Gaussian classification probability, Laplacian classification probability, and Sobel classification probability, and physical quantities are extracted from the final output features, the physical quantities are converted into the gating expert weights corresponding to the Gaussian expert path, Laplacian expert path, and Sobel expert path through a gating network, the Gaussian classification probability, Laplacian classification probability, and Sobel classification probability are weighted and fused to obtain a bearing fault classification probability and converted into a predicted fault class.

[0063] By the above technical solution, the present application provides a physically knowledge-driven transfer learning cross-condition bearing fault classification method, which has at least the following beneficial effects:

[0064] (1) The application effectively overcomes the problems of low recognition rate, poor cross-condition generality and insufficient classification accuracy of traditional bearing fault classification methods in strong noise environment, and through the introduction of a physical knowledge driven transfer learning mechanism, the application can realize accurate identification of bearing faults under complex noise conditions, significantly improve the accuracy and stability of the fault classification results, and thus achieve the purpose of improving the generality and overall identification performance of the cross-condition bearing fault classification model;

[0065] (2) The application realizes accurate modeling of key physical features of bearing vibration signals by combining physical knowledge constraints with a multi-scale convolution structure, the PDE multi-scale convolution calculates energy values using Gaussian, Sobel and Laplace kernels and direction modulation is performed by a learnable diffusion tensor; the physical gating fusion mechanism adaptively adjusts the branch weights according to physical quantities such as energy, gradient and curvature, and realizes effective fusion and noise-resistant enhancement of multi-scale physical features;

[0066] (3) The application adopts a physically driven Transformer to map the physical quantities in the features extracted by the PDE multi-scale convolution module to attention bias, guide the model to focus on feature patterns with physical meaning under different working conditions, and enhance the robustness to strong noise and non-stationary signals through a physical gating feedforward network;

[0067] (4) The application performs transfer learning based on a physical adaptive MMD domain alignment loss, dynamically adjusts the alignment strength according to the physical differences between the source domain and the target domain, and combines source domain contrast learning to improve intra-class aggregation and inter-class separability, thereby enhancing the cross-condition transfer capability;

[0068] (5) The application constructs a physical guided multi-expert classifier to model different physical features through Gaussian, Laplace and Sobel expert paths respectively, and dynamically selects the optimal expert through a physical quantity driven gating network, realizes fault discrimination with physical interpretation, and significantly improves the stability and accuracy of cross-condition bearing fault classification. BRIEF DESCRIPTION OF DRAWINGS

[0069] The accompanying drawings, which are included to provide a further understanding of the application and constitute a part of this application, illustrate certain illustrative embodiments of the application and together with the description serve to explain the application. In the drawings:

[0070] Figure 1 A physical knowledge driven transfer learning cross-condition bearing fault classification method flowchart of the application;

[0071] Figure 2 A cross-condition bearing fault classification model overall framework diagram constructed by the application;

[0072] Figure 3 The comparative bar chart of the bearing fault classification performance of the method proposed in the application and other existing methods under-4dB noise condition;

[0073] Figure 4 The target domain t-SNE feature visualization result of the method proposed in the application under the 1hp→2hp load condition with a signal-to-noise ratio of-4dB;

[0074] Figure 5 The target domain t-SNE feature visualization result of the method proposed in the application under the 1hp→3hp load condition with a signal-to-noise ratio of-4dB;

[0075] Figure 6 The target domain t-SNE feature visualization result of the method proposed in the application under the 2hp→3hp load condition with a signal-to-noise ratio of-4dB;

[0076] Figure 7 The target domain t-SNE feature visualization result of the method proposed in the application under the 3hp→2hp load condition with a signal-to-noise ratio of-4dB;

[0077] Figure 8 The target domain t-SNE feature visualization result of the method proposed in the application under the 3hp→1hp load condition with a signal-to-noise ratio of-4dB;

[0078] Figure 9 The target domain t-SNE feature visualization result of the method proposed in the application under the 2hp→1hp load condition with a signal-to-noise ratio of-4dB. DETAILED DESCRIPTION

[0079] In order to make the above objectives, features and advantages of the present application more apparent, the present application will be described in further detail below with reference to the accompanying drawings and specific embodiments. The implementation process of how the present application applies technical means to solve technical problems and achieve technical effects can be fully understood and implemented by the description.

[0080] Those skilled in the art can understand that all or part of the steps in the embodiment method can be completed by programs instructing related hardware, therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product in the form of being implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.

[0081] Reference should be made to Figures 1-9, shows a specific embodiment of the embodiment, the embodiment constructs a source domain and a target domain data set by collecting bearing driving end vibration acceleration signals; a PDE multi-scale convolution module is constructed, depth separable convolution is performed based on Gaussian kernel, Laplacian kernel, four different direction Sobel kernels, and feature fusion is performed based on energy value calculation weight; a physically driven Transformer is constructed, physical quantities are extracted to construct a physically enhanced multi-head self-attention mechanism, and a traditional feedforward network is optimized in combination with a gating mechanism; a classifier based on multiple expert paths and physical gating is designed; a physically adaptive MMD domain alignment loss based on Gaussian kernel MMD is constructed, a source domain contrast learning loss for the source domain is constructed, and a total transfer loss is constructed, combined with a classification loss, to train a cross-condition bearing fault classification model; the trained cross-condition bearing fault classification model is used to output a bearing fault classification result for the measured bearing driving end vibration acceleration signal, and bearing fault classification under different working conditions is realized.

[0082] Please refer to Figure 1 The embodiment proposes a physically knowledge driven transfer learning cross-condition bearing fault classification method, which comprises the following steps:

[0083] S1, collecting driving end vibration acceleration signals of the bearing under different working conditions by using a multi-point arranged acceleration vibration sensor, and constructing an original vibration sequence under multiple working conditions in time order in combination with time index.

[0084] As a preferred embodiment of step S1, the specific process comprises:

[0085] S11, selecting multiple measuring points on the bearing seat surface of the test bench, and fixedly installing acceleration vibration sensors to ensure that the sensors are in close contact with the bearing seat and the signal transmission is stable and reliable;

[0086] S12, running the bearing under different working conditions (such as different load conditions 1 hp, 2 hp, 3 hp), and collecting driving end vibration acceleration signals of the bearing under normal state and each fault state (inner ring fault, outer ring fault, rolling element fault, etc.) under each working condition by the acceleration vibration sensor, the sampling frequency is set to a fixed value (in this embodiment, it is set to 12 kHz), the collected driving end vibration acceleration signals are used as original vibration data, and an original vibration sequence is constructed in time order. The collected driving end vibration acceleration signals contain typical bearing vibration components such as impact wave trains generated by bearing defects and low-frequency modulation caused by load. The constructed original vibration sequence includes driving end vibration acceleration signals arranged in time order and time index.

[0087] S2, data division: unsupervised data preprocessing.

[0088] The original vibration sequence under one working condition is taken as the source domain signal, and the original vibration sequence under another working condition is taken as the target domain signal, and then, in an unsupervised manner, a source domain training set is constructed from the source domain signal, and a target domain verification set and a target domain test set are constructed from the target domain signal.

[0089] As a preferred embodiment of step S2, the specific process includes:

[0090] S21, signal slicing and sample reconstruction. Each driving end vibration acceleration signal in the original vibration sequence from the source domain signal and the target domain signal is cut into a plurality of signal sample units of uniform length by using a sliding window cutting manner, to form a source domain original sample library and a target domain original sample library. In this embodiment, the length of each signal sample unit is set to 1000 .

[0091] S22, random division of the source domain and the target domain without label basis. According to the time index in the original vibration sequence, source domain samples and target domain samples are sampled from the source domain original sample library and the target domain original sample library, respectively; specifically:

[0092] A time section threshold is set , and for the source domain original sample library and the target domain original sample library constructed based on the original vibration sequence with a total length of ;

[0093] In the part of the target domain original sample library with the first time index, a number of signal sample units are randomly selected as the source domain samples, to form a source domain sample set , wherein is a sample number control parameter;

[0094] In the part of the target domain original sample library with the remaining time index, a number of signal sample units are randomly selected as the target domain samples, to form a target domain sample set .

[0095] This process does not use fault category labels, but only randomly intercepts signals according to time sections, to ensure that the target domain is strictly unsupervised, wherein are the source domain sample index and the target domain sample index, respectively, indicate the “source domain” and the “target domain”, respectively.

[0096] In this embodiment, the length of each signal sample unit is set to 1000 , .

[0097] S23, source domain label mapping and one-hot encoding. According to the fault type (including normal state and each fault state), the source domain sample is assigned a fault category label represents the total number of fault types, which is set to 9 in this embodiment), and one-hot encoding is performed to obtain a labeled source domain sample , which constitutes the source domain training set.

[0098] The one-hot form ensures that the output dimension of the classifier is fixed, and the label category corresponds to different bearing damage modes.

[0099] S24, uniform standardization of cross-condition data. The source domain samples are stacked into a matrix , where is a real number space, and the target domain samples form a matrix ;

[0100] The mean and standard deviation are calculated based on the source domain:

[0101] ;

[0102] ;

[0103] Subsequently, uniform standardization is performed to obtain the standardized source domain sample matrix and the standardized target domain sample matrix :

[0104] ;

[0105] Uniform standardization using source domain statistics can avoid the influence of vibration amplitude differences under different working conditions on feature learning, and ensure that the model focuses on "fault impact mode differences" rather than "energy differences caused by different loads".

[0106] S25, strict unsupervised target domain validation set and test set division. The standardized target domain sample matrix is constructed as a set represents the th standardized target domain sample), and its index is completely randomly shuffled:

[0107] ;

[0108] wherein is the newly generated index after random shuffling, represents a random permutation operation; based on , the set ​​The samples in the source domain are rearranged, and then divided into a target domain verification set and a target domain test set in a 1:1 ratio. No target domain label information is used in the division process to ensure that the target domain is strictly unsupervised.

[0109] The embodiment adopts a random permutation operation , and the action object is an index sequence rather than the sample itself. When the input is , the output is a permutation sequence with a length of , which satisfies:

[0110] (1) (the elements are not repeated or missing);

[0111] (2) the order is randomly disturbed;

[0112] The meaning of using is that the division is completely random and unsupervised without relying on the target domain label.

[0113] This step uniformly preprocesses and divides the source domain and target domain data without using the target domain label information, so that the target domain is strictly unsupervised in the process of transfer learning.

[0114] S3, build a bearing fault classification model across working conditions.

[0115] As a preferred embodiment of step S3, the specific process includes:

[0116] S31, initial feature extraction is performed using a backbone network to convert the input signal into a learnable tensor. Specifically:

[0117] First, one-dimensional convolution and pooling are used to extract wide-frequency time sequence features, including low-frequency structural vibration components and high-frequency impact components, which are represented by the formula:

[0118] ;

[0119] wherein represents the input signal, represents a one-dimensional convolution operation with a kernel size of 7x7, a step of 2, and 128 convolution kernels, , respectively represent the bias term and the output feature time step of the one-dimensional convolution, is a ReLU function, represents the features extracted by the first layer of one-dimensional convolution;

[0120] Then, maximum pooling is applied to compress the time dimension, retaining the dominant peak features, which is particularly important for rolling body impact type faults, and is mathematically represented as:

[0121] ;

[0122] wherein, represents a one-dimensional max-pooling operation with kernel size 2x2, , respectively represent the output feature and the output feature time step of the one-dimensional max-pooling;

[0123] The one-dimensional convolution is used again to refine the local features and keep the number of channels consistent, which is mathematically expressed as:

[0124] ;

[0125] wherein, represents a one-dimensional convolution operation with kernel size 3x3, step size 1, and 128 convolution kernels, is the bias term of the one-dimensional convolution, is the one-dimensional feature extracted by the one-dimensional convolution;

[0126] Reshape from one-dimensional feature to two-dimensional feature, which is mathematically expressed as:

[0127] ;

[0128] wherein, the two-dimensional feature is the learnable tensor, represents the length of the time dimension; represents the length of the pseudo-space dimension; is the number of feature channels. The axis direction of the time dimension and the pseudo-space dimension is defined as two main directions.

[0129] S32, constructing a PDE multi-scale convolution module based on physical priori. Specifically:

[0130] S321, first construct Gaussian kernel , Laplacian kernel , Sobel kernel of 0°, 45°, 90°, 135° four gradient directions , , , ;

[0131] (1) Design Gaussian kernel , which is used to suppress noise and preserve smooth structure, corresponding to the diffusion process in the heat conduction equation. In bearing vibration, low-frequency components reflect more slowly changing components such as rotor imbalance, foundation vibration, and load fluctuation. The designed Gaussian kernel is in the form of:

[0132] ;

[0133] (2) Design Laplacian kernel , approximate second-order derivative, highlight local mutation points, can emphasize the impact response of local peeling of rolling body and inner and outer rings. The designed Laplacian kernel is of the form:

[0134] ;

[0135] (3) Design four different direction Sobel kernels for different direction gradients. These directions actually correspond to different local slopes and impact rising and falling edges, which can capture different waveform patterns of fault impact. The four different direction Sobel kernels designed are of the form:

[0136] ; ; ;

[0137] .

[0138] In this embodiment, 3x3 is the preferred size for constructing the kernel. This size is the most common, small in computational overhead, and stable in training for standard discrete PDE / edge operators. Without changing the core idea of the present application, the kernel size can be extended to 5x5, 7x7, and other odd sizes, but it needs to meet the structural properties of the corresponding operator (such as normalization of smoothing kernel, zero and / or symmetry of derivative kernel, etc.). These can be equivalent alternatives to the present embodiment.

[0139] Among the three branches constructed: the Gaussian branch is used to suppress noise and retain smooth structure, which is equivalent to "diffusion, smoothing" modeling in the feature layer, and is more robust for low-frequency structural components under strong noise, providing a stable source for subsequent "physical energy index (low-frequency energy) + gating fusion"; the Laplacian branch highlights local mutation points and emphasizes the impact response of local peeling of rolling body and inner and outer rings, which is highly matched with the typical "impact-attenuation-reimpact" mutation structure of bearing fault, and is more sensitive to sharp impact and pitting peeling induced mutations; it is complementary to the smoothing branch based on Gaussian kernel (one preserves trend, one catches mutation), and is more likely to maintain discriminant structure under cross-condition / strong noise; the Sobel branch uses different local slopes and impact rising / falling edges to capture different waveform patterns of fault impact, and can adaptively strengthen "more significant impact edge direction" under different conditions / strong noise, improving robustness and interpretability, avoiding learning "pseudo-edge" without physical meaning by pure data-driven convolution, and enhancing cross-condition generalization stability.

[0140] S322, replicate the above Gaussian kernel, Laplacian kernel, and four different direction Sobel kernels in the channel dimension, so that they are Aligning, constituting basic kernels, further introducing learnable kernel parameters into each basic kernel to form Gaussian deep separable convolution kernels , Laplacian deep separable convolution kernels , four different direction Sobel deep separable convolution kernels ;

[0141] Introducing learnable diffusion tensor parameters to constitute a learnable diffusion tensor matrix :

[0142] ;

[0143] wherein, , are learnable diffusion tensor parameters controlling the diffusion strength of the two main directions (i.e. the axis direction of the time dimension and the pseudo-space dimension), is a learnable diffusion tensor parameter controlling the coupling between the two main directions;

[0144] The diffusion strength is the directional "smoothing / propagation" strength control of the PDE convolution kernel output / response, and for the learnable diffusion tensor matrix :

[0145] (1) The larger it is: it means that the diffusion along the x main direction (i.e. the axis direction of the time dimension) is stronger → the feature is more inclined to smooth and aggregate in this direction (more noise suppression, more emphasis on continuous structure);

[0146] (2) The larger it is: it means that the diffusion along the y main direction (i.e. the axis direction of the pseudo-space dimension) is stronger → the feature is more inclined to smooth and aggregate in this direction;

[0147] (3) Coupling diffusion is embodied: let the information not only propagate along pure x or pure y, but also propagate along the "oblique / rotated direction", so as to change the direction preference and response form of the kernel.

[0148] In general, bearing impact / spalling often presents as "abrupt edge + non-stationary waveform form", and under strong noise / cross-condition conditions, the fixed direction kernel may not be stable; after introducing the learnable diffusion tensor matrix, the model can adaptively select "which direction to smooth, which direction to retain the edge / abruptness", so as to balance noise suppression and impact structure preservation.

[0149] Use learnable diffusion tensor parameters to each deep separable convolution kernel (including , , , , , ) are processed as follows:

[0150] ;

[0151] Generating learnable PDE convolution kernels , the learnable PDE convolution kernels include a learnable Gaussian PDE convolution kernel , a learnable Laplacian PDE convolution kernel , four learnable Sobel-PDE convolution kernels in different directions , respectively, constitute a Gaussian convolution layer, a Laplacian convolution layer, four Sobel convolution layers in different directions, and then respectively perform Gaussian depth separable convolution, Laplacian depth separable convolution, and four Sobel depth separable convolutions in different directions on the learnable tensor to generate Gaussian features , Laplacian features , and four initial Sobel features in different directions ; this process is mathematically represented as follows:

[0152] ;

[0153] ;

[0154] ;

[0155] ;

[0156] ;

[0157] ;

[0158] wherein represents a depth separable convolution, , , , , , respectively represent Gaussian depth separable convolution, Laplacian depth separable convolution, and Sobel depth separable convolution in four different gradient directions of 0°, 45°, 90°, and 135°;

[0159] More specifically, the four initial Sobel features in different directions are linearly scaled by a scaling factor , , , , and then globally average-pooled as follows:

[0160] ;

[0161] where, denotes the linearly scaled initial Sobel feature, denotes four different gradient directions, denotes the pooled Sobel feature;

[0162] The pooled Sobel features of four different directions are concatenated:

[0163] ;

[0164] denotes the concatenated Sobel feature;

[0165] The concatenated Sobel feature is input into the fully connected layer and converted into weights :

[0166] ;

[0167] where, denotes the learnable weight matrix of the fully connected layer, which is used to learn the coupling manner between directions; denotes the bias term of the fully connected layer; denotes the softmax function, which forces the contribution of each direction feature to be normalized, forming a competitive mechanism;

[0168] Based on the weights of four directions , the initial Sobel features of four directions are weighted and fused to obtain the Sobel feature , which is mathematically represented as follows:

[0169] ;

[0170] This step can be regarded as "selecting which first-order gradient of which direction is more important", and the model can dynamically select the edge and abrupt structure of different directions.

[0171] Next, the PDE regularization term is calculated:

[0172] Define a set of basic PDE kernels with clear physical meaning :

[0173] ; where are the depthwise separable convolution kernels corresponding to the basic kernels.

[0174] Further define the PDE regularization term​ :

[0175] ;

[0176] where, is the regularization coefficient. The PDE regularization term enables the model to preserve the physical meaning of PDE convolution while allowing moderate learning.

[0177] S323, calculate the energy value of the high-frequency feature, Laplace feature and Sobel feature respectively to convert into the respective physical fusion weight, specifically:

[0178] In order to evaluate the signals of different sizes and amplitudes fairly, use "average absolute energy" as a unified physical index. The three energy values calculated are low-frequency branch energy , gradient branch energy , and second-order change energy , which are mathematically represented as follows:

[0179] ;

[0180] ;

[0181] ;

[0182] where, are the time dimension index, pseudo space dimension index, and feature channel index of , respectively; are the time dimension index, pseudo space dimension index, and feature channel index of , respectively; are the time dimension index, pseudo space dimension index, and feature channel index of , respectively;

[0183] Concatenate the three energy values into a three-dimensional physical vector :

[0184] ;

[0185] Learn the physical fusion weight of the three branch paths through a multi-layer perception (including two fully connected layers) and softmax :

[0186] ;

[0187] where, represents the sigmoid nonlinear activation; respectively represent the learnable weight matrix and bias term of the first fully connected layer of this multi-layer perception, which are used to convert the three-dimensional physical vector Map to latent space dimensions; Respectively represent the learnable weight matrix and bias term of the full connection layer of the second layer of the multi-layer perception, used for outputting the unnormalized weights of the three branches.

[0188] Through the above operation, the model regards the three types of PDE information of Gaussian feature, Laplace feature and Sobel feature as a competitive relationship, and determines which type of physical structure should be relied on according to physical knowledge.

[0189] S324, according to the physical fusion weight The Gaussian feature, Laplace feature and Sobel feature are weighted and fused, and then multiplied with the learnable tensor to generate the physical gate fusion feature :

[0190] ;

[0191] This step embodies the "physical knowledge driven", when the low frequency branch energy is large, it indicates that the structure vibration is dominant, and the model tends to the Gaussian feature branch; when the gradient branch energy or the second order change energy is large, it indicates that the impact and mutation are strong, and the model relies more on the Sobel feature branch and the Laplace feature branch.

[0192] The application combines physical knowledge constraints and multi-scale convolution structures to realize precise modeling of key physical features of bearing vibration signals, and the PDE multi-scale convolution calculates energy values using Gaussian, Sobel and Laplace kernels and direction modulation by a learnable diffusion tensor; the physical gate fusion mechanism adaptively adjusts the branch weight according to energy, gradient and curvature, etc. Physical quantities realize effective fusion and noise resistance enhancement of multi-scale physical features.

[0193] S33, construct a physically driven Transformer. Specifically:

[0194] S331, the physically driven Transformer receives the physical gate fusion feature as its input feature, extracts physical quantities; the extraction of physical quantities is specifically:

[0195] Reshape the physical gate fusion feature into an input feature sequence, the sequence length is , and the th element in the sequence is denoted as ; this embodiment takes as an example:

[0196] The sequence low frequency energy is calculated based on the average amplitude :

[0197] ;

[0198] in, To determine the absolute value sign;

[0199] Constructing sequence gradient energy based on the rate of change :

[0200] ;

[0201] ;

[0202] in, Indicates the first The input feature sequence is then used to... The amount of change in each input feature sequence;

[0203] Constructing second-order curvature energy based on the degree of mutation :

[0204] ;

[0205] .

[0206] in, It reflects the overall energy level, the load size, and the vibration intensity; This represents the rate of change of the waveform and whether the impact edge is obvious; It reflects the sharpness of the impact and can distinguish between early minor peeling and serious failure.

[0207] Ultimately forming physical quantities : .

[0208] S332. Traditional Transformers suffer from the following drawbacks in industrial time series: lack of physical priors, inability to interpret the direction of attention under various operating conditions, and susceptibility to noise dominance. To address these issues, this invention incorporates physical quantities... Injected into the Transformer attention structure, a physically enhanced multi-head self-attention mechanism and a physically gated feedforward network are designed, combined with residual connections and layer normalization to obtain the final output features; specifically:

[0209] S3321. Construct a physically enhanced multi-head self-attention mechanism, including:

[0210] (1) For physical quantities Perform a linear mapping to obtain the physical biases corresponding to the query, key, and value. , , :

[0211] ;

[0212] ;

[0213] ;

[0214] wherein, are the learnable projection matrices corresponding to the query, key, value, respectively, mapping the 3-dimensional physical quantities to the 128-dimensional Transformer space; are the learnable bias vectors corresponding to the query, key, value, respectively, providing adjustable baseline shifts on top of the linear mapping of physical quantities, so that stable physical biases can be generated when the physical quantity magnitude is close to or at zero. The dimensions of the bias vectors are consistent with the corresponding physical bias vectors. The physical quantities determine that the Transformer should pay more attention to the overall energy change (low frequency), more attention to the impact change speed (first derivative), and more attention to the sharp mutation (second derivative), which is the key source of the model's cross-condition migration ability.

[0215] The physical bias is extended to be the same as the physical gating fusion feature and is superimposed, obtaining the physically enhanced query , key , value :

[0216] ;

[0217] The physical statistical quantity determines from which angle the Transformer looks at the current feature, thereby embodying the modulation of physical prior on attention.

[0218] (3) Calculate multi-head self-attention:

[0219] ;

[0220] wherein, denotes the multi-head self-attention calculation, denotes the transpose operation, denotes the feature dimension of the key ; then, after the dropout layer, the physical gating fusion feature is connected in residual and layer normalization is performed to obtain the physical attention feature :

[0221] ;

[0222] wherein, denotes the layer normalization operation, denotes the dropout layer.

[0223] S3322, a physical gating feedforward network is constructed, comprising:

[0224] The physical attention feature is input into a traditional feedforward network layer (including two linear mappings), to obtain a feedforward feature :

[0225] ;

[0226] wherein, respectively represent the learnable weight matrix and bias term of the first linear mapping of this feedforward network layer, for projecting the physical attention feature to an intermediate hidden layer; respectively represent the learnable weight matrix and bias term of the second linear mapping of this feedforward network layer, for mapping the feature projected to the intermediate hidden layer back to the output space.

[0227] This feedforward network layer is completely data-driven, lacking physical meaning. The present application designs a fusion mechanism based on physical quantities as follows:

[0228] (1) The physical quantity is mapped into a physical embedding vector of the same dimension as the physical gating fusion feature:

[0229] ;

[0230] wherein, represents a mapping layer, , respectively represent the learnable weight matrix and bias term of this mapping layer;

[0231] Through a nonlinear activation, a gating weight is obtained:

[0232] ;

[0233] wherein, represents a sigmoid nonlinear activation, , respectively represent the learnable weight matrix and bias term of this sigmoid nonlinear activation, represents the dimension of the gating weight;

[0234] Based on the gating weight , the physical attention feature and the feedforward feature are weightedly fused and residual connection calculation is performed, to obtain a final output feature , and this process is mathematically represented as follows:

[0235] ;

[0236] wherein, is element-wise multiplication.

[0237] Gating weights Each dimension determines whether to trust the output of the feedforward network or the output of the physically enhanced multi-head self-attention mechanism more. When the physical quantity shows a certain impact or energy pattern, the corresponding dimension of the gate may be more biased towards the feedforward feature , so as to make the model migrate to a feature space more suitable for the physical scene.

[0238] The application adopts a physically driven Transformer, maps the physical quantity in the feature extracted by the PDE multi-scale convolution module as an attention bias, guides the model to pay attention to the feature mode with physical meaning under different working conditions, and enhances the robustness to strong noise and non-stationary signals through a physically gated feedforward network.

[0239] The application constructs a physically driven Transformer based on PDE multi-scale convolution, forms a migration model, and realizes unified representation learning of vibration signals in source and target domains.

[0240] S34, construct a classifier. Specifically:

[0241] S341, respectively construct a Gaussian expert path, a Laplace expert path and a Sobel expert path based on a multi-layer perception (including one mapping layer), and convert the final output feature from the physically driven Transformer into a Gaussian classification probability , a Laplace classification probability , and a Sobel classification probability .

[0242] ;

[0243] ;

[0244] ;

[0245] wherein, respectively represent the mapping layers corresponding to the Gaussian expert path, the Laplace expert path and the Sobel expert path; respectively represent the learnable weight matrix and the bias term of the mapping layer corresponding to the Gaussian expert path; respectively represent the learnable weight matrix and the bias term of the mapping layer corresponding to the Laplace expert path; respectively represent the learnable weight matrix and the bias term of the mapping layer corresponding to the Sobel expert path.

[0246] Gauss expert is better at identifying failure modes related to low-frequency structure; Sobel expert is more sensitive to impact edges, periodic impacts; Laplacian expert pays more attention to sharp mutations, such as severe spalling, pitting, etc.

[0247] The above three mapping layers The structure is the same, but the parameters are independent of each other and do not share weights. The input features of each mapping layer have been shown to decompose different physical patterns in the upstream step, and naturally converge to different discriminant boundaries through the respective "parameter-independent" branch in this classifier.

[0248] S342, extracting physical quantities from the final output features of the Transformer from the physical drive :

[0249] ;

[0250] wherein, respectively represent sequence low-frequency energy calculation, sequence gradient energy calculation, and second-order curvature energy calculation;

[0251] The physical quantity is converted into the gating expert weights corresponding to the Gauss expert path, the Laplacian expert path, and the Sobel expert path through a gating network (including two mapping layers) :

[0252] ;

[0253] wherein, represents a weight group composed of respectively represent the learnable weight matrix and bias term of the first mapping layer of the gating network; respectively represent the learnable weight matrix and bias term of the second mapping layer of the gating network.

[0254] When the sequence low-frequency energy of the vibration signal is high, the gating will be more biased towards the Gauss expert; when the impact is sharp, the weights of the Sobel and Laplacian experts will be increased, so that the classification decision has a clear physical explanation, thereby realizing the physical-guided expert fusion: the bearing fault classification probability , the Laplacian classification probability , and the Sobel classification probability are weighted and fused to obtain the bearing fault classification probability :

[0255] ;

[0256] The bearing fault classification probability is converted into a predicted fault category.

[0257] The application realizes fault discrimination with physical interpretation by constructing a physical guided multi-expert classifier to model different physical characteristics through Gaussian expert paths, Laplace expert paths and Sobel expert paths respectively, and dynamically selecting the optimal expert by a physical quantity driven gating network, thereby significantly improving the stability and accuracy of bearing fault classification across working conditions.

[0258] S4, constructing a source domain contrast learning loss for the source domain, constructing a total transfer loss between the source domain and the target domain, introducing a PDE regularization term, combining a classification loss to construct a total loss, training the cross-working-condition bearing fault classification model based on the source domain training set and the target domain validation set, gradually optimizing the model parameters and adjusting the model structure until convergence, and obtaining an optimal cross-working-condition bearing fault classification model.

[0259] As a preferred embodiment of step S4, the specific process includes:

[0260] S41, calculating the sequence low-frequency energy , sequence gradient energy , and second-order curvature energy difference of the source domain and the target domain respectively by using the cross-working-condition bearing fault classification model on the source domain training set and the target domain validation set.

[0261] ;

[0262] This vector measures the gap between the source domain and the target domain in the core physical behavior.

[0263] Input the physical difference quantity into a multi-layer perception (including two fully connected layers) to obtain a physical weight (value 0-1):

[0264] ;

[0265] wherein, respectively represent the learnable weight matrix and bias term of the first layer of the fully connected layer of the multi-layer perception, which is used to project the physical difference quantity from three dimensions to the hidden layer feature space; respectively represent the learnable weight matrix and bias term of the second layer of the fully connected layer of the multi-layer perception, which is used to map the hidden layer feature representation to a scalar.

[0266] When the cross-working-condition difference is large, is large, which means that stronger domain alignment constraints are needed; when the source domain and the target domain are close in physical characteristics, is small, so as not to damage the discrimination structure by too strong alignment. This is the "physical adaptive alignment strength adjustment" that traditional MMD cannot achieve. ​

[0267] S42. Normalize the final output features from the source and target domains and select a Gaussian kernel. The formula for calculating the Gaussian kernel is:

[0268] ;

[0269] In the above formula Represents the Gaussian kernel function. This indicates that two quantities need to be calculated for the Gaussian kernel, corresponding to the final output features from the source and target domains in this invention. Indicates bandwidth parameter, This represents the natural exponential function. This represents the calculation of the square of the L2 norm;

[0270] Calculate the square of the MMD between the final output features from the source domain and the target domain. as follows:

[0271] ;

[0272] in, These represent the final output features from the source domain and the target domain, respectively. , Let represent the probability distributions of the final output features from the source domain and the target domain, respectively. Indicates from Two samples were drawn independently. and Calculate them in the Gaussian kernel The similarity is calculated, and the expected value of that similarity is taken. In the definition of MMD and This usually represents two independent and identically distributed samples; and They represent from Two samples were drawn independently from the sample;

[0273] Its discrete estimate is:

[0274] ;

[0275] Indicates the first from the source domain The final output feature, Indicates the first term of the target domain One final output feature; Indicates the use of calculation The batch size is the number of samples, which is the number of source domain features participating in the statistics in one iteration (the same as the number of target domain features participating in the statistics).

[0276] The discrete estimation of the square of the obtained Gaussian kernel MMD That is, the physical adaptive MMD domain alignment loss.

[0277] S43, using the cross-condition bearing fault classification model to extract each final output feature of the source domain training set, first through the projection head for dimension reduction, and then normalized, so that all features are located in the unit sphere:

[0278] ;

[0279] ;

[0280] wherein, is the projection head, is the output feature of the projection head, , respectively, the learnable weight matrix and the bias term of the projection head, denotes the L2 norm calculation, denotes the spherical feature obtained by normalizing the output feature of the projection head, is the dimension of the output feature of the projection head.

[0281] S44, a similarity matrix is constructed between the spherical features:

[0282] ;

[0283] wherein, denotes the similarity calculation operation, , respectively, the spherical features , ; is the temperature coefficient.

[0284] S45, let is the index set of the same fault category as the sample The fault category of the sample is provided by the fault category label of the source domain training set, and the InfoNCE contrast loss can be written as:

[0285] ;

[0286] wherein, denotes the sample of a different fault category from the sample .

[0287] The InfoNCE contrast loss allows the same fault category to be close to each other in the feature space, and different fault categories to be far away, thereby improving the cross-condition discrimination ability. The InfoNCE contrast loss is the source domain contrast learning loss.

[0288] S46, the physical weight is applied to the physical adaptive MMD domain alignment loss, and combined with the source domain contrast learning loss as the total transfer loss:

[0289] ;

[0290] Wherein, Total transfer loss; Contrast loss weight.

[0291] S47, the classification loss, the cross-entropy loss is calculated between the predicted fault category generated by the cross-condition bearing fault classification model on the source domain training set and the corresponding fault category label.

[0292] S48, the total loss Is:

[0293] ;

[0294] Wherein, Classification loss; , Total transfer loss , PDE regularization term Corresponding weight, used to adjust the importance of the two in the total loss.

[0295] The source domain training set in the embodiment, wherein the data is data enhanced before use, and the target domain verification data in the target domain verification set is not used in the model training process Any label, just participate in The calculation in It is a typical unsupervised domain adaptation.

[0296] The cross-condition bearing fault classification model built, its structure can be seen from Figure 2During model training, firstly, the model parameters are initialized, and the source domain working condition data and target domain working condition data from the source domain training set and the target domain verification set are subjected to initial feature extraction by the backbone network to be converted into learnable tensors, and then in the shared feature space of the source domain and the target domain, Gaussian deep separable convolution, Laplace deep separable convolution, and four different direction Sobel deep separable convolution are respectively performed based on the PDE multi-scale convolution module, wherein the initial Sobel features of the four gradient directions extracted by the Sobel deep separable convolution are fused (this fusion method can also be understood as gated fusion), and then physical feature driven multi-scale gated fusion is performed: the Gaussian features, Laplace features, and Sobel features are weighted and fused according to the physical fusion weight, and the features obtained after weighted fusion are subjected to element-by-element multiplication operation with the learnable tensors to generate physical gated fusion features. The physical gated fusion features are input into the physically driven Transformer, physical quantity automatic extraction is first performed, and then a physical enhanced multi-head self-attention mechanism and a physical gated feedforward network are constructed, combined with addition (residual connection) and layer normalization, to obtain the final output features from the source domain and the target domain (as source domain working condition features and target domain working condition features); according to the source domain working condition features and the target domain working condition features, a physical difference quantity is calculated, a physical self-adaptive MMD domain alignment loss is constructed and weighted; combined with the fault class labels (source domain working condition labels) in the source domain training set, a source domain contrast learning loss is calculated, and at the same time, a classification loss is calculated by using a classifier to calculate the cross-entropy loss between the predicted fault class generated by the source domain and the corresponding fault class label.

[0297] The present application performs transfer learning based on the physical self-adaptive MMD domain alignment loss, dynamically adjusts the alignment strength according to the physical difference between the source domain and the target domain, and combines source domain contrast learning to improve intra-class aggregation and inter-class separability, thereby enhancing the cross-working condition transfer capability. In addition, the present application effectively overcomes the problems of low recognition rate in strong noise environment, poor cross-working condition generality, and insufficient classification accuracy of traditional bearing fault classification methods. By introducing a physical knowledge driven transfer learning mechanism, the present application can realize accurate identification of bearing faults under complex noise conditions, significantly improve the accuracy and stability of the fault classification results, and thus achieve the purpose of improving the generality and overall identification performance of the cross-working condition bearing fault classification model.

[0298] S5, real-time acquisition of bearing driving end vibration acceleration signals, input into the optimal cross-working condition bearing fault classification model, and the predicted fault class output by the model is taken as the bearing fault classification result, so as to realize bearing fault classification under different working condition conditions.

[0299] In a specific implementation, target domain test data in a target domain test set can be input into the optimal cross-condition bearing fault classification model, and the output prediction fault category is used for model performance evaluation, and evaluation indexes (accuracy, precision, recall, and F1 score) are output and combined with t-SNE feature visualization to verify the cross-condition fault classification result.

[0300] The t-SNE feature visualization uses different colors to scatter plot different fault categories, and it can be observed intuitively that after the physical knowledge driven cross-condition migration, whether the same fault category is obviously clustered and different categories are clearly separated in the target domain, and the evaluation result is better in terms of obvious clustering and clear separation. Finally, the bearing fault classification report is generated by combining the bearing fault classification result, the evaluation index, and the t-SNE feature visualization result.

[0301] The physical knowledge driven migration learning cross-condition bearing fault classification method proposed in the present application has been introduced; in addition, the present embodiment is used to verify the performance of the method proposed in the present application, and experiments are performed on the bearing fault data set provided by West Chester University:

[0302] The West Chester University data set includes 1 hp, 2 hp, and 3 hp three power conditions, a total of 27 “.mat” data files of each fault state data of the inner ring, the outer ring, and the rolling body, and three “.mat” files of normal state data. Each file contains a bearing drive end vibration acceleration signal collected by placing an acceleration vibration sensor above the bearing seat on the drive end of the motor under the condition of a sampling frequency of 12 kHz, and all fault bearings are damaged by electric spark machining. The bearing type collected in the West Chester University public bearing fault data set is three fault bearings and one normal bearing, and the fault positions of the fault bearings are the bearing outer ring, the bearing rolling body, and the bearing inner ring. The data files are sampled, and the length of the signal is 2048 each time. The data set obtained by sampling is divided into a source domain training set with fault category labels and a target domain verification set and a target domain test set without fault category labels. The source domain training set accounts for 60%, the target domain verification set accounts for 20%, and the target domain test set accounts for 20%.

[0303] Table 1 and Figure 3 The table below shows the cross-condition bearing fault classification performance comparison of the method (the present model) proposed in the present application and four other existing methods (models: UTL-SOCNN-FBNM, BDTN, DANN, and DDAN) under the condition of -4 dB noise and different load conditions 1 hp, 2 hp, and 3 hp. It can be seen that compared with the four other existing models, the present model has advantages in accuracy, precision, recall, and F1 score under strong noise conditions.

[0304] Table 1. Cross-condition bearing fault classification performance comparison of the proposed method and other existing methods under-4dB noise condition

[0305]

[0306] In the present application, the number of training times is set to 100, the batch size is set to 128, and the learning rate uses the Adam algorithm with an initial value of 0.0001. The cross-condition transfer task of the present embodiment uses the load condition as the domain division basis. The symbol "a hp→b hp" represents that the a hp load condition data constitutes the source domain (provides labels for supervised classification training), and the b hp load condition data constitutes the target domain (does not use target domain labels, and only uses its unlabeled samples to participate in alignment / adaptation training). Finally, the transfer performance is evaluated on the target domain test set.

[0307] Figures 4-9 The target domain t-SNE feature visualization results of the proposed method under the condition of signal-to-noise ratio-4dB, 1hp→2hp load condition, 1hp→3hp load condition, 2hp→3hp load condition, 3hp→2hp load condition, 3hp→1hp load condition, and 2hp→1hp load condition are shown in Table 1. It can be seen that in the feature space constructed by the present application under strong noise conditions, the fault samples in the target domain are significantly separated, so the transfer learning method proposed by the present application is reliable.

[0308] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0309] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor, or other system that can take instructions from an instruction execution system, apparatus or device, or in conjunction with these instructions execution systems, apparatus or devices.

[0310] The above embodiments have been described in detail, and the principles and embodiments of the present application have been described by applying specific examples. The above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific embodiments and application scope will be changed, and the above description should not be understood as a limitation of the present application.

Claims

1. A physical knowledge driven transfer learning cross-condition bearing fault classification method, characterized in that, The application relates to a method for realizing bearing fault classification under different working conditions. The bearing vibration acceleration signals at the driving end under different working conditions are collected by using acceleration vibration sensors arranged at multiple points, and original vibration sequences under multiple working conditions are constructed in time sequence by combining time indexes; From the constructed original vibration sequences under multiple working conditions, an original vibration sequence under an arbitrary working condition is selected as a source domain signal, and an original vibration sequence under a working condition is selected from the original vibration sequences under the remaining working conditions which are not selected as target domain signals, so as to construct a source domain training set, a target domain verification set and a target domain test set in an unsupervised form; the source domain training set contains fault category labels, and the target domain verification set and the target domain test set do not contain fault category labels; A bearing fault classification model under different working conditions is constructed, and the bearing fault classification model under different working conditions comprises: a backbone network which converts input signals into learnable tensors; a PDE multi-scale convolution module which respectively performs deep separable convolution on the learnable tensors based on a Gaussian kernel, a Laplacian kernel and four different direction Sobel kernels, constructs a PDE regularization term for total loss calculation and generates physical gate fusion features in combination with a physical gate mechanism; a physically driven Transformer which extracts physical quantities from the physical gate fusion features and designs a physical enhancement multi-head self-attention mechanism and a physical gate feedforward network to obtain final output features; a classifier which converts the final output features into predicted fault categories; a total loss is constructed to train the bearing fault classification model under different working conditions, and the trained bearing fault classification model under different working conditions is used as an optimal bearing fault classification model under different working conditions; The bearing driving end vibration acceleration signals are collected in real time, input into the optimal bearing fault classification model under different working conditions, and the predicted fault categories output by the bearing fault classification model under different working conditions are used as bearing fault classification results, so that bearing fault classification under different working conditions is realized.

2. The physical knowledge-driven transfer learning cross-condition bearing fault classification method according to claim 1, characterized in that: The bearing driving end vibration acceleration signals under different working conditions are collected by using acceleration vibration sensors arranged at multiple points, and original vibration sequences under multiple working conditions are constructed in time sequence by combining time indexes, and the method comprises the following steps: A plurality of measuring points are selected on the bearing seat surface, and acceleration vibration sensors are fixedly installed; The bearing is operated under different working conditions, the acceleration vibration sensors collect driving end vibration acceleration signals in normal states and in each fault state under each working condition, the sampling frequency is set to a fixed value, the collected driving end vibration acceleration signals are used as original vibration data, and original vibration sequences are constructed in time sequence.

3. The physical knowledge-driven transfer learning cross-condition bearing fault classification method according to claim 1, characterized in that: The source domain training set, the target domain verification set and the target domain test set are constructed in an unsupervised form, and the method comprises the following steps: In a sliding window clipping mode, each driving end vibration acceleration signal in the original vibration sequences from the source domain signal and the target domain signal is clipped into a plurality of signal sample units with a uniform length, so as to construct a source domain original sample library and a target domain original sample library; According to the time indexes in the original vibration sequences, source domain samples and target domain samples are sampled from the source domain original sample library and the target domain original sample library; According to the fault types, the source domain samples are assigned with fault category labels, and one-hot encoding is performed to obtain labeled source domain samples, so as to construct a source domain training set. The mean and the standard deviation are calculated based on the source domain samples, and the source domain samples and the target domain samples are standardized according to the mean and the standard deviation; The indexes of the standardized target domain samples are randomly shuffled, and are divided into a target domain verification set and a target domain test set in a 1:1 ratio.

4. The physical knowledge-driven transfer learning cross-working condition bearing fault classification method according to claim 3, characterized in that: The source domain samples and the target domain samples are respectively sampled from the source domain original sample library and the target domain original sample library according to the time index in the original vibration sequence, and the source domain samples and the target domain samples are standardized according to the mean and the standard deviation calculated based on the source domain samples. Set time segment threshold For time-based indexes with a total length of The source domain original sample library and the target domain original sample library constructed from the original vibration sequences: In the target domain original sample library, the first... Random selection within a portion of the time index Each signal sample unit serves as a source domain sample. This is a parameter controlling the number of samples; the remaining samples in the original sample library of the target domain. Random selection within a portion of the time index Each signal sample unit is used as a target domain sample.

5. The physical knowledge driven transfer learning cross-condition bearing fault classification method according to claim 1, characterized in that: The PDE multi-scale convolution module performs depth separable convolution on the learnable tensor based on a Gaussian kernel, a Laplacian kernel and four different direction Sobel kernels to generate Gaussian features, Laplacian features and Sobel features, and simultaneously calculates a PDE regularization term; The four different direction Sobel kernels are Sobel kernels of four gradient directions of 0°, 45°, 90° and 135°; The physical gating mechanism calculates energy values of the Gaussian features, the Laplacian features and the Sobel features respectively to convert into respective physical fusion weights, performs weighted fusion, and then performs element-by-element multiplication operation with the learnable tensor to generate physical gating fusion features; The physical gating mechanism calculates energy values of the Gaussian features, the Laplacian features and the Sobel features respectively to convert into respective physical fusion weights, performs weighted fusion, and then performs element-by-element multiplication operation with the learnable tensor to generate physical gating fusion features; The physical gating mechanism calculates energy values of the Gaussian features, the Laplacian features and the Sobel features respectively to convert into respective physical fusion weights, performs weighted fusion, and then performs element-by-element multiplication operation with the learnable tensor to generate physical gating fusion features; 6. The physical knowledge-driven transfer learning cross-condition bearing fault classification method according to claim 5, characterized in that: The physical gating mechanism calculates energy values of the Gaussian features, the Laplacian features and the Sobel features respectively to convert into respective physical fusion weights, performs weighted fusion, and then performs element-by-element multiplication operation with the learnable tensor to generate physical gating fusion features. The PDE multi-scale convolution module performs depth separable convolution on the learnable tensor based on a Gaussian kernel, a Laplacian kernel and four different direction Sobel kernels to generate Gaussian features, Laplacian features and Sobel features, and simultaneously calculates a PDE regularization term; Introducing a learnable diffusion tensor parameter to form a learnable dilated tensor matrix : ; wherein, , are learnable diffusion tensor parameters controlling the diffusion strength in the two principal directions, respectively, is a learnable diffusion tensor parameter controlling the coupling between the two principal directions. Utilizing learnable diffusion tensor parameters for each depth separable convolution kernel the following processing is performed: ; Generating learnable pde convolution kernels , the learnable pde convolution kernels include a learnable gaussian pde convolution kernel, a learnable laplacian pde convolution kernel, four learnable sobel-pde convolution kernels in different directions, respectively constituting a gaussian convolution layer, a laplacian convolution layer, four sobel convolution layers in different directions, and then respectively performing gaussian depth separable convolution, laplacian depth separable convolution, and four sobel depth separable convolutions in different directions on the learnable tensor to generate gaussian features, laplacian features, and four initial sobel features in different directions; The Gaussian depth separable convolution kernel, the Laplacian depth separable convolution kernel and the four different direction Sobel depth separable convolution kernels are formed by copying the Gaussian kernel, the Laplacian kernel and the four different direction Sobel kernels in the channel dimension to align with the feature channel number of the learnable tensor to constitute the basis kernel, and further introducing learnable kernel parameters; The four different direction initial Sobel features are linearly scaled and globally averaged pooled, and then are spliced and input into a fully connected layer to convert into weights corresponding to the four directions, and the four direction initial Sobel features are weighted fused based on the weights to obtain the Sobel features; The sum of squares of L2 norms of the base kernel corresponding to each depth separable convolution kernel and each learnable PDE convolution kernel is calculated as a PDE regularization term.

7. The physical knowledge-driven transfer learning cross-condition bearing fault classification method according to claim 5, characterized in that: The energy values of the high-frequency features, Laplacian features, and Sobel features are calculated respectively to convert into respective physical fusion weights, including: The absolute values of the feature values at each spatial position of the high-frequency features, Laplacian features, and Sobel features are averaged to obtain respective energy values, which are sent to a multi-layer perception to be converted into physical fusion weights corresponding to the high-frequency features, Laplacian features, and Sobel features through a softmax operation.

8. The physical knowledge driven transfer learning cross-condition bearing fault classification method according to claim 5, characterized in that: The physical gating fusion features are converted into sequence low-frequency energy, sequence gradient energy, and second-order curvature energy, including: The physical gating fusion feature is reshaped into an input feature sequence, and the sequence length is , and the element at the th position in the sequence is denoted as ; Calculating sequence low frequency energy based on average amplitude : ; wherein is the absolute value symbol; Building sequence gradient energy based on varying speed : ; ; wherein, represents a change in the first input feature sequence to the second input feature sequence; represents a change in the first input feature sequence to the second input feature sequence;​ Constructing second order curvature energy based on degree of mutation : ; 。 9. The physical knowledge-driven transfer learning cross-condition bearing fault classification method according to claim 1, characterized in that: The total loss is constructed to train the cross-condition bearing fault classification model, including: The final output features extracted by the cross-condition bearing fault classification model from the source domain training set are first reduced in dimension by a projection head, and then normalized so that all features are located on a unit sphere, and then a similarity matrix is constructed to calculate an InfoNCE contrast loss as the source domain contrast learning loss; The sequence low-frequency energy difference, sequence gradient energy difference, and second-order curvature energy difference between the source domain training set and the target domain validation set are calculated by the cross-condition bearing fault classification model, and the three differences are combined to construct a physical difference quantity, which is input into a multi-layer perception to be converted into a physical weight; The final output features from the source domain training set and the target domain validation set are normalized and a Gaussian kernel is selected, and then the Gaussian kernel MMD of the two is calculated, and the discrete estimation of the square of the obtained Gaussian kernel MMD is taken as the physical adaptive MMD domain alignment loss; The physical weight is applied to the physical adaptive MMD domain alignment loss, and combined with the source domain contrast learning loss as the total transfer loss; The cross-entropy loss between the predicted fault class generated by the cross-condition bearing fault classification model for the source domain training set and the corresponding fault class label is calculated as a classification loss. The weighted sum of the classification loss, the total transfer loss, and the PDE regularization term is taken as the total loss.

10. The physical knowledge driven transfer learning cross-condition bearing fault classification method according to claim 1, characterized in that: The classifier includes a multi-layer perception based Gaussian expert path, Laplacian expert path, and Sobel expert path, which converts the final output features from the physical driven Transformer into Gaussian classification probability, Laplacian classification probability, and Sobel classification probability, and extracts physical quantities from the final output features, which are converted into gating expert weights corresponding to the Gaussian expert path, Laplacian expert path, and Sobel expert path by a gating network, and the Gaussian classification probability, Laplacian classification probability, and Sobel classification probability are weighted and fused to obtain a bearing fault classification probability and convert it into a predicted fault class.

Citation Information

Patent Citations

  • Intelligent aero-engine bearing fault migration diagnosis method based on multi-source data and attention mechanism

    CN120654092A

  • Method and system for fault diagnosis of rolling bearing

    GB202302142D0