Small sample bearing fault diagnosis method based on double-feature space fusion

By using a dual-feature space fusion method, bearing fault features are extracted using a one-dimensional multi-scale dynamic network and a two-dimensional RPE-ViT network. This solves the problem of insufficient accuracy and generalization ability in bearing fault diagnosis under small sample conditions, and achieves efficient fault identification and classification.

CN121030516APending Publication Date: 2025-11-28LANZHOU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511128712.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In bearing fault diagnosis under small sample conditions, existing technologies have poor fault tolerance in single-channel feature space and are easily affected by external interference. Traditional machine learning methods are sensitive to noise, and deep learning models have difficulty capturing the dependencies of long sequence signals, resulting in insufficient diagnostic accuracy and generalization ability.

Method used

A dual-feature space fusion method is adopted, which extracts local fault features through a one-dimensional multi-scale dynamic network and extracts globally dependent fault features through a two-dimensional RPE-ViT network, and then performs feature fusion to finally classify faults.

Benefits of technology

It achieves high-accuracy fault diagnosis under small sample conditions, with fewer parameters and faster diagnosis speed, and can effectively solve the problems of poor fault tolerance and susceptibility to external interference in single-channel feature space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121030516A_ABST
    Figure CN121030516A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample bearing fault diagnosis method based on double feature space fusion, and the method comprises the steps: firstly obtaining a one-dimensional vibration signal of a bearing, and converting the one-dimensional vibration signal into a multi-window spectrum frequency domain signal and a time frequency domain signal; inputting the multi-window spectrum frequency domain signal into a one-dimensional multi-scale dynamic network to extract local fault features, and inputting the time frequency domain signal into a two-dimensional RPE-ViT network to extract global dependency fault features; and finally, fusing the local fault features and the global dependency fault features, and classifying the bearing faults according to the fused features. According to the method, high fault diagnosis accuracy can be achieved under the condition of small samples, and meanwhile, the method also has less parameter quantity and higher diagnosis speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bearing fault diagnosis, and more particularly to a small sample bearing fault diagnosis method based on double feature space fusion. BACKGROUND

[0002] As the core load-bearing component of rotating machinery, rolling bearings are widely used in key industrial fields such as aerospace, energy and power, and intelligent manufacturing, and their performance directly affects the operation reliability and production efficiency of equipment. Building an efficient and accurate bearing fault diagnosis system is not only a key technical link to ensure the continuous and stable operation of rotating machinery, but also an important prerequisite for promoting industrial intelligent transformation and achieving safe production.

[0003] Currently, traditional methods mainly use machine learning algorithms such as support vector machines (SVM), decision trees (DT), and random forests (RF) to learn from complex vibration data and establish a mapping relationship between fault features and states, thereby achieving relatively reliable fault recognition.

[0004] However, SVM achieves classification by constructing an optimal hyperplane, but its diagnostic performance is highly dependent on the selection of kernel functions, and the adaptability of different kernel functions such as Gaussian kernel and polynomial kernel requires repeated debugging based on expert experience. DT builds a classification model based on feature splitting, and RF improves stability by integrating multiple decision trees, but both are sensitive to outliers and noise in the data, and under the influence of complex electromagnetic interference, environmental noise, and other factors in industrial sites, they are prone to model overfitting or feature misjudgment. In addition, such methods are generally limited by shallow network structures, making it difficult to mine deep feature relationships in nonlinear vibration signals, such as weak impact signals in early bearing faults, modulation phenomena of fault feature frequencies, and other complex information, resulting in insufficient accuracy and generalization ability of fault diagnosis.

[0005] To solve the above problems, deep learning models are widely used in bearing fault diagnosis, achieving high-precision classification of bearings under varying operating conditions; however, for long sequence data, the limited perception ability of CNN often leads to information loss or ambiguity. In addition, due to the local connection and parameter sharing characteristics of CNN, it is difficult to capture dependencies in long sequence signals. Furthermore, in industrial practice, bearings are often replaced before or just after failure, resulting in a lack of failure samples. When the sample size is too small, complex networks have difficulty extracting rich features and are prone to overfitting.

[0006] Therefore, how to solve the problems of poor fault tolerance and susceptibility to external interference in single-channel feature space under small sample conditions is a problem that needs to be solved by those skilled in the art. SUMMARY

[0007] Therefore, in order to at least partially solve the above problems in the prior art, the present application provides a small sample bearing fault diagnosis method based on double feature space fusion.

[0008] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0009] A small sample bearing fault diagnosis method based on double feature space fusion, the steps comprising,

[0010] Obtaining a one-dimensional vibration signal of a bearing, and converting the one-dimensional vibration signal into a multi-window spectrum frequency domain signal and a time-frequency domain signal, respectively;

[0011] Inputting the multi-window spectrum frequency domain signal into a one-dimensional multi-scale dynamic network to extract local fault features, and inputting the time-frequency domain signal into a two-dimensional RPE-ViT network to extract global dependent fault features;

[0012] Fusing the local fault features and the global dependent fault features, and classifying faults according to the fused features.

[0013] As a preferred, converting the one-dimensional vibration signal of the bearing into the multi-window spectrum frequency domain signal comprises:

[0014] Sampling the one-dimensional vibration signal of the bearing to obtain a discrete time domain vibration signal;

[0015] Defining K window functions, and weighting the time domain vibration signal of each discrete point using the kth window function, k = 1, 2,..., K;

[0016] Performing fast Fourier transform on the obtained weighted signal to obtain a frequency domain signal, and calculating a single window power spectrum estimate according to the frequency domain signal;

[0017] Weighted averaging the K single window power spectrum estimates to obtain a multi-window spectrum frequency domain signal.

[0018] As a preferred, the K window functions satisfy the following conditions:

[0019]

[0020] In the formula, w i (t) represents the ith window function, w j (t) represents the jth window function, δ ij represents a Kronecker function, and i and j represent the serial number index of the window function.

[0021] As a preferred, the fast Fourier transform on the obtained weighted signal comprises:

[0022]

[0023] In the formula, f mrepresents the mth frequency point, f m = m·f s / N, N represents the signal length, n represents the nth discrete time domain vibration signal, x(n) represents the weighted signal, f s is the sampling frequency.

[0024] As preferred, the bearing one-dimensional vibration signal is converted into a time-frequency domain signal, comprising:

[0025] The bearing one-dimensional vibration signal is normalized;

[0026] The normalized data points are mapped to a polar coordinate system, and the angles corresponding to each data point are calculated;

[0027] A symmetric Gram matrix is constructed based on the angles;

[0028] The matrix element values are mapped to the gray values or colors of the image pixels, so as to visualize the Gram matrix into an image.

[0029] As preferred, the angle corresponding to each data point is calculated by an inverse cosine function,

[0030]

[0031] In the formula, represents the normalized data points.

[0032] As preferred, the calculation method of the elements in the Gram matrix is:

[0033]

[0034] In the formula, GASF ts is the element of the tth row and s th column in the Gram matrix, representing the relationship between the data points corresponding to time points t and s.

[0035] As preferred, the one-dimensional multi-scale dynamic network comprises:

[0036] A dynamic weight generator is configured to perform global average pooling on the input data, and sequentially generate a weight vector through a first connection layer, a Relu activation function, a second connection layer and a Softmax function, wherein the dimension of the weight vector is the same as the number of multi-scale convolution branches;

[0037] A multi-scale convolution branch is configured to independently convolve the input data using different convolution kernel sizes to generate corresponding feature maps;

[0038] A dynamic weighted fusion is configured to weight the weight vector and the output of each multi-scale convolution branch, output a dynamic convolution kernel through normalization and an activation function, and multiply the dynamic convolution kernel output by each branch with the original input data to obtain respective fusion features.

[0039] Cascade residual fusion, sequentially passing the respective fusion features through a Leaky Relu function, and fusing in a cascade residual manner into a final output.

[0040] As a preference, the two-dimensional RPE-ViT network comprises an encoder and a decoder;

[0041] The encoder comprises a multi-head self-attention module, a layer normalization module, and a multi-layer perception module;

[0042] The decoder is identical in structure to the encoder, and a masked multi-head attention module is added before the multi-head self-attention module.

[0043] As a preference, the two-dimensional RPE-ViT network is used to extract global dependent fault features, comprising:

[0044] The time-frequency domain signal is divided into multiple patches, and an embedding vector of each patch is generated;

[0045] A relative position encoding is added to each embedding vector;

[0046] The encoded feature vector is input into the two-dimensional RPE-ViT network for feature extraction.

[0047] The small sample bearing fault diagnosis method based on double feature space fusion provided by the present application can achieve high fault diagnosis accuracy under small sample conditions through double channel feature extraction, can solve the problems of poor fault tolerance effect and easy interference of single channel feature space, has less parameter quantity and faster diagnosis speed, and can realize efficient bearing fault diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0049] Figure 1 A multi-scale dynamic network structure diagram is provided for the present application.

[0050] Figure 2 A PE-ViT structure diagram is provided for the present application.

[0051] Figure 3 A fault diagnosis flowchart is provided for the present application.

[0052] Figure 4The accuracy and loss value result chart of the test set provided by the present application.

[0053] Figure 5 The TSNE visualization result chart provided by the present application.

[0054] Figure 6 The sample comparison result chart provided by the present application, wherein (a) is a comparison result chart under different training samples, and (b) is a sample comparison result chart of different methods. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0056] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details that are set forth in the following description, in other manners different from those described herein, and it can be apparent to those skilled in the art that the present application is not limited to the specific embodiments disclosed herein and can be practiced with or without the other implementations of the present application.

[0057] To solve the problems in the prior art, the embodiments of the present application disclose a small sample bearing fault diagnosis method based on double feature space fusion, which comprises the following steps:

[0058] Obtain a one-dimensional vibration signal of a bearing, and convert the one-dimensional vibration signal into a multi-window spectrum frequency domain signal and a time-frequency domain signal, respectively;

[0059] Input the multi-window spectrum frequency domain signal into a one-dimensional multi-scale dynamic network to extract local fault features, and input the time-frequency domain signal into a two-dimensional RPE-ViT network to extract global dependent fault features.

[0060] Fuse the local fault features and the global dependent fault features, and perform fault classification according to the fused features.

[0061] The fault diagnosis method provided by the present application can achieve a high fault diagnosis accuracy under a small sample condition, and has a small parameter amount and a fast diagnosis speed.

[0062] In one embodiment, after obtaining a one-dimensional vibration signal of a bearing, the one-dimensional vibration signal is preprocessed and converted into a multi-window spectrum frequency domain signal and a time-frequency domain signal.

[0063] In the present embodiment, the preprocessing refers to normalizing the one-dimensional vibration signal, and the expression is:

[0064]

[0065] wherein X represents the normalized one-dimensional vibration signal; X max represents the maximum value of the one-dimensional vibration signal in the diagnostic model input data set; X min represents the minimum value of the one-dimensional vibration signal in the diagnostic model input data set. Further,

[0066] 1. In this embodiment, the IMTSAFFT technology is used to perform frequency domain analysis and processing on the received signal. The IMTSAFFT performs weighted processing on the signal through multiple orthogonal window functions, reduces the variance of the spectrum estimation by statistical averaging, clearly separates different frequency components in the signal, identifies the frequency bands where the noise and effective signal are located, and enhances the feature expression ability of the original signal.

[0067] Specifically, the bearing one-dimensional vibration signal is converted into a multi-window spectrum frequency domain signal, and the steps include:

[0068] (1) Sampling the bearing one-dimensional vibration signal x(t) (t is time) to obtain a discrete time domain vibration signal x(n), n represents a discrete time point; the sampling process is represented as: x(n) = x(nT s ), T s represents the sampling period;

[0069] (2) Defining K window functions as w1(t), w2(t), …, w k (t), and the K window functions satisfy the following conditions:

[0070]

[0071] wherein w i (t) represents the i-th window function, w j (t) represents the j-th window function, δ ij represents the Kronecker function, and i and j represent the serial number index of the window function.

[0072] (3) Weighting the time domain vibration signal x(n) of each discrete point using the k-th window function, k = 1, 2, …, K;

[0073] x k (n) = x(n) · w k (n), k = 1, 2, …, K

[0074] (4) Performing fast Fourier transform on the obtained weighted signal to obtain a frequency domain signal, i.e.:

[0075]

[0076] wherein f mrepresents the mth frequency point, f m = m f s / N, N represents the signal length, n represents the nth discrete time domain vibration signal, x(n) represents the weighted signal, f s is the sampling frequency, and j represents the imaginary unit.

[0077] According to the frequency domain signal, a single window power spectrum estimation is calculated;

[0078]

[0079] (5) Weighted average of K single window power spectrum estimations to obtain a multi-window spectrum frequency domain signal.

[0080]

[0081] Wherein the weight coefficient is taken as 1 / K.

[0082] 2, the embodiment uses GASF transform technology to deeply characterize the original signal, ingeniously encodes one-dimensional time series data into two-dimensional polar coordinate image, not only losslessly retains the inherent time dynamic and amplitude intensity information of the signal, but also highlights the hidden feature mode through geometric structure, provides a more information-intensive and more distinctive visual carrier for subsequent analysis.

[0083] Specifically, the one-dimensional vibration signal of the bearing is converted into a time-frequency domain signal, and the steps include:

[0084] The one-dimensional vibration signal of the bearing is normalized and scaled to the range (-1, 1); assuming that a given one-dimensional fault signal sequence is X = [x1, x2,... x n ], then the sequence after normalization is

[0085] The normalized data points are mapped to the polar coordinate system, and the angles corresponding to each data point are calculated by the inverse cosine function; represents the normalized data points.

[0086] Based on the angle, a symmetric Gram matrix G is constructed to represent the correlation between different time points, and the Gram matrix G is an N*N symmetric matrix; the calculation method is:

[0087]

[0088] In the formula, GASF ts is the element of the tth row and s th column of the Gram matrix, which represents the relationship between the data points corresponding to the time points t and s.

[0089] From the above, the complete formula of the GASF matrix can be obtained as:

[0090]

[0091] Finally, the Gram matrix G is visualized as an image, each element in the matrix corresponds to a pixel point in the image, and the color or gray value of the pixel point is determined by the numerical value of the element.

[0092] In an embodiment, the multi-window spectral frequency domain signal is input into a one-dimensional multi-scale dynamic network to extract local fault features, and the time-frequency domain signal is input into a two-dimensional RPE-ViT network to extract global dependent fault features.

[0093] 1) In this embodiment, a multi-scale dynamic network is set up to make up for the problem of insufficient local feature extraction in RPE-ViT in fault diagnosis.

[0094] In this application, the one-dimensional multi-scale dynamic network processes one-dimensional frequency domain signals and includes dynamic convolution modules and feature fusion modules connected in turn; its structure is as shown in Figure 1 , wherein,

[0095] The dynamic convolution module uses one-dimensional convolution for feature extraction, and the convolution kernel size is set to 3*1, 5*1 and 7*1. Different sizes of convolution kernels have their own functions. Small size convolution kernel captures high frequency detail features, such as impact frequency generated by local damage of bearing; medium size convolution kernel extracts medium frequency periodic features, corresponding to fault feature frequency of bearing elements; large size convolution kernel obtains low frequency trend features, reflecting slow changes of equipment running state;

[0096] Specifically, the dynamic convolution module includes three sub-branches, each branch including a dynamic weight generator and a convolution output. The process of the dynamic weight generator part is as follows:

[0097] First, global average pooling is performed.

[0098]

[0099] In the formula: L is the sequence length, C is the number of channels, x in is the input data, x pool is the output after global pooling.

[0100] Second, the convolution kernel weight π is output.

[0101] π=Softmax{FC2[Relu(FC1(x pool ))]}

[0102] Finally, k convolution output features are applied to the input x in , weighted fusion is performed using π, and BN and activation output are performed.

[0103]

[0104] Branch 1 has a convolution kernel size of 3, a dynamic convolution kernel of W1, and the output of the convolution is y1 = x * W1.

[0105] Branch 2 has a convolution kernel size of 5, a dynamic convolution kernel of W2, and the output of the convolution is y2 = x * W2.

[0106] Branch 3 has a convolution kernel size of 7, a dynamic convolution kernel of W3, and the output of the convolution is y3 = x * W3.

[0107] Furthermore, the feature fusion module has the following formula:

[0108] y = S{L[y3(L[y2(L[y1(x)])])]}

[0109] In the formula: L represents the Leaky ReLU operation, and S represents the cascaded residual fusion operation.

[0110] 2) In this embodiment, the two-dimensional RPE-ViT network structure is as follows: Figure 2 As shown, the feature extraction process includes:

[0111] The time-frequency domain signal is divided into N patches, and an embedding vector is generated for each patch. i∈{1,2,...,n};

[0112] Add relative position encoding to each embedding vector; this includes defining a relative position encoding matrix. Where M is the maximum relative distance, and d is the embedding dimension; for any two patches i and j, their relative distance is k = ji, and the corresponding relative position encoding is...

[0113] The encoded feature vector is then input into the two-dimensional RPE-ViT network for feature extraction.

[0114] In this embodiment, the two-dimensional RPE-ViT network includes an encoder and a decoder;

[0115] The encoder includes a multi-head self-attention module, a layer normalization module, and a multilayer perceptron module;

[0116] The decoder has the same structure as the encoder, and a masked multi-head attention module is added before the multi-head self-attention module.

[0117] The multi-head self-attention module serves as the core computational unit. It constructs the correlation matrix between any two elements in the input signal sequence by computing multiple attention heads in parallel: MultiHead(Q,K,V)=Concat(head1,…,head)h W o Each attention head passes through the head. i =Attention(QW i Q ,KW i K VW i V This mechanism enables the capture of feature dependencies at different scales. For example, in bearing fault diagnosis, it can effectively correlate periodic impacts in time-domain signals with characteristic frequencies in the frequency domain.

[0118] The multilayer perceptron module undertakes the task of feature nonlinear transformation, and its structure can be represented as follows:

[0119] MLP(x)=Dropout(W2·GELU(W1x+b1)+b2)

[0120] The network employs a two-layer fully connected network to perform dimensionality increases and decreases in the feature space. The inserted GELU activation function introduces nonlinearity through x·Φ(x) (Φ(x) being the cumulative distribution function of the standard normal distribution), which better preserves gradient continuity compared to ReLU. Experiments show that this improves network convergence speed by 15%-20%. The Dropout function randomly zeros out some neurons with a certain probability, effectively suppressing overfitting and reducing the false detection rate by approximately 8% on the bearing fault dataset. This synergistic architecture of attention mechanism and nonlinear transformation enables RPE-ViT to demonstrate superior performance compared to traditional multi-scale dynamic networks in long-distance dependency modeling of bearing fault features.

[0121] In one embodiment, the local fault features and the globally dependent fault features are fused, and the fault is classified by the Softmax function based on the fused features.

[0122] In this invention, feature extraction focuses on the model's ability to capture local features of fault data, thereby innovatively constructing a multi-scale dynamic network architecture for the local perception module. By fully utilizing its weight sharing and local connectivity characteristics, it achieves efficient extraction of local detailed features from fault data.

[0123] The weight-sharing mechanism unique to CNNs enables multi-scale dynamic networks to reuse the same convolution kernel parameters when processing similar features at different locations, significantly reducing the number of model parameters while ensuring the consistency of feature extraction. In addition, the local connectivity of CNNs gives the network the ability to focus on local regions of data, enabling it to accurately identify subtle texture changes and geometric features of faults such as cracks on bearing surfaces and wear on gear teeth.

[0124] For example, in the scenario of rotating machinery fault diagnosis, abnormal vibration signals before equipment failure often exhibit short-term and high-frequency characteristics. The time attention mechanism can automatically enhance the attention to such key time period data, preventing useful information from being submerged in a large amount of normal data.

[0125] Furthermore, the multi-scale design expands the feature extraction dimensions of the network. By setting convolutional kernels of different sizes, features are extracted from fault data at multiple granularities. Small-sized convolutional kernels excel at capturing fine local details, while large-sized kernels can acquire more macroscopic structural information. The combination of the two allows the network to capture both the microscopic features of fault points and understand their positional relationship within the overall structure. This organic integration of multi-scale and temporal attention mechanisms with a multi-scale dynamic network infrastructure effectively fills the gap in local perception of the RPE-ViT model, significantly enhancing the robustness and accuracy of the fault diagnosis model and providing more reliable technical support for the intelligent operation and maintenance of industrial equipment.

[0126] The second part fully leverages the advantages of the RPE-ViT model to construct a global long-distance feature dependency modeling system with a multi-head self-attention mechanism at its core.

[0127] As a core component of ViT, the multi-head self-attention mechanism breaks through the limitations of traditional convolutional neural networks in capturing long-distance dependencies. It can calculate the correlation weights between various positions in the input sequence in parallel from different dimensions. Through the collaborative work of multiple "heads," it comprehensively captures feature dependencies at different scales and levels in fault data. For example, in the fault diagnosis scenario of complex electromechanical systems, the potential correlations between different types of data such as equipment vibration signals, temperature changes, and current fluctuations can all be analyzed through the multi-head self-attention mechanism to achieve deep correlation analysis across channels and time, effectively uncovering fault feature patterns hidden behind the data.

[0128] The RPE-ViT model simplifies the parameter scale of the ViT model by rationally designing the number of network layers and neurons in the MLP. The MLP, through nonlinear transformation, compresses model complexity while preserving the expressive power of key features, achieving a lightweight model. Experiments have verified that the optimized model significantly improves inference speed while maintaining diagnostic accuracy, making it more suitable for real-time online fault diagnosis scenarios.

[0129] Furthermore, in the fault classification stage, a carefully designed classification head serves as the "decision center" of the model's output layer. It transforms the feature vectors processed by the multi-head self-attention mechanism and MLP into specific fault category prediction results. The classification head adopts a classic architecture combining fully connected layers with the Softmax activation function. Through training on a large amount of labeled data, it learns the feature mapping relationships of different fault types, enabling it to accurately identify various states of the equipment, such as normal operation, minor faults, and severe faults.

[0130] In a preferred embodiment, after acquiring a one-dimensional vibration signal, performing preprocessing and conversion, it is divided into a training sample set, a validation sample set, and a test sample set; the fault diagnosis model is trained using the training sample set to obtain the optimal model, and the training process is as follows: Figure 3 As shown.

[0131] To enhance the model's ability to acquire feature information, this embodiment divides the original preprocessed and transformed data into slices. Preferably, the displacement and length of two adjacent slice samples are set to 512 and 1024, respectively. 700 samples are taken to form the training dataset and 100 samples are taken to form the test dataset, that is, the training dataset and the test dataset are divided in a 7:1 ratio.

[0132] During the model training process, this embodiment uses the Adam optimizer for parameter updates, specifically including:

[0133] Initialize the Adam optimizer;

[0134] Training samples are input into the model for fault feature learning and diagnosis. During the learning process, the Adam optimizer adjusts the learning rate for each training iteration, continuously updating the loss value and accuracy. Specifically, the Adam optimizer needs to determine if the model has converged during optimization; that is, if the diagnostic accuracy stabilizes as the number of iterations of the dual-branch bearing fault diagnosis model based on multi-scale dynamic networks and ViT increases, the model is considered to have converged. At this point, training is terminated, and the optimal diagnostic model parameters are saved.

[0135] To further improve classification accuracy, a cross-entropy loss function is introduced as the optimization objective to enhance the model's ability to distinguish easily confused fault categories. For example, when distinguishing between fault types with high feature similarity, such as bearing outer race faults and inner race faults, the classification head can enhance its sensitivity to subtle feature differences through optimized training, thereby achieving high-precision fault classification and diagnosis, and providing a reliable basis for equipment maintenance decisions.

[0136] The cross-entropy loss function obtains the average cross-entropy loss over the entire training set by summing or averaging the cross-entropy losses of each sample.

[0137] The cross-entropy loss function makes the model focus more on samples that are difficult to classify, thereby improving the model's classification and generalization abilities. The specific expression is shown below:

[0138]

[0139] Where: y i It is the one-hot encoding of the true label of the i-th sample; n is the number of samples in the training set. It is the probability of the i-th type predicted by the model.

[0140] To verify the effectiveness of the method provided in the embodiments of the present invention in terms of fault diagnosis performance, two specific embodiments are presented below. Furthermore, fault diagnosis experiments were conducted in a small sample environment using PyTorch-python in these two specific embodiments.

[0141] Example 1:

[0142] 1.1 Dataset Description:

[0143] The CWRU dataset contains four different fault states, specifically divided as follows: Dataset A: 1797 r / min, 0 HP load; Dataset B: 1772 r / min, 1 HP load; Dataset C: 1750 r / min, 2 HP load; Dataset D: 1730 r / min, 3 HP load. This dataset simulates the damage conditions that occur in actual bearing operation. The damage locations include the rolling elements (Ball), the inner ring, and the outer ring at the 6 o'clock position (Outer@6), with damage diameters of 0.007, 0.014, and 0.021 inches, respectively. Based on the bearing damage state, the dataset can be divided into 10 states, with specific labels as shown in Table 1, including normal states and fault states with different damage degrees. In this experiment, the collected data will be divided into training and testing samples at a ratio of 9:1 to ensure the effectiveness of model training and evaluation. The displacement and length of two adjacent slice samples are set to 512 and 1024, respectively.

[0144] Table 1. CWRU Dataset Label Assignment Table

[0145]

[0146] 1.2 Model structural parameters

[0147] This invention was conducted on a computer with a Windows 10 system, an i5-13400F processor, and 16GB of RAM. Python was used as the development tool, and TyTorch was used as the deep learning framework. The key parameters set in the experiment are as follows: the learning rate was 0.001; to suppress overfitting, the dropout value of the Dropout layer was set to 0.5; the SGD optimization algorithm was used to update the network training parameters in real time; the number of iteration batches was set to 50, the number of sample batches was set to 64, the patch size was set to 64, the encoder was set to 6, the MLP hidden size was set to 128, and the number of attention heads was set to 8.

[0148] 1.3 Comparison Methods

[0149] To verify whether this invention has good fault diagnosis performance in experiments with varying noise and operating conditions, it will be compared with four existing excellent methods: the multi-scale dynamic network method, the ResNet18 method, the LSTM-multi-scale dynamic network method, and the CWT-ViT method. The multi-scale dynamic network method consists of an input layer, a wide convolutional layer, a fully connected layer, and a classification layer. The ResNet18 method consists of an input layer, residual blocks, pooling layers, a fully connected layer, and a classification layer. The LSTM-multi-scale dynamic network method has a similar structure to this invention, consisting of a dual-branch LSTM and a multi-scale dynamic network. The CWT-ViT method consists of CWT data transformation technology and a ViT network, with the classification head in ViT completing the fault type classification.

[0150] 1.4 Fault Diagnosis Results and Analysis under Normal Conditions

[0151] Figure 4 Figures (a) and (b) show the fault diagnosis accuracy curve and loss value change curve after 50 iterations of the present invention. It can be seen that the fault diagnosis model has excellent convergence speed and tends to stabilize within a small number of batches. At the same time, the accuracy of both the training set and the test set eventually reaches 100%, and the loss value drops to below 0.005. It can be concluded that the bearing fault diagnosis model has a significant effect on feature extraction, and requires fewer training batches, shorter training time, and higher diagnostic accuracy.

[0152] To further verify the fault classification effectiveness of this invention, this application employs T-SNE dimensionality reduction technology to visualize the classification results, as shown below. Figure 5 As shown, 'a' corresponds to the original data, and 'b' corresponds to the classification result. After feature extraction and fault classification processing according to this invention, the originally chaotic and difficult-to-distinguish fault data becomes highly clustered of similar data, completely separated from dissimilar data, and the spatial location of each data point is clearly discernible. This result demonstrates that the model proposed in this paper possesses powerful fault identification and classification capabilities.

[0153] The final results of the four performance metrics of the model in this paper are shown in Table 2. As can be seen from the table, the values ​​of the four performance metrics are all 99.83% or above, indicating that the model can accurately identify the positive class and correctly capture the positive class samples when processing a given task, while reducing false alarms.

[0154] Table 2 Description of Fault Types in the Dataset

[0155]

[0156]

[0157] 1.5 Results and Analysis of Small Sample Experiments

[0158] While most methods achieve high diagnostic accuracy with sufficient fault samples, their learning performance is limited when the number of fault samples is limited due to insufficient feature extraction capabilities. In actual industrial production, bearing failures occur relatively infrequently, resulting in a limited number of available fault samples. This necessitates that fault diagnosis methods maximize feature extraction capabilities under limited training sample conditions. To verify the effectiveness of the proposed method in fault diagnosis with a small number of samples, we constructed training sets containing 10, 13, 15, 20, and 25 samples, and a test set containing 5 samples, for each fault type. We randomly selected the required number of samples from the collected 800 samples to form the training and test sets. To reduce bias from random sampling, we repeated the experiment five times and used the average as the final result. The specific data are shown in Table 3.

[0159] Table 3 Results of small sample experiments

[0160]

[0161] As shown in the table, when the number of samples is 10, the average fault diagnosis rates of multi-scale dynamic networks, ResNet-18, LSTM-multi-scale dynamic networks, and CWT-ViT are 81.05%, 91.21%, 97.21%, and 91.73%, respectively. The fault diagnosis accuracy of this invention is 98.64%, representing improvements of 17.59%, 7.43%, 1.43%, and 6.91% compared to the other four methods. The accuracy of each method improves with increasing sample size, demonstrating the sensitivity of fault diagnosis methods to sample size. When the sample size increases to 20, the fault diagnosis accuracy of this invention reaches 100%, indicating that it can correctly identify all fault categories at this sample size. However, the multi-scale dynamic network method still has a false diagnosis rate of 11.15%, while the LSTM-multi-scale dynamic network method has a diagnosis accuracy of 98.86%, which is closest to the performance of this invention. This indicates that the dual-branch architecture is highly advantageous in addressing the problem of poor fault diagnosis performance caused by small sample sizes.

[0162] Example 2:

[0163] 2.1 Gearbox Dataset Description:

[0164] To further verify the effectiveness of the method proposed in this invention, bearing data from the Southeast University gearbox dataset were also used for verification. This dataset contains bearing vibration signals under four different operating conditions: normal, inner ring fault, rolling element fault, and combined inner and outer ring fault. A 20Hz-0V operating condition was set, and data from eight channels were collected; this experiment used two channels of data.

[0165] The experiment selected 70 samples from each fault type as the training set and 10 samples from each fault type as the test set, dividing them into training and test sets in a 7:1 ratio. Details of the dataset allocation are shown in Table 4.

[0166] Table 4. Gearbox Dataset Label Assignment Table

[0167] Label Damage location Sampling frequency / Hz Training set Test set 0 Healthy 5120 70 10 1 Rolling element 5120 70 10 2 Inner ring 5120 70 10 3 Mixed fault 5120 70 10

[0168] 2.2 Fault Diagnosis Results and Analysis under Different Sample Conditions

[0169] The only difference between this experiment and the previous one is the dataset; all other network model parameters are the same.

[0170] To verify the effectiveness and generalization performance of this invention under small sample conditions in the SEU dataset, experiments will continue to be conducted with 10, 13, 15, 20, and 25 samples per class, and the results are as follows. Figure 6As shown in (a), it can be seen that the accuracy of the invention increases with the increase of the sample size, but the standard deviation decreases, indicating the sensitivity of the invention to data. The accuracy is 99.17% and 99.20% when the sample size is 15 and 20, respectively.

[0171] To verify the superiority of this invention, a comparative experiment was conducted with a comparative method, and the results are as follows: Figure 6 As shown in (b), when the sample size is 10, the accuracy rates of the comparison methods are 78.13%, 82.76%, 74.28%, and 81.89%, respectively, while the accuracy rate of the present invention is 94.17%. This demonstrates that the method can still complete the fault diagnosis task even in noisy environments with a small sample size.

[0172] 2.3 Noise Resistance Performance Experiment Results and Analysis

[0173] To further verify the fault diagnosis capability of this invention in a high-noise environment, this section uses the Southeast University dataset to test the model. -4, -2, 0, 2, and 4 dB were added to the original vibration signal. The model performance index values ​​are shown in Table 5.

[0174] Table 5 Noise Resistance Performance of Gearbox Dataset

[0175] Signal-to-noise ratio / dB Accuracy Precision Recall F1-score -4 87.75% 88.82% 87.75% 87.88% -2 96.50% 96.57% 96.50% 96.47% 0 98.00% 98.06% 98.00% 98.00% 2 99.75% 99.75% 99.75% 99.75% 4 99.50% 99.50% 99.50% 99.50%

[0176] As shown in the table, the performance indicators of the model proposed in this paper can still maintain a high fault diagnosis rate under strong noise environment. When the noise intensity is less than -2dB, the fault diagnosis accuracy, precision, recall, and F1-score are all above 96.47%. When the noise intensity reaches 4dB, all four performance indicators are above 99%, indicating that the model has strong noise resistance.

[0177] In summary, this invention proposes a novel bearing fault diagnosis method. It can solve the problems of low model fault diagnosis effectiveness and poor generalization ability caused by strong noise interference and large variations in operating conditions in rolling bearings. The main conclusions are as follows.

[0178] (1) A dual-branch network architecture was designed, which can enhance the complementarity and fault tolerance of network features.

[0179] (2) A multi-scale dynamic network was designed to compensate for the lack of local feature extraction capability of the RPE-ViT model.

[0180] (3) Multiple experiments were conducted using the CWRU dataset and the Southeast University bearing dataset. The final experimental results show that the present invention not only has a good fault diagnosis effect in small sample environments, but also has strong generalization performance.

[0181] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0182] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A small sample bearing fault diagnosis method based on double feature space fusion, characterized by the steps of The method comprises the following steps: obtaining a one-dimensional vibration signal of a bearing, and converting the one-dimensional vibration signal into a multi-window spectrum frequency domain signal and a time-frequency domain signal respectively; inputting the multi-window spectrum frequency domain signal into a one-dimensional multi-scale dynamic network to extract local fault features, and inputting the time-frequency domain signal into a two-dimensional RPE-ViT network to extract global dependent fault features; fusing the local fault features and the global dependent fault features, and classifying bearing faults according to the fused features.

2. The small sample bearing fault diagnosis method according to claim 1, characterized in that, The method for converting the one-dimensional vibration signal of the bearing into the multi-window spectrum frequency domain signal comprises the following steps: sampling the one-dimensional vibration signal of the bearing to obtain a discrete time domain vibration signal; defining K window functions, and weighting the time domain vibration signal of each discrete point by using the kth window function, wherein k = 1, 2, …, K; performing fast Fourier transform on the obtained weighted signal to obtain a frequency domain signal, and calculating a single-window power spectrum estimate according to the frequency domain signal; performing weighted averaging on the K single-window power spectrum estimates to obtain a multi-window spectrum frequency domain signal.

3. The small sample bearing fault diagnosis method of claim 2, wherein, The K window functions satisfy the following conditions: where w i (t) represents the i-th window function, w j (t) represents the j-th window function, δ ij represents the Kronecker function, and i and j represent the ordinal index of the window function.

4. The small sample bearing fault diagnosis method of claim 2, wherein, The method for performing fast Fourier transform on the obtained weighted signal comprises the following steps: where f m represents the mth frequency point, f m = m·f s / N, N represents the signal length, n represents the nth discrete time-domain vibration signal, x(n) represents the weighted signal, f s is the sampling frequency.

5. The small sample bearing fault diagnosis method of claim 1, wherein, The method for converting the one-dimensional vibration signal of the bearing into the time-frequency domain signal comprises the following steps: normalizing the one-dimensional vibration signal of the bearing; mapping the normalized data points to a polar coordinate system, and calculating the angles corresponding to the data points; constructing a symmetric Gram matrix based on the angles; mapping the matrix element values to the gray values or colors of image pixels to visualize the Gram matrix into an image.

6. The small sample bearing fault diagnosis method of claim 5, wherein, The angles corresponding to the data points are calculated by an inverse cosine function, In the formula, represents the normalized data points.

7. The small sample bearing fault diagnosis method of claim 5, wherein, The calculation method of the elements in the Gram matrix is as follows: where GASF ts is the element of the Gram matrix in the tth row and sth column, representing the relationship between the data points corresponding to time points t and s.

8. The small sample bearing fault diagnosis method of claim 1, wherein, The one-dimensional multi-scale dynamic network comprises: a dynamic weight generator configured to perform global average pooling on input data, and sequentially generate a weight vector through a first connection layer, a Relu activation function, a second connection layer and a Softmax function, wherein the dimension of the weight vector is the same as the number of multi-scale convolution branches; a multi-scale convolution branch configured to independently convolve the input data using different convolution kernel sizes to generate corresponding feature maps; a dynamic weighted fusion configured to weight the weight vector and the outputs of each multi-scale convolution branch, and output a dynamic convolution kernel through normalization and an activation function; and multiply the dynamic convolution kernel output by each branch with the original input data to obtain respective fusion features; a cascaded residual fusion configured to sequentially pass the respective fusion features through a Leaky Relu function, and fuse the respective fusion features in a cascaded residual manner to obtain a final output.

9. The small sample bearing fault diagnosis method of claim 1, wherein, The two-dimensional RPE-ViT network comprises an encoder and a decoder; the encoder comprises a multi-head self-attention module, a layer normalization module and a multi-layer perception module; the decoder has the same structure as the encoder, and a mask multi-head attention module is added before the multi-head self-attention module.

10. The small sample bearing fault diagnosis method of claim 1, wherein, The method for extracting global dependent fault features by using the two-dimensional RPE-ViT network comprises the following steps: dividing the time-frequency domain signal into a plurality of patches, and generating an embedding vector for each patch; adding relative position encoding to each embedding vector; inputting the encoded feature vectors into the two-dimensional RPE-ViT network for feature extraction.