Rolling bearing fault diagnosis method based on multi-scale residual network and improved GRU

CN118094371BActive Publication Date: 2026-09-25CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410100889.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2026-09-25
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

但由于转速、负载和故障类型的不同,所获取的滚动轴承振动信号也表现出多尺度特性,而传统的CNN感受野单一,不能有效地处理多尺度特征提取问题

Benefits of technology

[0053]本发明一种基于多尺度残差网络和改进GRU的滚动轴承故障诊断方法,技术效果如下:1)本发明采用层次结构和通道级的信息融合构建多尺度网络,使其在对不同尺度的特征进行深度挖掘的同时,能够有效减少模型参数,提升网络运行效率,并使用残差连接加强特征表征能力,解决深度神经网络训练过程中出现的梯度消失和梯度爆炸问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118094371B_ABST
    Figure CN118094371B_ABST
Patent Text Reader

Abstract

The method comprises the following steps: acquiring original vibration data of a rolling bearing during operation, and dividing the original vibration data into a training data set and a test data set; inputting the training data set into a multi-scale residual network for preliminary feature extraction; inputting local feature information extracted by the multi-scale residual network into an improved GRU; finally, the obtained feature information is subjected to alpha-Dropout and global average pooling processing, and then input into a softmax layer for fault classification; the parameters of the multi-scale residual network and the improved GRU are optimized to obtain a trained fault diagnosis model; and the test data set is input into the fault diagnosis model for fault diagnosis to determine the health condition of the rolling bearing. The multi-scale structure is introduced into the CNN to effectively extract the local features of the rolling bearing signal; meanwhile, the GRU is improved so that it can better capture the long-term dependence and time correlation in the sequence data, and improve the fault diagnosis performance and convergence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for diagnosing rolling bearing faults, specifically a method for diagnosing rolling bearing faults based on multi-scale residual networks and improved GRUs. Background Technology

[0002] With the rapid development of modern industry, rotating machinery is widely used in various fields such as manufacturing, transportation, and aerospace. Rolling bearings, as key components of rotating machinery, play a crucial role in ensuring the safe and efficient operation of mechanical equipment. Rolling bearings typically operate under high-intensity conditions in complex environments, which can lead to bearing damage and consequently, significant economic losses or even personal injury. Therefore, effective fault diagnosis of rolling bearings is receiving increasing attention.

[0003] In recent years, deep learning has attracted widespread attention across various fields due to its powerful feature mining capabilities, providing a new perspective for fault diagnosis. Compared with traditional fault diagnosis methods, deep learning can adaptively extract fault information from vibration signals, avoiding information loss caused by manual processing. Therefore, numerous deep learning models have been applied in the field of fault diagnosis, with Convolutional Neural Networks (CNNs) and Gated Recurrent Units (GRUs) achieving significant results. However, due to differences in rotational speed, load, and fault type, the acquired rolling bearing vibration signals exhibit multi-scale characteristics, while traditional CNNs have a single receptive field and cannot effectively handle multi-scale feature extraction. Furthermore, GRUs also suffer from long training times, poor convergence, and gradient explosion issues. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a rolling bearing fault diagnosis method based on a multi-scale residual network and an improved GRU. By introducing a multi-scale structure into the convolutional neural network (CNN), local features of the rolling bearing signal are effectively extracted, and residual connections are used to enhance feature representation capabilities. Simultaneously, the gated recurrent unit (GRU) is improved to better capture long-term dependencies and temporal correlations in sequential data, thereby improving fault diagnosis performance and convergence.

[0005] The technical solution adopted in this invention is as follows:

[0006] A rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU includes the following steps:

[0007] Step 1: Obtain raw vibration data during the operation of the rolling bearing;

[0008] Step 2: Divide the original vibration signal obtained in Step 1 into samples of a specified length, and divide them into training dataset and test dataset;

[0009] Step 3: Input the training dataset into the multi-scale residual network for preliminary feature extraction;

[0010] Step 4: Input the local feature information extracted by the multi-scale residual network into the improved GRU to further obtain temporal features;

[0011] Step 5: The final feature information is processed by α-Dropout and Global Average Pooling (GAP) and then input into the softmax layer for fault classification.

[0012] Step 6: Repeat steps 3 to 5 to fine-tune the parameters of the multi-scale residual network and the improved GRU until the diagnostic accuracy of the training dataset reaches a stable level, thus obtaining the trained fault diagnosis model.

[0013] Step 7: Input the test dataset into the trained fault diagnosis model to perform fault diagnosis and determine the health status of the rolling bearing.

[0014] In step 3, the multi-scale residual network is composed of multiple cascaded basic residual blocks. The basic residual blocks split the input signal according to the hierarchical structure and perform independent convolution operations on each branch. Then, the convolution output is fused at the channel level to obtain feature information at different scales.

[0015] Preliminary feature extraction is performed in a multi-scale residual network, including the following steps:

[0016] S3.1: In the basic residual block, given an input signal X, the basic feature map x is first obtained through a 32×32 convolution. The wide convolution kernel can provide a larger receptive field to better obtain global information.

[0017] S3.2: Subsequently, x is uniformly divided into n feature map subsets at the channel level, using x i Let i ∈ {1,2,...,n}; each feature map subset has the same spatial size and the number of channels is 1 / n, compared to the basic feature map x.

[0018] S3.3: After that, except for x1, each x i Each has a corresponding 5×5 convolution, denoted as C. i (), C i The output of () is y i In addition, each x i Before inputting Ci(), it will be compared with y. i-1are superposed. Therefore, y i is expressed as:

[0019]

[0020] S3.4: In addition, in order to better fuse feature information, the obtained new feature map y i is subjected to a concatenation operation to obtain y m , and a 1×1 convolution is used to obtain a feature map y having the same spatial size and number of channels as the basic feature map x.

[0021] S3.5: Then, x and y are superposed to implement skip connection, so as to obtain a mixed feature map y with multi-scale attributes s . S3.6: Finally, the mixed feature map y s is processed by a Batch Normalization (BN) layer and a Rectified Linear Unit (ReLU) activation layer to obtain an output Y.

[0022] In said step S3.6, batch normalization (BN) is a technique used to accelerate the neural network training process and improve model performance. It normalizes the input of each mini-batch, adjusts its mean to 0 and variance to 1, so as to stabilize the data distribution. Specifically, for an input x with a mean μ and a standard deviation σ, the output obtained after batch normalization is calculated by the formula:

[0023]

[0024] wherein γ and β are a scaling factor and a shifting factor respectively, and ε is a very small constant to avoid the case where the denominator is 0.

[0025] Further, the Rectified Linear Unit (ReLU) is a commonly used activation function, which maps negative input values to zero and keeps non-negative input values unchanged. The definition of the ReLU function is as follows:

[0026] ReLU(x)=max(0,x)

[0027] wherein x is an input value.

[0028] In step 3, preferably, in the basic residual block, the number n of splits of the basic feature map x is used as a control parameter of the scale dimension, and a larger n may allow learning features with richer receptive field sizes. Because when 1<j≤i, Ci() can accept feature information from all feature map subsets x j To simplify the structure and reduce computation, the present invention sets n to 4.

[0029] In step 4, the improved GRU redesigned the GRU's reset gate to share weights with the update gate, thereby reducing the number of parameters and improving network operating efficiency. The Self-Normalizing Exponential Linear Unit (SELU) was used as the new activation function of the GRU, thereby enhancing the network's self-normalization property during training.

[0030] The function expression for SELU is:

[0031]

[0032] Where x is the input value, e x It is an exponential function, λ = 1.0507, α = 1.6733.

[0033] The calculation process for the improved GRU is described as follows:

[0034] s t =σ(W s [h t-1 ,x t ]+b s )

[0035]

[0036]

[0037]

[0038] Where: x t This is the input at the current time step, where the subscripts t and t-1 represent the hidden layers at the current and previous time steps, respectively. σ and SELU are the activation functions, and W... s and b s Let s represent the shared weight matrix and the bias value, respectively. t a represents the combination of the current input and the memory from the previous moment. t This represents information retained from the memory of the previous moment. h represents the candidate state vector. t h represents the output of the hidden layer at the current time t; t-1 This represents the output of the hidden layer t-1 at the previous time step. This represents the Hadamard product.

[0039] Step 5 includes the following steps:

[0040] S5.1: Regularize the output of the improved GRU using α-Dropout to reduce model complexity and improve its generalization ability;

[0041] S5.2: Use GAP to perform dimensionality reduction on the features after S5.1 regularization, so as to reduce the high-dimensional feature vector to a one-dimensional vector required by the classifier.

[0042] S5.3: Input the final feature vector into the softmax layer for fault classification.

[0043] In S5.1, α-Dropout introduces a parameter α to control the dropout probability of each neuron, and dynamically adjusts the dropout probability based on the importance of the neuron. Specifically, assuming y i This represents the output of the i-th neuron and its corresponding dropout probability d. i It can be represented as:

[0044]

[0045] Where: p represents the initial global dropout probability, which is used to control the proportion of neurons dropped during training, and the hyperparameter α is responsible for adjusting the relationship between the importance of neurons and the probability of dropping neurons.

[0046] In S5.2, GAP differs from traditional pooling layers; it calculates the average value of each channel as the channel output. This allows GAP to reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP can be determined by the following formula:

[0047]

[0048] in, Let N represent the output of the j-th channel of the previous network, N represent the number of output channels, L represent the length of each channel, and V represent the length of each channel. j This represents the dimensionality-reduced output of the j-th channel.

[0049] In S5.3, the final feature vector is input into the softmax layer for fault classification. Softmax is a commonly used activation function, mainly used for predicting sample categories in multi-class classification problems. Specifically, when the input vector Z has n dimensions, the probability distribution P of the i-th dimension is... i It can be written as:

[0050]

[0051] Among them, Z i and Z j Let i and j represent the i-th and j-th dimensions of the input vector Z, respectively.

[0052] In step 6, the multi-scale residual network, the improved GRU, and the subsequent optimization and classification network constitute the fault diagnosis model. Steps 3 to 5 are repeated to fine-tune the parameters of the multi-scale residual network and the improved GRU until the diagnostic accuracy of the training dataset reaches a stable level, thus obtaining the trained fault diagnosis model.

[0053] This invention provides a rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU. The technical effects are as follows: 1) This invention uses hierarchical structure and channel-level information fusion to construct a multi-scale network, which can effectively reduce model parameters and improve network operating efficiency while deeply mining features at different scales. It also uses residual connections to enhance feature representation capabilities and solve the gradient vanishing and gradient explosion problems that occur during deep neural network training.

[0054] 2) This invention newly designs the reset gate of GRU, which shares weights with the update gate to reduce the number of parameters and improve network operating efficiency. It also introduces SELU into the activation function to enhance the self-normalization property of the network during training, promote the signal to approach zero mean and unit variance during training, and enhance the anti-interference performance against uncorrelated signals.

[0055] 3) This invention also uses SELU as a new activation function and introduces self-normalization properties to construct an improved GRU in order to capture temporal features more effectively.

[0056] 4) The multi-scale residual network in this invention can deeply mine features at different scales, and adopts hierarchical structure and channel-level information fusion to effectively reduce model parameters and improve network operating efficiency. Attached Figure Description

[0057] Figure 1 This is a flowchart of a rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU proposed in this invention.

[0058] Figure 2 This is a schematic diagram illustrating the signal segmentation and dataset partitioning involved in this invention.

[0059] Figure 3 This is a diagram of the basic residual block structure of the multi-scale residual network proposed in this invention.

[0060] Figure 4 This is a diagram of the improved GRU structure proposed in this invention.

[0061] Figure 5 This is a visualization distribution diagram of the original data in a preferred embodiment of the present invention.

[0062] Figure 6 This is a visualization distribution map of the features extracted in a preferred embodiment of the present invention.

[0063] Figure 7 This is a confusion matrix diagram of diagnostic results according to a preferred embodiment of the present invention. Detailed Implementation

[0064] To more clearly and completely illustrate the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0065] A rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRUs is illustrated in the flowchart below. Figure 1 As shown, the main steps include:

[0066] S1: Obtain raw vibration data of the rolling bearing during operation through a data acquisition system;

[0067] The data acquisition system mainly consists of a motor, transmission mechanism, motor driver, vibration sensor, data acquisition board, computer and its supporting software, and is used to acquire the vibration signals generated by the rolling bearing during operation.

[0068] S2: Divide the collected vibration signal into samples of a specified length and divide them into training dataset and test dataset;

[0069] The vibration signals acquired by the data acquisition system are single-dimensional time-series signals. Before fault diagnosis, they need to be segmented into sample signals of specific lengths and numbers. Specific segmentation methods include... Figure 2 As shown, for a signal with a total length of L, assuming the segmentation length is l, we can obtain a number of samples n = L / l. Then, based on this, the samples are divided into a training dataset and a test dataset.

[0070] S3: Input the training dataset into the multi-scale residual network for preliminary feature extraction;

[0071] Although the vibration signals of bearings with different fault types have different characteristics, they are not easy to distinguish without processing. Multi-scale residual networks can be used to mine in depth at different scales and obtain detailed local feature information.

[0072] S4: Input the local feature information extracted by the multi-scale residual network into the improved GRU to further obtain temporal features;

[0073] The feature information extracted by the multi-scale residual network is mainly the local feature information of the vibration signal. The vibration signal is a time series. Using the improved GRU for time series feature extraction can make the feature differentiation of rolling bearing vibration signals in different states higher.

[0074] S5: The final feature information is processed by α-Dropout and GAP, and then input into the softmax layer for fault classification.

[0075] The feature information extracted by the multi-scale residual network and the improved GRU has high discriminative power. After optimization with α-Dropout and GAP, a set of one-dimensional feature data is obtained. These data come from rolling bearings with different fault states. Softmax is often used for multi-classification problems and can distinguish different fault state data to achieve fault classification.

[0076] S6: Repeat S3-S5 to fine-tune the parameters of the multi-scale residual network and the improved GRU until the diagnostic accuracy of the training dataset reaches a stable level, thus obtaining the trained fault diagnosis model.

[0077] S7: Input the test dataset into the trained fault diagnosis model to perform fault diagnosis and determine the health status of the rolling bearing.

[0078] Furthermore, in S3, the multi-scale residual network is composed of multiple cascaded basic residual blocks, as shown in the basic residual block structure diagram. Figure 3 As shown, specifically, given an input signal X, a basic feature map x is first obtained through a 32×32 convolution. A wide convolutional kernel can provide a larger receptive field to better capture global information. Then, x is evenly divided into n feature map subsets at the channel level, and x is used as the basis for further processing. i This means that i∈{1,2,...,n}. Each subset of feature maps has the same spatial size as the basic feature map x, but the number of channels is 1 / n. Then, each x except x1... i Each has a corresponding 5×5 convolution, denoted as C. i (), C i The output of () is y i In addition, each x i Before inputting Ci(), it will be compared with y. i-1 The layers are superimposed. Therefore, y i It can be represented as:

[0079]

[0080] Furthermore, in order to better integrate feature information, the resulting new feature map y is... i Perform a splicing operation to obtain y m Then, a 1×1 convolution is used to obtain a feature map y with the same spatial size and number of channels as the basic feature map x. Next, x and y are concatenated to implement a skip connection, resulting in a hybrid feature map y with multi-scale properties. s Finally, the hybrid feature map y sAn output Y is obtained after processing by a Batch Normalization (BN) layer and a Rectified Linear Unit (ReLU) activation layer.

[0081] Preferably, in the basic residual block, the number of splits n of the basic feature map x is used as a control parameter for the scale dimension, and a larger n may allow learning features with richer receptive field sizes. Because when 1<j≤i, C​i​() can receive feature map subsets x from all j feature information. To simplify the structure and reduce computation, the present invention sets n to 4.

[0082] Further, in said step S4, the structural diagram of the improved GRU is as Figure 4 shown. The reset gate of GRU is redesigned to share weights with the update gate, so as to reduce the number of parameters and improve the operation efficiency of the network, and SELU is adopted as the new activation function of GRU, thereby improving the self-normalization property of the network during training, promoting the signal to approach zero mean and unit variance during training, and enhancing the anti-interference performance against irrelevant signals.

[0083] The functional expression of SELU is:

[0084]

[0085] wherein, x is the input value, e x is the exponential function, λ=1.0507, α=1.6733.

[0086] The calculation process of the improved GRU is described as follows:

[0087] s t =σ(W s [h t-1 ,x t +b s )

[0088]

[0089]

[0090]

[0091] wherein: x t is the input at the current time step, subscripts t and t-1 represent the hidden layer at the current time and the previous time respectively, σ and SELU are activation functions, W s and b s represent the shared weight matrix and bias value respectively, s t represents the combination of the current input and the previous moment memory, a tThis represents information retained from the memory of the previous moment. h represents the candidate state vector. t h represents the output of the hidden layer at the current time t. t-1 This represents the output of the hidden layer t-1 at the previous time step. This represents the Hadamard product.

[0092] Preferably, the activation function tanh used in traditional GRUs has a derivative that tends to zero when the input exceeds a certain range. This means the network struggles to update weights via gradient descent, potentially leading to gradient vanishing and reduced network stability. The function curves of tanh and SELU are shown in the figure. Therefore, improving GRU by using SELU as the activation function not only incorporates self-normalization properties into the network but also avoids gradient vanishing, improving the network's robustness and convergence.

[0093] Furthermore, S5 can be subdivided into the following three steps:

[0094] s51: Regularize the output of the improved GRU using α-Dropout to reduce model complexity and improve its generalization ability;

[0095] s52: Use GAP to reduce the dimensionality of the regularized features so as to reduce the high-dimensional feature vector to a one-dimensional vector required by the classifier.

[0096] s53: Input the final feature vector into the softmax layer for fault classification.

[0097] Preferably, traditional regularization methods randomly drop neurons in network layers during training to reduce dependencies between neurons. However, using the same dropout probability for each neuron does not take into account the differences between neurons. In α-Dropout, a parameter α is introduced to control the dropout probability of each neuron, and the dropout probability is dynamically adjusted based on the importance of the neuron. Specifically, assuming y i This represents the output of the i-th neuron and its corresponding dropout probability d. i It can be represented as:

[0098]

[0099] Where p represents the initial global dropout probability, which is used to control the proportion of neurons dropped during training, and the hyperparameter α is responsible for adjusting the relationship between the importance of neurons and the probability of dropping neurons.

[0100] Preferably, GAP differs from traditional pooling layers in that it calculates the average value of each channel as the channel output. In this way, GAP can reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP can be determined by the following formula:

[0101]

[0102] in, Let N represent the output of the j-th channel of the previous network, N represent the number of output channels, L represent the length of each channel, and V represent the length of each channel. j This represents the dimensionality-reduced output of the j-th channel.

[0103] The following detailed description is provided in conjunction with specific embodiments:

[0104] A preferred embodiment of the present invention is the application of a rolling bearing fault diagnosis method based on a multi-scale residual network and an improved GRU on the bearing fault diagnosis dataset (CWRU Bearing Data Center) at Case Western Reserve University in the United States.

[0105] The CWRU bearing dataset consists of vibration data collected from the drive end under four load conditions by an accelerometer with a sampling frequency of 12 kHz. The vibration data is divided into fault and normal states. The fault state includes three damaged components: the inner ring, the ball bearing, and the outer ring. Each damaged component is further divided into three damage levels: 0.007, 0.014, and 0.021 inches. Therefore, the data for each load condition includes 10 operating states.

[0106] In this embodiment, the vibration data was divided into 118 samples, each containing 1024 continuous data points. Then, 78 samples were randomly selected from these 118 samples as the training set, and the remaining samples were used as the test set. Specific information about the dataset is shown in Table 1.

[0107] Table 1. Description of CWRU bearing dataset

[0108]

[0109]

[0110] Furthermore, the training set is input into the deep learning model of the multi-scale residual network and improved GRU proposed in this invention for training. After the diagnostic accuracy rises to a stable level, the trained model is obtained. Then, the test set data is input into the model for testing.

[0111] The t-Distributed Stochastic Neighbor Embedding (t-SNE) method was used to process the original data under 0hp load and the features extracted from the original data by the proposed model to verify the feature extraction capability of the proposed model. Figure 5 As shown, the different fault characteristics of the original data exhibit overlapping states and are indistinguishable. Conversely, as... Figure 6 As shown, the feature distribution extracted by the model proposed in this invention is separable and can be easily partitioned, indicating that the proposed model has excellent feature extraction capabilities. Furthermore, the confusion matrix of the test set at different rotational speeds is shown below. Figure 7 As shown, the proposed model accurately identifies 10 operating conditions under different load conditions, proving that it can accurately identify different bearing failure types.

[0112] Preferably, to further verify the superiority of the proposed model, CNN, GRU, and CNN-GRU were used as basic comparison models. Simultaneously, to demonstrate the superior feature extraction capability of the multi-scale residual network and the effectiveness of the GRU improvement, the multi-scale residual structure and the GRU improvement were used as ablation features, resulting in two ablation comparison models. For ease of recording, the multi-scale residual network is denoted as MSRN, and the improved GRU is denoted as EGRU. Detailed comparative experimental results are shown in Table 2.

[0113] Table 2 Experimental results of different models

[0114]

[0115] As shown in Table 2, the diagnostic accuracy of the model proposed in this invention is above 99.9% under all four load conditions, and is higher than the other three basic comparison models, demonstrating its superior fault diagnosis performance. Furthermore, although the diagnostic accuracy of the ablation comparison model is lower than that of the model proposed in this invention, it still has advantages compared to the basic comparison models. This proves the excellent feature extraction capability of the multi-scale residual network proposed in this invention and the effectiveness of its GRU improvement.

Claims

1. A rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU, characterized in that... Includes the following steps: Step 1: Obtain raw vibration data during the operation of the rolling bearing; Step 2: Divide the original vibration signal obtained in Step 1 into samples of a specified length, and divide them into training dataset and test dataset; Step 3: Input the training dataset into the multi-scale residual network for preliminary feature extraction; Step 4: Input the local feature information extracted by the multi-scale residual network into the improved GRU to further obtain temporal features; Step 5: The final feature information is processed by α-Dropout and global average pooling (GAP) and then input into the softmax layer for fault classification. Step 6: Repeat steps 3 to 5 to fine-tune the parameters of the multi-scale residual network and the improved GRU to obtain the trained fault diagnosis model. Step 7: Input the test dataset into the trained fault diagnosis model to perform fault diagnosis and determine the health status of the rolling bearing; In step 3, preliminary feature extraction is performed in the multi-scale residual network, including the following steps: S3.1: In the basic residual block, given an input signal X First, the basic feature map is obtained through a 32×32 convolution. x ; S3.2: Subsequently, x Evenly divided into channels n A subset of feature maps, used x i express, i ∈{1, 2, ..., n }; Mapping with basic features x In contrast, each feature map subset has the same spatial size and the number of channels is 1 / n ; S3.3: After that, except x Besides 1, each x i Each has a corresponding 5×5 convolution, denoted as C i (), C i The output of () is used y i In addition, each x i In the input Ci () will be before y i-1 Superimposed; therefore, y i Represented as: ; S3.4: In addition, the resulting new feature mapping y i Perform a splicing operation to obtain y m And obtain the basic feature mapping through a 1×1 convolution. x Feature mappings with the same spatial size and number of channels y ; S3.5: Then, x and y By performing superposition to achieve skip connections, a hybrid feature map with multi-scale attributes is obtained. y s S3.6: Finally, hybrid feature mapping y s The output is obtained after batch normalization (BN) and ReLU activation layer processing. Y ; In step 4, the improved GRU redesigned the GRU's reset gate to share weights with the update gate, and used the self-normalized exponential linear unit SELU as the new activation function of the GRU. The function expression for SELU is: ; in, x It is the input value. e x It is an exponential function, λ=1.0507, α=1.6733; The calculation process for the improved GRU is described as follows: ; ; ; ; in: x t This is the input for the current time step, the index. t and t -1 represents the hidden layer at the current time step and the previous time step, respectively, and σ and SELU are the activation functions. W s and b s These represent the shared weight matrix and the bias values, respectively. s t This represents the combination of the current input and the memory from the previous moment. a t This represents information retained from the memory of the previous moment. Represents the candidate state vector. h t Indicates the hidden layer at the current time. t The output; Indicates the hidden layer at the previous moment t -1 output, This represents the Hadamard product.

2. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 1, characterized in that: In S3.6, batch normalization (BN) is a technique used to accelerate the training process of neural networks and improve model performance. It normalizes each mini-batch of inputs, adjusting its mean to 0 and its variance to 1. Specifically, for the input... x Its mean is μ The standard deviation is σ The output is obtained after batch normalization. , The calculation formula is: ; in, and These are the scaling factor and the translation factor, It is a very small constant to avoid the case where the denominator is 0; The Rectified Linear Unit (ReLU) is a commonly used activation function that maps negative input values ​​to zero while leaving non-negative input values ​​unchanged. The ReLU function is defined as follows: ; in, x This is the input value.

3. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 2, characterized in that: In step 3, the basic feature mapping is performed in the basic residual block. x Number of splits n As a control parameter for the scale dimension, a larger n Allows learning to have richer features of receptive field size; because when 1 <j≤i hour, Ci () accepts subsets of all feature maps x j The characteristic information; in order to simplify the structure and reduce calculations, n Set it to 4.

4. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 1, characterized in that: Step 5 includes the following steps: S5.1: Regularize the output of the improved GRU using α-Dropout; S5.2: Dimensionality reduction operation is performed on the features after S5.1 regularization using GAP; S5.3: Input the final feature vector into the softmax layer for fault classification.

5. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 4, characterized in that: In S5.1, a parameter was introduced in α-Dropout. α This controls the dropout probability of each neuron and dynamically adjusts the dropout probability based on the importance of the neuron; let... y i Indicates the first i The output of each neuron and its corresponding dropout probability d i It can be represented as: ; in: p This represents the initial global dropout probability, used to control the proportion of neurons dropped during training; it's a hyperparameter. α It is responsible for adjusting the relationship between the importance of neurons and the probability of discarding neurons.

6. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 4, characterized in that: In S5.2, GAP calculates the average value of each channel as the channel output; in this way, GAP can reduce the spatial dimension of the feature map to a single dimension. The calculation of GAP is determined by the following formula: ; in, Indicates the previous network number j Output of each channel, N Indicates the number of output channels. L Indicates the length of each channel. V j Indicates the first j Dimensionally reduced output of each channel.

7. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 4, characterized in that: In S5.3, the final feature vector is input into the softmax layer for fault classification; softmax is used to predict the sample category in multi-class classification problems. Z The dimension is n At that time, the first i The probability distribution of each dimension P i Written as: ; in, Z i and Z j They represent the input vectors respectively. Z The i and j Each dimension.

8. The rolling bearing fault diagnosis method based on multi-scale residual networks and improved GRU according to claim 1, characterized in that: In step 6, the multi-scale residual network, the improved GRU, and the subsequent optimization and classification network constitute the fault diagnosis model. Steps 3 to 5 are repeated to fine-tune the parameters of the multi-scale residual network and the improved GRU to obtain the trained fault diagnosis model.