Bearing fault analysis method based on multiple attention and Mamba network

Through the combination of multiple attention and Mamba network, the problems of multi-sensor data fusion and long-distance feature extraction in high-speed train bearing fault diagnosis are solved, improving the accuracy and robustness of bearing fault analysis, and achieving an optimized balance between global and local information.

CN120256866APending Publication Date: 2025-07-04CENT SOUTH UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510334010.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the diagnosis of bearing faults of high-speed trains, the multi-sensor data fusion effect is poor, making it difficult to achieve local and global information balance of long-distance feature extraction, and the diagnosis is insufficient under variable working conditions and strong noise interference, resulting in low accuracy.

Method used

The bearing failure analysis method based on multiple attention and Mamba network is adopted, and the model parameters are optimized through neural architecture search network, combined with two-dimensional convolution kernel, time pyramid attention module, Mamba network, sequence fusion attention module and HAB attention network, to achieve global and local timing feature extraction of information in different time and different sensors of the same sensor, and optimize the balance of global and local information.

Benefits of technology

It improves the accuracy of bearing fault analysis, improves the diagnostic robustness under variable working conditions and strong noise interference, realizes sufficient feature extraction of information and time information between sensor channels, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256866A_ABST
    Figure CN120256866A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of high-speed train bearing fault analysis, and provides a bearing fault analysis method based on multiple attention and a Mamba network, and the method comprises the steps: obtaining an original fault signal of a high-speed train bearing; a neural architecture search network is utilized to optimize model parameters of the bearing fault analysis model; processing the original fault signal by using the bearing fault analysis model after model parameter optimization to obtain a fault analysis result of the high-speed train bearing; the bearing fault analysis model comprises a two-dimensional convolution kernel, a time pyramid attention module, an expansion mode searchable Mama network, a two-dimensional CNN network, an addition module, a sequence fusion attention module, a preprocessing module, an HAB attention network and an output module. According to the invention, the accuracy of bearing fault analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of high-speed train bearing fault analysis, and particularly relates to a bearing fault analysis method based on multi-attention and Mamba network. Background Technique

[0002] In the modern industrial equipment system, rotating machinery plays an irreplaceable key role. As one of the core rotating components in the bogie, the bearings of high-speed trains usually operate under harsh conditions such as high speed and heavy load, and their failure may trigger system-level failures or even catastrophic accidents. Therefore, reliable condition monitoring and fault diagnosis are of great significance for the safe operation of high-speed trains. However, current condition detection and fault diagnosis methods still face many challenges when dealing with massive monitoring data:

[0003] 1. In the fault diagnosis of high-speed trains, the train is usually equipped with multiple sensors to collect vibration signals and other operation data. However, existing methods have significant deficiencies in extracting spatio-temporal correlation features between different sensor data, especially in cross-sensor data fusion. Due to the heterogeneity of sensor data, traditional deep learning methods are difficult to effectively capture the internal relationship between multi-source data, resulting in poor fusion effect of multi-sensor data, which greatly affects the accuracy and reliability of the fault diagnosis model.

[0004] 2. For the long-distance feature extraction of bearing vibration signals, traditional deep learning methods often face significant limitations. When processing long-sequence data, existing models usually change the size of the sliding window to capture key long-distance information. However, if the window is too large, local details are easily lost, and if the window is too small, global context feature extraction is insufficient. How to achieve the optimal balance between local and global information in long-distance feature extraction is still an urgent problem to be solved in current fault diagnosis.

[0005] 3. In a complex environment with variable working conditions and strong noise interference, the robustness of high-speed train fault diagnosis is insufficient. In actual operation, working condition factors such as the running speed and load of the train, as well as track excitation noise and electromagnetic interference, will cause distributed drift of the signal and signal-to-noise ratio fluctuation. These factors make the fault features and noise components highly overlapped, thus affecting the robustness and accuracy of the diagnosis model.

[0006] How to improve the diagnosis accuracy in an uncertain environment is still the key challenge to enhance the application effect of intelligent fault diagnosis technology. These characteristics make it difficult for intelligent fault diagnosis models to generalize to data with different working conditions of the same type of equipment. Therefore, how to complete the multi-source information fusion of information between sensors with limited available data, achieve the optimal balance between local and global in long-distance feature extraction, and enhance the diagnosis robustness under variable working conditions and strong noise interference has always been an urgent problem to be solved in intelligent fault analysis.

[0007] For the demand of deeply extracting the information connection between multiple channels of each sensor and the time series information of a single channel in the existing sample data, a direct and effective method is to introduce an attention mechanism to automatically focus on important features and improve the accuracy of fault diagnosis. However, the current mainstream attention mechanism methods are usually single-channel attention mechanisms or time attention mechanisms, which can often only capture the single front-back dependence relationship of the channel sequence and the time sequence, and it is difficult to distinguish whether to process different times of the same sensor first or the information of different sensors at the same time first, resulting in an incomplete understanding of the channel and time sequence by the model and a low accuracy rate in bearing fault analysis. Summary of the Invention

[0008] An embodiment of the present application provides a bearing fault analysis method based on multiple attention and Mamba network, which can solve the problem of low accuracy rate in bearing fault analysis.

[0009] An embodiment of the present application provides a bearing fault analysis method based on multiple attention and Mamba network, including:

[0010] Obtain the original fault signal of the high-speed train bearing;

[0011] Use the neural architecture search network to optimize the model parameters of the bearing fault analysis model;

[0012] Use the bearing fault analysis model with optimized model parameters to process the original fault signal to obtain the fault analysis result of the high-speed train bearing;

[0013] The bearing fault analysis model includes a two-dimensional convolution kernel, a time pyramid attention module, a Mamba network with a searchable unfolding method, a two-dimensional CNN network, an addition module, a sequence fusion attention module, a preprocessing module, a HAB attention network, and an output module;

[0014] The two-dimensional convolution kernel is used to adjust the number of channels of the original fault signal. The time pyramid attention module is used to extract preliminary time series features from the data output by the two-dimensional convolution kernel. The Mamba network is used to extract the global time series features of the preliminary time series features. The two-dimensional CNN network is used to extract the local time series features of the preliminary time series features. The addition module is used to perform an addition process on the global time series features and the local time series features. The sequence fusion attention module is used to perform fusion and dimensionality reduction on the multi-channel data output by the addition module. The preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module. The HAB attention network is used to extract the time series information of the data output by the preprocessing module. The output module is used to map the time series information and output the fault analysis result.

[0015] Optionally, the temporal pyramid attention module is used to extract the preliminary temporal features Out through the following formula TPA :

[0016]

[0017] C = concat(A1, A2, A3, dim = 1)

[0018] Out TPA = sigmoid(sigmoid(Relu(conv(C)))) * inupt + input

[0019] Where input represents the data output by the two-dimensional convolution kernel, None represents a placeholder, X1 represents the searchable viewing distance, Adaptiveavgpool2d represents the adaptive average pooling operation, concat represents the concatenation operation, dim represents the dimension, conv represents the convolution operation, Relu represents the Relu function, and sigmoid represents the sigmoid function;

[0020] The size of A1 is B * C * H * (1 × X), the size of A2 is B * C * H * (2 × X), the size of A3 is B * C * H * (4 × X), and B, C, H are the sizes of the data output by the two-dimensional convolution kernel.

[0021] Optionally, the Mamba network is used to extract the global temporal features Out through the following formula mamba :

[0022]

[0023] Input mamba = Out TPA .view i

[0024] Where mamba(·) represents the Mamba network, view i represents the i-th search order, i = 1, 2, 3, 4.

[0025] Optionally, the two-dimensional CNN network is a pseudo two-dimensional convolution kernel.

[0026] Optionally, the addition module is used to add the global temporal features and the local temporal features through the following formula:

[0027] Out C = Out mamba + CNN(Out TPA )

[0028] Where Out CRepresents the multi-channel data output by the addition module, CNN(Out TPA ) represents the data output by the two-dimensional CNN network.

[0029] Optionally, the sequence fusion attention module is used to fuse and reduce the dimension of the multi-channel data output by the addition module through the following formula:

[0030] Out ronghe = mean(Out C * Att, dim = 2)

[0031] Att = AdaptiveAvgpool2d (1,None) split(W)

[0032] W = Relu(Conv (X2,1) (X in ))

[0033] X in = stack(Out C , dim = 2)

[0034] Among them, Out ronghe represents the data output by the sequence fusion attention module, mean represents the average reduction operation, AdaptiveAvgpool2d represents the adaptive average pooling operation, None represents the placeholder, split represents the operation of taking the upper half segment according to the dimension of dim = 2, Relu represents the Relu function, Conv represents the convolution operation, X2 represents the size of the convolution kernel for performing this convolution operation, stack represents the stacking operation, and dim represents the dimension.

[0035] Optionally, the preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module through the following formula:

[0036] Out pred = Maxpool(Relu(BN(Conv1d(Out ronghe ))))

[0037] Among them, Out pred represents the data output by the preprocessing module, Maxpool represents the max pooling function, Relu represents the Relu function, BN represents the normalization function, and Conv1d represents the convolution operation.

[0038] Optionally, the HAB attention network is used to extract the temporal information Out of the data output by the preprocessing module through the following formula T :

[0039] Out T = Out CAB+Out W-MSA +Out pred

[0040] Out CAB =Att CAB *Out pred

[0041] Att CAB =sigmoid(Conv1d(CAB t ))

[0042] CAB t =Relu(Conv1d(AdaptiveAvgpool1d (1) Out pred ))

[0043] Out W-MSA =softmax(K*Q)*V

[0044] K, Q, V = FC(Out pred .view(-1, X3, c))

[0045] Among them, sigmoid represents the sigmoid function, Conv1d represents the convolution operation, Relu represents the Relu function, AdaptiveAvgpool1d (1) represents the shape after adaptive average pooling, softmax represents the softmax function, K represents the K matrix corresponding to the window-based multi-head attention mechanism, Q represents the Q matrix corresponding to the window-based multi-head attention mechanism, V represents the V matrix corresponding to the window-based multi-head attention mechanism, FC represents the fully connected layer, view(-1, X3, c) represents reshaping the tensor into (-1, X3, c), X3 represents the window size, and c represents the input channel dimension.

[0046] Optionally, the output module includes an average pooling layer and a fully connected layer connected in sequence. The input end of the average pooling layer is the input end of the output module, and the output end of the fully connected layer is the output end of the output module.

[0047] The above solution of this application has the following beneficial effects:

[0048] In the embodiments of the present application, the Mamba network and the two-dimensional CNN network that can be searched by the expansion method are used to extract the global and local temporal features of the information of the same sensor at different times and the information of different sensors at the same time. Then, the HAB attention network is further used to optimize the global and local temporal features. At the same time, the neural architecture search network is used to optimize the model parameters to obtain the optimal balance of global and local information, so that the feature extraction of the information between the channels of the sensor and the time information is sufficient, and the global and local information is balanced in the long-distance feature extraction. Furthermore, when the bearing fault analysis is performed based on the temporal information output by the HAB attention network, the accuracy rate is higher than that of the traditional bearing fault diagnosis method, achieving the effect of improving the accuracy rate of the bearing fault analysis.

[0049] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the implementation examples or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of the bearing fault analysis method based on multiple attention and the Mamba network provided by an embodiment of the present application;

[0052] Figure 2 It is a schematic structural diagram of the bearing fault analysis model provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic diagram of the search scheme provided by an embodiment of the present application;

[0054] Figure 4 It is a schematic structural diagram of the test bench model in the example;

[0055] Figure 5a It is a schematic diagram of the sensor layout position in the example Figure 1 ;

[0056] Figure 5b It is a schematic diagram of the sensor layout position in the example Figure 2 ;

[0057] Figures 6a - 6f are schematic diagrams of the results of the confusion matrix and accuracy rate of the bearing fault analysis method in the example, where a, b, c, d, e, f respectively represent the experimental model of this experiment, Self-CNN, GRU-CNN, LSTM-CNN, Bi-LSTM, Bi-Transformer. Detailed implementation manners

[0058] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0059] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0060] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0061] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0062] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0063] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0064] Aiming at the problem of low accuracy in current bearing fault analysis, the embodiments of the present application provide a bearing fault analysis method based on multiple attention and Mamba network. This method extracts the global and local temporal features of information at different times of the same sensor and information of different sensors at the same time through the expandable searchable Mamba network and two-dimensional CNN network. Then, the global and local temporal features are further optimized through the HAB attention network. At the same time, the neural architecture search network is used to optimize the model parameters to obtain the optimal balance of global and local information, so that the feature extraction of information between channels and time information of the sensor is sufficient, and the global and local information is balanced in long-distance feature extraction. Furthermore, the accuracy of bearing fault analysis based on the temporal information output by the HAB attention network is higher than that of traditional bearing fault diagnosis methods, achieving the effect of improving the accuracy of bearing fault analysis.

[0065] The following will exemplarily illustrate the bearing fault analysis method based on multiple attention and Mamba network provided by the present application in combination with specific embodiments.

[0066] As Figure 1 shown, the bearing fault analysis method based on multiple attention and Mamba network provided by the embodiments of the present application includes the following steps:

[0067] Step 11, obtaining the original fault signal of the high-speed train bearing.

[0068] In some embodiments of the present application, the fault vibration signal (i.e., the above-mentioned original fault signal) of the high-speed train bearing can be collected by multiple vibration sensors arranged at different positions on the high-speed train bogie.

[0069] Step 12, using the neural architecture search network to optimize the model parameters of the bearing fault analysis model.

[0070] In the related art, the neural architecture search network aims to automate the design of neural network architectures rather than manually designing the architectures. Its goal is to continuously explore different network structures in the search space to find the architecture with the optimal performance. In some embodiments of the present application, the model parameters of the bearing fault analysis model can be specifically optimized by using reinforcement learning. Exemplarily, the reinforcement learning Actor-Critic algorithm agent can be used to achieve the automated search of the neural architecture under the interaction of multiple factors of the model and optimize the model parameters of the bearing fault analysis model. Among them, the model parameters here include but are not limited to all learnable parameters in the bearing fault analysis model.

[0071] It should be noted that the specific optimization process will be introduced in detail in combination with the model structure later.

[0072] Step 13: Process the original fault signal using the bearing fault analysis model with optimized model parameters to obtain the fault analysis results of the high-speed train bearing.

[0073] The bearing fault analysis model will be exemplarily described below in conjunction with specific embodiments.

[0074] As Figure 2 shown, the above-mentioned bearing fault analysis model includes a two-dimensional convolution kernel, a temporal pyramid attention module, a Mamba network with a searchable unfolding method, a two-dimensional CNN network, an addition module, a sequence fusion attention module, a preprocessing module, a HAB attention network, and an output module.

[0075] The above-mentioned two-dimensional convolution kernel is used to adjust the number of channels of the original fault signal. Specifically, the size of the two-dimensional convolution kernel can be 1 to adjust the number of channels of the original fault signal.

[0076] The above-mentioned temporal pyramid attention module is used to extract preliminary temporal features from the data output by the two-dimensional convolution kernel. It should be noted that the temporal pyramid attention module is a temporal pyramid attention module with searchable granularity and can extract the temporal features of a single channel.

[0077] Specifically, the temporal pyramid attention module is mainly used to extract the preliminary temporal feature Out through the following formula TPA :

[0078]

[0079] C = concat(A1, A2, A3, dim = 1)

[0080] Out TPA = sigmoid(sigmoid(Relu(conv(C)))) * inupt + input

[0081] where input represents the data output by the two-dimensional convolution kernel; None represents a placeholder, usually indicating that the size of this dimension is determined by the actual size of the input data; X1 represents the searchable look-ahead distance, which can be specifically determined which granularity of look-ahead distance to select during the optimization of the model parameters of the bearing fault analysis model (for example, in the 0th search space of Table 1, the searchable parameter X1 is 1); Adaptiveavgpool2d represents the adaptive average pooling operation, where 2d means that the input data of the adaptive average pooling operation is 2-dimensional data. It should be noted that 2d in the following text also refers to the data dimension being 2-dimensional, and 1d refers to the data dimension being 1-dimensional; concat represents the concatenation operation; dim represents the dimension; conv represents the convolution operation; Relu represents the Relu function; sigmoid represents the sigmoid function.

[0082] The size of A1 is B*C*H*(1×X), the size of A2 is B*C*H*(2×X), and the size of A3 is B*C*H*(4×X). B, C, and H are the sizes of the data output by the two-dimensional convolutional kernel. B is the batch size, C is the number of channels, and H is the height.

[0083] In some embodiments of the present application, the above-mentioned temporal pyramid attention module is obtained by improving the position attention module. After the data output by the two-dimensional convolutional kernel is input into the temporal pyramid attention module, the temporal pyramid attention module is mainly used to perform the following processing on it:

[0084] The data is divided into blocks with three different granularities to obtain three different granularities of attention, and then unified processing is performed. In order to adapt to the tasks of the present application, in the multi-channel preprocessing stage, in order to avoid the mixing of multi-channel temporal information, the idea of pseudo-2D is adopted to obtain multi-granularity self-attention. The Adaptiveavgpool2d of the traditional self-attention mechanism (1) is respectively modified to Adaptiveavgpool2d (None,X) , Adaptiveavgpool2d (None,2X) , Adaptiveavgpool2d (None,4X) , so that the calculation of attention is limited to different times of the same sequence, avoiding the confusion of different sequences and different times, and ensuring the performance of the model:

[0085] A1 = Adaptiveavgpool2d (None,X) (input)

[0086] A2 = Adaptiveavgpool2d (None,2X) (input)

[0087] A3 = Adaptiveavgpool2d (None,4X) (input)

[0088] After that, the results extracted from each channel are merged to obtain multi-channel data similar to a 2D image. The specific method is as follows:

[0089] C = concat(A1, A2, A3, dim = 1)

[0090] Out TPA = sigmoid(sigmoid(Relu(conv(C)))) * inupt + input

[0091] Among them, in the concat operation, A1, A2, and A3 are uniformly adjusted to B*C*H*(4×X), and then concatenated according to the dimension dim = 1.

[0092] The above Mamba network is used to extract the global temporal features of the preliminary temporal features, the two-dimensional CNN network is used to extract the local temporal features of the preliminary temporal features, and the addition module is used to add the global temporal features and the local temporal features.

[0093] Among them, the Mamba network is used to extract the global temporal feature Out through the following formula mamba :

[0094]

[0095] Input mamba = Out TPA .view i

[0096] Among them, mamba(·) represents the Mamba network, and view i represents the i-th search order, where i = 1, 2, 3, 4.

[0097] It should be noted that the input data of the expandable searchable Mamba network is two-dimensional, and one-dimensional Mamba processing is performed according to the search direction through the search method. In the embodiments of the present application, there are 4 search orders for the input data of the Mamba network (specifically, the 4 search schemes shown in Figure 3 ), and multiple search methods enable multiple ways of processing the input Mamba network, so that the Mamba network can extract the global temporal features of the information of the same sensor at different times and the information of different sensors at the same time.

[0098] The above two-dimensional CNN network is a two-dimensional CNN network with searchable granularity, which can specifically be a pseudo-two-dimensional convolution kernel. The pseudo-two-dimensional convolution kernel: compared with the 2*2 two-dimensional convolution kernel, multiple different scales of convolution kernels such as 1*3, 1*5, 1*7, 1*9, etc. are used in combination to extract simultaneously, and the extraction of two-dimensional features can also be achieved, but essentially it is the fusion of the results of two one-dimensional feature extractions. In order to prevent feature confusion between channels, a pseudo-two-dimensional convolution of one size among 1*1, 1*3, 1*5, 1*7, 1*9, 1*11 can be used for extraction. It should be noted that in the present application, the selection of the size of the pseudo-two-dimensional convolution can be determined during the optimization process of the model parameters of the bearing fault analysis model.

[0099] It is worth mentioning that since the two-dimensional CNN network is a two-dimensional CNN network with searchable granularity, the two-dimensional CNN network can extract the local temporal features of information at different times of the same sensor and information of different sensors at the same time.

[0100] The above addition module is used to add the global temporal feature and the local temporal feature through the following formula:

[0101] Out C = Out mamba + CNN(Out TPA )

[0102] where, Out C represents the multi-channel data output by the addition module, and CNN(Out TPA ) represents the data output by the two-dimensional CNN network (i.e., the local temporal feature).

[0103] The above sequence fusion attention module is used to fuse and reduce the dimension of the multi-channel data output by the addition module, and the preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module.

[0104] In some embodiments of the present application, the above sequence fusion attention module is mainly used to fuse and reduce the dimension of the multi-channel data output by the addition module through the following formula:

[0105] Out ronghe = mean(Out C * Att, dim = 2)

[0106] Att = AdaptiveAvgpool2d (1,None) split(W)

[0107]

[0108] X in = stack(Out C , dim = 2)

[0109] where, Out rongheThe data representing the output of the sequence fusion attention module, mean represents the operation of averaging for dimensionality reduction, AdaptiveAvgpool2d represents the adaptive average pooling operation, 2d indicates that the input of the adaptive average pooling operation is two-dimensional data, None represents a placeholder, split represents the operation of taking the upper half segment along the dimension dim = 2 (the operation of taking the upper half segment can be understood as only retaining the upper half segment of the data. For example, the data (B*C*H*W) becomes (B*C / 2*H*W), where B is the batch size, C is the number of channels, H is the height, and W is the width), Relu represents the Relu function, Conv represents the convolution operation, X2 represents the size of the convolution kernel for performing this convolution operation (this convolution operation can use a pseudo-two-dimensional convolution of one of the sizes 1*1, 2*1, 3*1, 4*1, 5*1, 6*1. Specifically, the selection of which size of pseudo-two-dimensional convolution can be determined during the optimization process of the model parameters of the bearing fault analysis model), stack represents the stacking operation, and dim represents the dimension.

[0110] The above preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module through the following formula:

[0111] Out pred = Maxpool(Relu(BN(Conv1d(Out ronghe ))))

[0112] Among them, Out pred represents the data output by the preprocessing module, Maxpool represents the max pooling function, Relu represents the Relu function, BN represents the normalization function, Conv1d represents the convolution operation, and 1d indicates that the input of the convolution operation is one-dimensional data.

[0113] In practical applications, due to the small number of sequences, directly padding and then using convolution processing will definitely cause some sequences to fuse with other sequences with a smaller number of other sequences. Therefore, in this application, the data is first copied along the channel and then concatenated along the dimension dim = 2:

[0114] X in = stack(Out C , dim = 2)

[0115] After that, a searchable pseudo-two-dimensional convolution is used for information merging of multiple sequences. Among them, the searchable pseudo-two-dimensional convolutions are Conv2d(1*1), Conv2d(2*1), Conv2d(3*1), Conv2d(4*1), Conv2d(5*1), Conv2d(6*1). Specifically, the selection of which size of pseudo-two-dimensional convolution can be determined during the optimization process of the model parameters of the bearing fault analysis model.

[0116] Next, use Relu and AdaptiveAvgpool2d again (1,None) to obtain the attention coefficient Att and multiply it with the original data:

[0117] W = Relu(Conv (p,1) (X in ))

[0118] Att = AdaptiveAvgpool2d (1,None) split(W)

[0119] Among them, the split operation is to take the upper half segment according to the dimension dim = 2, so the output will be features with the same size as the input Out C of the same size.

[0120] Finally, reduce the dimension of the obtained data by taking the mean according to the dimension dim = 2, and then the final output Out ronghe can be obtained:

[0121] Out ronghe = mean(Out C * Att, dim = 2)

[0122] After the sequence fusion attention module outputs one-dimensional data Out ronghe the one-dimensional data Out ronghe can be preprocessed by the preprocessing module to adjust the sequence length of the module.

[0123] The above HAB attention network is used to extract the temporal information of the data output by the preprocessing module.

[0124] In some embodiments of the present application, the HAB attention network is mainly used to extract the temporal information Out of the data output by the preprocessing module through the following formula T :

[0125] Out T = Out CAB + Out W-MSA + Out pred

[0126] Out CAB = Att CAB * Out pred

[0127] Att CAB = sigmoid(Conv1d(CAB t ))

[0128] CAB t= ReLU(Conv1d(AdaptiveAvgpool1d (1) (Out pred )))

[0129] Out W-MSA = softmax(K * Q) * V

[0130] K, Q, V = FC(Out pred .view(-1, X3, c))

[0131] Where sigmoid represents the sigmoid function, Conv1d represents the convolution operation, 1d means the input of the convolution operation is one-dimensional data, ReLU represents the ReLU function, AdaptiveAvgpool1d (1) represents the shape after adaptive average pooling, softmax represents the softmax function, K represents the K matrix corresponding to the window-based multi-head attention mechanism, Q represents the Q matrix corresponding to the window-based multi-head attention mechanism, V represents the V matrix corresponding to the window-based multi-head attention mechanism, FC represents the fully connected layer, view(-1, X3, c) means reshaping the tensor to (-1, X3, c), -1 means the size of this dimension is automatically calculated by pytorch to ensure the total number of elements remains unchanged, X3 represents the window size, and c represents the input channel dimension. X3 is searchable and can be specifically determined during the optimization process of the model parameters of the bearing fault analysis model.

[0132] In practical applications, the data Out output by the preprocessing module pred is divided into two branches. One branch is input into the channel attention mechanism (CAB) module, which extracts the temporal features of the input data from the channel perspective. What is obtained is the time feature, and it focuses on the global feature extraction:

[0133] CAB t = ReLU(Conv1d(AdaptiveAvgpool1d (1) (Out pred )))

[0134] Att CAB = sigmoid(Conv1d(CAB t ))

[0135] Out CAB = Att CAB * Out pred

[0136] The other branch is input into the window-searchable window-based self-attention mechanism (W-MSA) module. In W-MSA, the input data is divided into windows, and K, Q, and V are calculated for each window. Then, multi-head attention is calculated, followed by transformation, and finally the output is obtained:

[0137] K, Q, V = FC(Out pred .view(-1, X3, c))

[0138] Out W-MSA = softmax(K * Q) * V

[0139] Out T = Out CAB + Out W-MSA + Out pred

[0140] The above output module is used to map the time series information and output the fault analysis result.

[0141] In some embodiments of the present application, the above output module includes an average pooling layer and a fully connected layer connected in sequence. The input end of the average pooling layer is the input end of the output module, and the output end of the fully connected layer is the output end of the output module.

[0142] Among them, the average pooling layer is used to perform a rotation operation on the time series information output by the HAB attention network, and the fully connected layer is mainly used to map the data output by the average pooling layer and output the fault analysis result. The fault analysis result can be: there are outer ring cracks, outer ring pitting, roller cracks, etc. in the high-speed train bearing.

[0143] Next, an exemplary description is given of the process of optimizing the model parameters of the bearing fault analysis model using the neural architecture search network.

[0144] First, action selection is performed through the policy network (actor) model of the reinforcement learning Actor-Critic algorithm. Action selection is performed according to the current states of multiple factors, generating one action at a time and adding it to the action set. The initial action set is [0, 0, 0, 0, 0]. According to the current action set, the corresponding network is generated. If the action set is not filled with 5 elements, the remaining unfilled positions are composed of 0s.

[0145] Then, put the generated action set into the model, and use the accuracy rate after 30 generations of model training as the reward value. Return and record the action probability (actor probably), reward, state and other parameters of the model into the experience pool. If the current actions have been read in, that is, the action set is full, start training. Randomly extract experiences from the experience pool to train the actor and the value (Critic network) respectively. After training is completed, initialize the action set and repeat this process until convergence (where the selected network remains unchanged) or reach the specified number of iterations. Among them, the training data includes sample data of different fault categories, and the sample data includes the fault vibration signals of high-speed train bearings under different fault categories, and the fault category corresponding to each fault vibration signal. Specifically, a sliding window with a length of 1024 can be used to intercept the fault vibration signals without repetition, 80 samples are taken from each fault category, and the total number of samples is 80 * 19. The training set and the validation set are divided according to the ratio of 7:3.

[0146] Exemplarily, in the process of optimizing the model parameters of the bearing fault analysis model using the neural architecture search network, it is possible to continuously explore in the five search spaces shown in Table 1 to determine the model parameters of the bearing fault analysis model. It should be noted that for each search space, Table 1 only shows some model parameters of the bearing fault analysis model.

[0147] Table 1 Model parameters corresponding to five search spaces

[0148]

[0149]

[0150] The bearing fault analysis method of the present application will be exemplarily described below with specific examples.

[0151] In this example, the bearing fault analysis method of the present application is applied to the axle box bearing fault diagnosis experimental platform of the comprehensive test bench of the high-speed train bogie for verification. Specifically, it is applied to the rolling bearing fault analysis. The present application uses the axle box bearing data collected by the high-speed train bogie comprehensive test bench for example verification. Taking the NTN CRI-2692 double-row tapered roller bearing for high-speed trains as the research object, specimens of the outer ring, rollers and cages containing normal, cracked and pitted are selected and installed on the bogie of the entire rolling test bench. The test bench model is as Figure 4 shown, and the sensor layout positions are as Figure 5a and Figure 5bAs shown, different types of bearing faults are respectively installed on an experimental bench with sensors for data acquisition. Three measurement points are arranged on each bearing, with a total of three sensors. The motor speed is at intervals of 500 RPM, and the brake output is at intervals of 20%. Vibration data of the single-side drive system of the experimental bench is collected under the conditions of rotational speeds from 1000 to 3000 RPM and torque loads from 0 to 40%. All experiments use a sampling frequency of 12k, a sampling time of 2s, and three groups of samples are taken for each working condition.

[0152] Based on the collected bearing fault vibration signal dataset, the proposed method is used to analyze the bearing faults. First, the original fault signal of the bearing is obtained and preprocessed by a two-dimensional convolution kernel. Second, the preprocessed signal is input into the time pyramid attention module (TPA) with searchable granularity to perform preliminary temporal features of each channel. Then, the signal is divided into two branches. One branch is input into a pseudo two-dimensional convolution with a searchable convolution kernel size to extract local temporal information of the same sequence, and the other branch is input into a mamba network with a searchable unfolding method to perform information interaction between different sensors and further global temporal extraction. After that, it is input into a sequence fusion attention module based on attention with searchable granularity for dimensionality reduction to facilitate further temporal feature extraction. After adjusting the sequence length through the preprocessing adjustment module, it is finally input into the HAB attention network with searchable scale for global and local feature extraction, and finally fault analysis is performed. Subsequently, a comparative experiment is carried out. By comparing this experimental model (i.e., the bearing fault analysis model proposed in this application) with different models, the excellent performance of this model is verified.

[0153] To better compare the diagnostic accuracy under experimental conditions, some authoritative methods in recent years are reproduced in this experiment. This experimental model is compared with the self-attention mechanism-convolutional neural network (Self-CNN), gated recurrent unit-convolutional neural network (GRU-CNN), long short-term memory network-convolutional neural network (LSTM-CNN), bidirectional long short-term memory network (Bi-LSTM), and Bi-Transformer (Bi-Transformer is a model that combines Bi-LSTM and Transformer). In particular, to more prominently highlight the anti-noise ability of the model, noise with a signal-to-noise ratio of -6 dB is selected for testing. Each method is tested 5 times, and the average value is taken as the experimental result. The experimental settings are uniformly 100 iterations, and the parameter configuration is: the batch size is 32, and the learning rate is 1e-5.

[0154] Finally, the accuracy of the fault analysis is output. The higher the accuracy of the bearing fault analysis, the more accurate its analysis. The comparison results are shown in Table 2. Figures 6a1 to 6f2As shown, in Table 2, Accuracy is the accuracy rate, Precision is the precision rate, Recall is the recall rate, and F1 is the F1 score.

[0155] Table 2 Test Results of Bearing Fault Analysis in Comparative Experiments

[0156]

[0157] In summary, under time-varying working conditions and limited data, the bearing fault analysis method of the present application realizes the preliminary feature extraction of information between channels through the time pyramid attention module with searchable granularity, introduces the mamba network with searchable unfolding method and the two-dimensional CNN network with searchable granularity to realize the information of the same sensor at different times, and the targeted global and local feature extraction of information of different sensors at the same time; through the HAB attention network with searchable scale, the optimization of global feature extraction and local feature extraction is further realized; finally, with the help of the reinforcement learning Actor-Critic algorithm agent, the automatic search of the neural architecture under the interaction of multiple factors of the model is realized to obtain the optimal balance of global and local information, solving the problems of insufficient feature extraction of information between channels and time information of sensors in the domain generalization bearing fault diagnosis method, the imbalance between global and local information in long-distance feature extraction, and the low accuracy of bearing fault diagnosis in a noisy environment.

[0158] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art of the present technology, several improvements and refinements can be made without departing from the principle described in the present application, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A bearing fault analysis method based on multiple attention and Mamba network, characterized in that Including: Obtain the original fault signal of the high-speed train bearing; Use the neural architecture search network to optimize the model parameters of the bearing fault analysis model; Use the bearing fault analysis model with optimized model parameters to process the original fault signal to obtain the fault analysis result of the high-speed train bearing; The bearing fault analysis model includes a two-dimensional convolution kernel, a temporal pyramid attention module, a Mamba network with a searchable unfolding method, a two-dimensional CNN network, an addition module, a sequence fusion attention module, a preprocessing module, a HAB attention network, and an output module; The two-dimensional convolution kernel is used to adjust the number of channels of the original fault signal. The temporal pyramid attention module is used to extract preliminary temporal features from the data output by the two-dimensional convolution kernel. The Mamba network is used to extract the global temporal features of the preliminary temporal features. The two-dimensional CNN network is used to extract the local temporal features of the preliminary temporal features. The addition module is used to perform an addition process on the global temporal features and the local temporal features. The sequence fusion attention module is used to perform fusion and dimensionality reduction on the multi-channel data output by the addition module. The preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module. The HAB attention network is used to extract the temporal information of the data output by the preprocessing module. The output module is used to map the temporal information and output the fault analysis result.

2. The bearing fault analysis method according to claim 1, wherein The time pyramid attention module is used to extract the preliminary temporal features Out through the following formula TPA : C = concat(A1, A2, A3, dim = 1) Out TPA = sigmoid(sigmoid(Relu(conv(C)))) * inupt + input Wherein, input represents the data output by the two-dimensional convolution kernel, None represents a placeholder, X1 represents the searchable look-ahead distance, Adaptiveavgpool2d represents an adaptive average pooling operation, concat represents a concatenation operation, dim represents the dimension, conv represents a convolution operation, Relu represents the Relu function, and sigmoid represents the sigmoid function; The size of A1 is B*C*H*(1×X), the size of A2 is B*C*H*(2×X), the size of A3 is B*C*H*(4×X), and B, C, H are the sizes of the data output by the two-dimensional convolution kernel.

3. The bearing fault analysis method according to claim 2, characterized in that, The Mamba network is used to extract the global temporal feature Out through the following formula mamba :[[]]END]] Input mamba = Out TPA· view i Among them, mamba(·) represents the Mamba network, and view i represents the i-th search order, where i = 1, 2, 3, 4.

4. The bearing fault analysis method according to claim 3, characterized in that The two-dimensional CNN network is a pseudo two-dimensional convolution kernel.

5. The bearing fault analysis method according to claim 4, wherein The addition module is used to perform an addition process on the global temporal features and the local temporal features through the following formula: Out C = Out mamba + CNN(Out TPA ) Among them, Out C represents the multi-channel data output by the addition module, and CNN(Out TPA ) represents the data output by the two-dimensional CNN network.

6. The bearing fault analysis method according to claim 5, wherein The sequence fusion attention module is used to perform fusion and dimensionality reduction on the multi-channel data output by the addition module through the following formula: Out ronghe = mean(Out C * Att, dim = 2) Att = AdaptiveAvgpool2d (1,None) split(W) X in = stack(Out C , dim = 2) Among them, Out ronghe represents the data output by the sequence fusion attention module, mean represents the average reduction operation, AdaptiveAvgpool2d represents the adaptive average pooling operation, None represents a placeholder, split represents the operation of taking the upper half segment according to the dimension dim = 2, Relu represents the Relu function, Conv represents the convolution operation, X2 represents the size of the convolution kernel for performing this convolution operation, stack represents the stacking operation, and dim represents the dimension.

7. The bearing fault analysis method according to claim 6, wherein The preprocessing module is used to preprocess the one-dimensional data output by the sequence fusion attention module through the following formula: Out pred = Maxpool(Relu(BN(Conv1d(Out ronghe )))) Among them, Out pred represents the data output by the preprocessing module, Maxpool represents the max pooling function, Relu represents the Relu function, BN represents the normalization function, and Conv1d represents the convolution operation.

8. The bearing fault analysis method according to claim 7, characterized in that The HAB attention network is used to extract the temporal information Out of the data output by the preprocessing module through the following formula T :[[]]END]] Out T = Out CAB + Out W-MSA + Out pred Out CAB = Att CAB * Out pred Att CAB = sigmoid(Conv1d(CAB t )) CAB t = Relu(Conv1d(AdaptiveAvgpool1d (1) (Out pred ))) Out W-MSA = softmax(K * Q) * V K, Q, V = FC(Out pred· view(-1, X3, c)) Among them, sigmoid represents the sigmoid function, Conv1d represents the convolution operation, Relu represents the Relu function, AdaptiveAvgpool1d (1) represents the shape after adaptive average pooling, softmax represents the softmax function, K represents the K matrix corresponding to the window-based multi-head attention mechanism, Q represents the Q matrix corresponding to the window-based multi-head attention mechanism, V represents the V matrix corresponding to the window-based multi-head attention mechanism, FC represents the fully connected layer, view(-1, X3, c) represents reshaping the tensor into (-1, X3, c), X3 represents the window size, and c represents the input channel dimension.

9. The bearing fault analysis method according to claim 8, wherein, The output module includes an average pooling layer and a fully connected layer connected in sequence. The input end of the average pooling layer is the input end of the output module, and the output end of the fully connected layer is the output end of the output module.

Citation Information

Cited By

  • Predictive maintenance method for precision degradation of machine tool feed shaft based on continuously optimized Mamba network

    CN120634533A

  • A machine tool feed shaft precision degradation predictive maintenance method based on continuous optimization of mamba network

    CN120634533B

  • Equipment cross-domain fault diagnosis method based on multiple sensors and causal contrast decoupling

    CN122490447A

  • Device cross-domain fault diagnosis method based on multi-sensor and causal contrast decoupling

    CN122490447B