Bearing fault diagnosis method based on bidirectional Mama network
By using a two-way Mamba network combined with one-dimensional timing processing network and two-dimensional CNN network in the fault diagnosis of high-speed train bearings, the problem of low fault diagnosis accuracy in the prior art is solved, and more efficient feature extraction and diagnostic accuracy is achieved.
Patent Information
- Application Number
- CN202510234775.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing high-speed train bearing fault diagnosis methods are insufficient when processing massive monitoring data, and it is difficult to maintain high-precision diagnostic effects under different working conditions. The feature extraction is insufficient, which affects the accuracy of fault diagnosis.
The bearing fault diagnosis method based on the bidirectional Mamba network is adopted, and fault feature extraction is extracted by introducing a combination of one-dimensional timing processing network and two-dimensional CNN network. The time attention network and the bidirectional Transformer network are used to further strengthen feature extraction, data compression and feature extraction are performed through the bidirectional Mamba network, and finally fault diagnosis results are generated through the output network.
It improves the accuracy of bearing fault diagnosis, can capture the information of channel sequence and time series more fully, ensures that the feature extraction is sufficient, and enhances the diagnostic ability of the model under different operating conditions.
Smart Images

Figure CN119935555A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of high-speed train bearing fault diagnosis, and in particular, relates to a bearing fault diagnosis method based on a bidirectional Mamba network. Background Art
[0002] In industrial production, intelligent fault diagnosis technology for rotating machinery plays a key role. As one of the core rotating components in the bogie, the bearing of a high-speed train not only bears the weight of the train, but is also directly related to the smoothness and safety of the train operation. Therefore, timely and effective condition monitoring and fault diagnosis of bearings are of vital importance to ensure the safe operation of trains. However, current condition detection and fault diagnosis methods still face many challenges when faced with massive monitoring data.
[0003] For high-speed trains: 1. Trains generate a large amount of monitoring data during operation, and these data contain information on a variety of operating conditions. However, existing diagnostic models often face the problem of insufficient generalization ability when processing these complex data, and it is difficult to maintain high-precision diagnostic results under different operating conditions, which in turn affects the actual application effect and reliability of the fault diagnosis system. 2. The dependencies between massive amounts of data distributed in time and channels are often easily overlooked, which further leads to problems such as insufficient feature extraction and increases the difficulty of fault diagnosis. These characteristics make it difficult for intelligent fault diagnosis models to generalize to data under different operating conditions of the same type of equipment. To this end, how to complete the fault diagnosis task of deep mining the relationship between time and channels of existing data with limited available data is an urgent problem to be solved in intelligent fault diagnosis.
[0004] In order to meet the demand for deep extraction of time and channel relationships in existing sample data, a direct and effective method is to introduce an attention mechanism to automatically focus on important features and improve the accuracy of fault diagnosis. However, the current mainstream attention mechanism methods are usually single channel attention mechanisms or time attention mechanisms, which can only capture the dependencies between channel sequences and time series, resulting in an incomplete understanding of channels and time series by the model, resulting in low accuracy of bearing fault diagnosis. Summary of the invention
[0005] The embodiment of the present application provides a bearing fault diagnosis method based on a bidirectional Mamba network, which can solve the problem of low accuracy in high-speed train bearing fault diagnosis.
[0006] The embodiment of the present application provides a bearing fault diagnosis method based on a bidirectional Mamba network, comprising:
[0007] Collect fault vibration signals of high-speed train bearings;
[0008] The fault vibration signal is input into the fault diagnosis model for analysis and processing to obtain the fault diagnosis result of the high-speed train bearing;
[0009] The fault diagnosis model includes a two-dimensional CNN network, a temporal attention network, a bidirectional Transformer network, a one-dimensional time series processing network, a first fusion network based on channel attention, a data compression network, a bidirectional Mamba network, a second fusion network based on channel attention, and an output network connected in sequence;
[0010] The two-dimensional CNN network is used to extract the characteristic information of the fault vibration signal, the time attention network is used to process the characteristic information output by the two-dimensional CNN network to obtain the attention characteristic information, the bidirectional Transformer network is used to extract the characteristics of the attention characteristic information from both the time and sequence directions, the one-dimensional time series processing network is used to extract the one-dimensional characteristic information of the characteristics output by the bidirectional Transformer network, the first fusion network is used to fuse multiple one-dimensional characteristic information output by the one-dimensional time series processing network, the data compression network is used to compress the fusion result output by the first fusion network, the bidirectional Mamba network is used to extract the characteristics of the compressed data output by the data compression network, the second fusion network is used to fuse multiple features output by the bidirectional Mamba network, and the output network is used to map the fusion result output by the second fusion network to output the fault diagnosis result.
[0011] Optionally, the two-dimensional CNN network includes multiple pseudo-two-dimensional convolution kernels and a first addition module, the input end of each pseudo-two-dimensional convolution kernel receives the fault vibration signal, the output end of each pseudo-two-dimensional convolution kernel is connected to the input end of the first addition module, the output end of the first addition module is connected to the input end of the temporal attention network, and the sizes of the multiple pseudo-two-dimensional convolution kernels are different.
[0012] Optionally, the temporal attention network includes: a first branch network for extracting attention information of different sequences at the same time, a second branch network for extracting attention information of different times in the same sequence, a first multiplication module, a second multiplication module, and a second addition module;
[0013] Among them, the input end of the first branch network and the input end of the second branch network both receive feature information output by the two-dimensional CNN network, the output end of the first branch network and the output end of the second branch network are both connected to the input end of the first multiplication module, the output end of the first multiplication module is connected to the input end of the second multiplication module, the output end of the two-dimensional CNN network is connected to the input end of the second multiplication module, the output end of the second multiplication module and the output end of the two-dimensional CNN network are both connected to the input end of the second addition module, and the output end of the second addition module is connected to the input end of the bidirectional Transformer network.
[0014] Optionally, the first branch network includes a first pooling layer and a first attention information extraction module connected in sequence, and the second branch network includes a second pooling layer and a second attention information extraction module connected in sequence;
[0015] The first attention information extraction module is used to extract the information through the formula A1 = Sigmoid (Relu (W1 * Data t )) obtain the attention information A1 of different sequences at the same time;
[0016] The second attention information extraction module is used to extract the information through the formula A2 = Sigmoid (Relu (W2 * Data c )) obtain the attention information A2 of the same sequence at different times;
[0017] Among them, Sigmoid represents the sigmoid function, Relu represents the Relu function, Data t The data output by the first pooling layer, W1 represents Data t The weight of Data c The data output by the second pooling layer, W2 represents Data c The weight of the first pooling layer is 1*none, and the pooling window size of the second pooling layer is none*1. None is a placeholder.
[0018] Optionally, the one-dimensional time series processing network includes multiple one-dimensional convolution kernels, the input end of each one-dimensional convolution kernel receives the features output by the bidirectional Transformer network, the output end of each one-dimensional convolution kernel is connected to the input end of the first fusion network, and the sizes of the multiple one-dimensional convolution kernels are different.
[0019] Optionally, the first fusion network and the second fusion network both fuse the received data using the following formula and output a fusion result:
[0020] s = Stack(input)
[0021] A c =Sigmoid(Relu(Maxpool(s)+Avgpool(s)))
[0022] stack out =A c *s
[0023] Out = sum(stack out , dim=1)
[0024] Among them, input represents the received data, Stack represents the stacking operation, s represents the data obtained by stacking the input along the channel dimension, Sigmoid represents the sigmoid function, Relu represents the Relu function, Maxpool represents the maximum pooling function, Avgpool represents the average pooling function, and A c Represents the channel attention coefficient, dim represents the dimension, sum represents the summation function, and Out represents the fusion result.
[0025] Optionally, the data compression network is used to perform compression processing on the fusion result output by the first fusion network at multiple multiples to obtain compressed data at multiple multiples.
[0026] Optionally, the bidirectional Mamba network includes a plurality of bidirectional Mamba layers, and an output end of each bidirectional Mamba layer is connected to an input end of the second fusion network;
[0027] The multiple bidirectional Mamba layers are used to receive compressed data at multiple multiples, and the multiple bidirectional Mamba layers correspond one to one to the compressed data at multiple multiples.
[0028] Optionally, the output network includes an average pooling layer, a reshape layer, and a fully connected layer connected in sequence.
[0029] Optionally, the fault vibration signal is collected by a plurality of vibration sensors arranged on the bogie of the high-speed train.
[0030] The above solution of the present application has the following beneficial effects:
[0031] In an embodiment of the present application, a method combining a one-dimensional time series processing network and a two-dimensional CNN network is introduced to extract fault features, so that the information of the channel sequence and the time series can be fully captured; in terms of feature extraction between different channels, a time attention network is introduced to further enhance the information captured by the channel sequence and the time series, thereby ensuring sufficient feature extraction; at the same time, a bidirectional Transformer network is introduced, and a KQV of a multi-head attention mechanism is extracted using a dual-directional convolution operation of channel-time, so that the features of different channels and different times are fully extracted while ensuring that the channel information is not chaotic; by introducing a bidirectional Mamba network, parallel input data with different compression degrees is used to ensure the unity of information loss filling and information accuracy of SSM during long-distance modeling, thereby ensuring that the data in the original fault vibration signal can be fully extracted, thereby enabling the output network to greatly improve the accuracy of fault diagnosis when performing fault diagnosis based on the fully extracted data. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 A flowchart of a bearing fault diagnosis method based on a bidirectional Mamba network provided in an embodiment of the present application;
[0034] Figure 2 A schematic diagram of the structure of a fault diagnosis model provided in one embodiment of the present application;
[0035] Figure 3 is a structural schematic diagram of the test bench model in the example;
[0036] Figure 4a Schematic diagram of sensor placement in the example Figure 1 ;
[0037] Figure 4b Schematic diagram of sensor placement in the example Figure 2 ;
[0038] Figures 5a-5e are schematic diagrams of the confusion matrix and accuracy results of the bearing fault diagnosis method in the example, where a, b, c, d, and e represent the one-dimensional basic model, bidirectional multi-granularity Transformer, two-dimensional preprocessing + channel-time attention, bidirectional multi-granularity Mamba under channel fusion, and the final model of this experiment, respectively. DETAILED DESCRIPTION
[0039] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0040] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0041] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0042] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0043] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0044] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0045] In response to the current problem of low accuracy in high-speed train bearing fault diagnosis, an embodiment of the present application provides a bearing fault diagnosis method based on a bidirectional Mamba network. The method extracts fault features by introducing a combination of a one-dimensional time series processing network and a two-dimensional CNN network, which can fully capture the information of channel sequences and time series; in terms of feature extraction between different channels, a time attention network is introduced to further enhance the capture of channel sequence and time series information, thereby ensuring sufficient feature extraction; at the same time, a bidirectional Transformer network is introduced to extract the KQV of the multi-head attention mechanism using convolution operations in the dual directions of channel-time, thereby fully extracting features of different channels and different times while ensuring that the channel information is not chaotic; by introducing a bidirectional Mamba network, parallel input data with different compression degrees is used to ensure the unity of information loss filling and information accuracy of SSM during long-distance modeling, thereby ensuring that the data in the original fault vibration signal can be fully extracted, thereby enabling the output network to greatly improve the accuracy of fault diagnosis when performing fault diagnosis based on the fully extracted data.
[0046] The bearing fault diagnosis method based on the bidirectional Mamba network provided by the present application is exemplarily described below in conjunction with specific embodiments.
[0047] like Figure 1 As shown, the bearing fault diagnosis method based on the bidirectional Mamba network provided in the embodiment of the present application includes the following steps:
[0048] Step 11, collecting fault vibration signals of high-speed train bearings.
[0049] In some embodiments of the present application, the fault vibration signal may be collected by multiple vibration sensors disposed at different positions on the high-speed train bogie. That is, the fault vibration signal is collected by multiple vibration sensors disposed on the high-speed train bogie.
[0050] It is understandable that, in order to facilitate data processing, a fixed-length sliding window may be used to intercept the fault vibration signal, and the intercepted data may be used for subsequent processing.
[0051] Step 12: Input the fault vibration signal into the fault diagnosis model for analysis and processing to obtain the fault diagnosis result of the high-speed train bearing.
[0052] The above fault diagnosis results may indicate that the high-speed train bearings have outer ring cracks, outer ring pitting, roller cracks, etc.
[0053] It can be understood that the above-mentioned fault diagnosis model is a trained model. That is, before using the fault diagnosis model for fault diagnosis, the model needs to be trained using training data. The training data contains sample data of different fault categories, and the sample data includes the fault vibration signal of the high-speed train bearing under different fault categories, and the fault category corresponding to each fault vibration signal. Specifically, a sliding window with a length of 1024 can be used to perform non-repetitive interception of the fault vibration signal, and 80 samples are taken from each fault category. The total number of samples is 80*19, and the training set, test set, and verification set are divided according to the ratio of 8:1:1. In some optional examples, supervised learning can be used to complete the training of the fault diagnosis model.
[0054] The specific structure of the above fault diagnosis model is exemplified below.
[0055] like Figure 2 As shown, the fault diagnosis model includes a two-dimensional CNN network, a temporal attention network, a bidirectional Transformer network, a one-dimensional time series processing network, a first fusion network based on channel attention, a data compression network, a bidirectional Mamba network, a second fusion network based on channel attention, and an output network connected in sequence.
[0056] Among them, the two-dimensional CNN network is used to extract the characteristic information of the fault vibration signal; the time attention network is used to process the characteristic information output by the two-dimensional CNN network to obtain the attention characteristic information; the bidirectional Transformer network is used to extract the characteristics of the attention characteristic information from both the time and sequence directions; the one-dimensional time series processing network is used to extract the one-dimensional characteristic information of the characteristics output by the bidirectional Transformer network; the first fusion network is used to fuse multiple one-dimensional characteristic information output by the one-dimensional time series processing network; the data compression network is used to compress the fusion result output by the first fusion network; the bidirectional Mamba network is used to extract the characteristics of the compressed data output by the data compression network; the second fusion network is used to fuse multiple features output by the bidirectional Mamba network; the output network is used to map the fusion result output by the second fusion network and output the fault diagnosis result.
[0057] The above-mentioned two-dimensional CNN network is mainly used to process the data (i.e., the fault vibration signal in step 11) by channel and perform preliminary information exchange, and use a multi-granularity two-dimensional CNN to extract features from each channel and output feature information of the fault vibration signal. Specifically, the two-dimensional CNN network includes multiple pseudo-two-dimensional convolution kernels and a first addition module, the input end of each pseudo-two-dimensional convolution kernel receives the fault vibration signal, the output end of each pseudo-two-dimensional convolution kernel is connected to the input end of the first addition module, the output end of the first addition module is connected to the input end of the time attention network, and the sizes of the multiple pseudo-two-dimensional convolution kernels are different from each other.
[0058] Pseudo-two-dimensional convolution kernel: Compared with the 2*2 two-dimensional convolution kernel, the use of 1*3, 3*1, 1*5, 5*1, 1*7, 7*1, 1*9, 9*1 and other convolution kernels of different scales can also achieve the extraction of two-dimensional features, but in essence it is the fusion of the results of two one-dimensional feature extractions. In order to prevent feature confusion between channels, four pseudo-two-dimensional convolution kernels of 1*3, 1*5, 1*7, and 1*9 can be used for extraction.
[0059] In some embodiments of the present application, the above-mentioned multiple pseudo-two-dimensional convolution kernels receive the fault vibration signal of the high-speed train bearing in parallel, and perform preliminary feature extraction on the fault vibration signal of the high-speed train bearing respectively, so as to integrate multi-scale information through convolution kernels of different sizes. As an optional example, the number of pseudo-two-dimensional convolution kernels can be 4. The above-mentioned first addition module is mainly used to add the features extracted by multiple pseudo-two-dimensional convolution kernels to obtain feature information of the fault vibration signal.
[0060] The above-mentioned temporal attention network is mainly used to extract the attention of different time points in the same sequence and the attention of the same time point in different sequences respectively and perform unified processing to ensure the effective capture of the temporal relationship of the same sequence and the dependency between different sequences, thereby avoiding the confusion of sequence information. Specifically, the temporal attention network includes: a first branch network for extracting attention information of different sequences at the same time, a second branch network for extracting attention information of different times in the same sequence, a first multiplication module, a second multiplication module and a second addition module. Among them, the input end of the first branch network and the input end of the second branch network both receive the feature information output by the two-dimensional CNN network, the output end of the first branch network and the output end of the second branch network are both connected to the input end of the first multiplication module, the output end of the first multiplication module is connected to the input end of the second multiplication module, the output end of the two-dimensional CNN network is connected to the input end of the second multiplication module, the output end of the second multiplication module and the output end of the two-dimensional CNN network are both connected to the input end of the second addition module, and the output end of the second addition module is connected to the input end of the bidirectional Transformer network.
[0061] The first branch network includes a first pooling layer and a first attention information extraction module connected in sequence, and the second branch network includes a second pooling layer and a second attention information extraction module connected in sequence. The first pooling layer and the second pooling layer both perform rotation operations on the data input thereto.
[0062] Specifically, the feature information output by the two-dimensional CNN network is input into the first branch network and the second branch network in parallel, so that the first branch network extracts attention information of different sequences at the same time, and the second branch network extracts attention information of different times in the same sequence.
[0063] Among them, the first attention information extraction module is used to extract the information through the formula A1 = Sigmoid (Relu (W1 * Data t )) obtains the attention information A1 of different sequences at the same time; the second attention information extraction module is used to extract the attention information through the formula A2 = Sigmoid (Relu (W2 * Data c )) obtains the attention information A2 of the same sequence at different times. Sigmoid represents the sigmoid function, which compresses the dynamic range of the input activation vector to [0, 1]. Relu represents the Relu function. Data t The data output by the first pooling layer, W1 represents Data t The weight of Data c The data output by the second pooling layer, W2 represents Data c The weight of the first pooling layer is 1*none, and the pooling window size of the second pooling layer is none*1. None is a placeholder, which usually indicates that the size of this dimension is determined by the actual size of the input data.
[0064] None here means that the size of the pooling window in this dimension is determined by the actual size of the input data. Usually, none is used to refer to a dimension that is not fixed, while the size of the other dimension is fixed or determined by design. For example, 100*100 enters (none, 1) and comes out as 100*1.
[0065] The first multiplication module is mainly used to multiply the attention information A1 of different sequences at the same time and the attention information A2 of different times in the same sequence. The second multiplication module is mainly used to multiply the data output by the first multiplication module and the feature information output by the two-dimensional CNN network. The second addition module is mainly used to add the data output by the second multiplication module and the feature information output by the two-dimensional CNN network. Specifically, the expression of the attention feature information Out output by the temporal attention network is:
[0066] Out=Data in +(A1*A2)*Data in
[0067] In the above formula, Data in It is the feature information output by the two-dimensional CNN network.
[0068] In order to better capture the relationship between sequences in two-dimensional data, the multi-head attention mechanism of the above-mentioned bidirectional Transformer network uses unidirectional 2D convolution kernels of various sizes to extract time and sequence information respectively. Specifically, when extracting time information horizontally, a 1×3 convolution kernel is used; when extracting sequence information vertically, a 3×1 convolution kernel is used. This design ensures that time information and sequence information can be effectively extracted at multiple granularity levels, avoiding information mixing.
[0069] It should be noted that when extracting vertical information, since the size of the convolution kernel is close to the number of channels, the data extracted by the convolution operation will contain different amounts of channel information. For example, when a 5×1 convolution kernel is applied to 8-channel data, one part of the data contains 5 channel information, while the other part contains only 3 channel information. To avoid this imbalance, the data is copied and spliced in the channel dimension to ensure the consistency of vertical channel feature extraction.
[0070] Specifically, the process of feature extraction by the bidirectional Transformer network is as follows:
[0071] Step 1.1, extract the KQV of Transformer multi-head attention from the attention feature information output by the temporal attention network at different times in the same sequence through (3*1, 5*1, 7*1) convolution groups;
[0072] Step 1.2, the attention feature information output by the temporal attention network is subjected to (1*3, 1*5, 1*7) convolution groups from different sequences at the same time to extract the KQV of Transformer multi-head attention.
[0073] Step 1.3, under the 7*1 convolution in step 1.1, since some of the extracted data only contain data from 1 channel, and some contain data from 7 channels, the input data is spliced along the channel, changing the data from 9*1024 to 18*1024, and then using the convolution kernel to extract, the first 9 rows of data obtained are used as KQV.
[0074] Step 1.4: Finally, all the results are summarized to get KQV with different channel combinations but the same number of channels:
[0075] K=K t1 +K t2 +K t3 +K c1 +K c2 +K c3
[0076] Q=Q t1 +Q t2 +Q t3 +Qc1 +Q c2 +Q c3
[0077] V=V t1 +V t2 +V t3 +V c1 +V c2 +V c3
[0078] ATT=softmax(K*Q)*V
[0079] In the above formula, K t1 , K t2 , K t3 , Q t1 , Q t2 , Q t3 , V t1 , V t2 , V t3 They are extracted from the convolution kernels in the time direction (1*3, 1*5, 1*7), K c1 , K c2 , K c3 , Q c1 , Q c2 , Q c3 , V c1 , V c2 , V c3 They are extracted from the channel (i.e. sequence) direction (3*1, 5*1, 7*1) convolution kernels, softmax is the softmax function, and ATT is the "attention" matrix (Attention), which determines how to weight the value matrix by calculating the similarity between the query (Query) and the key (Key). Here, ATT is the weighted value matrix after processing by the softmax function.
[0080] The above-mentioned one-dimensional time series processing network is mainly used to extract one-dimensional feature information. This application adopts a two-dimensional one-dimensional fusion network, aiming to give full play to the advantages of the two-dimensional network in information interaction between channels, and at the same time use the powerful ability of the one-dimensional network in extracting single sequence time series relations to build an efficient fault diagnosis network. The network first performs preliminary information preprocessing through a two-dimensional network to ensure sufficient communication of information between channels and perform preliminary time series relationship analysis. Specifically, the above-mentioned one-dimensional time series processing network includes multiple one-dimensional convolution kernels, and the input end of each one-dimensional convolution kernel receives the features output by the bidirectional Transformer network. The multiple one-dimensional convolution kernels process the features output by the bidirectional Transformer network in parallel to obtain multiple one-dimensional feature information. The output end of each one-dimensional convolution kernel is connected to the input end of the first fusion network, and the sizes of the multiple one-dimensional convolution kernels are different. As an optional example, the number of one-dimensional convolution kernels can be 4.
[0081] In some embodiments of the present application, the first fusion network is mainly used to realize intelligent fusion of features. Specifically, the first fusion network fuses the received data (i.e., multiple one-dimensional feature information) through the following formula and outputs the fusion result:
[0082] s = Stack(input)
[0083] A c =Sigmoid(Relu(Maxpool(s)+Avgpool(s)))
[0084] stack out =A c *s
[0085] Out = sum(stack out , dim=1)
[0086] Among them, input represents the received data (i.e., multiple one-dimensional feature information), Stack represents the stacking operation, s represents the data obtained by stacking the input along the channel dimension, Sigmoid represents the sigmoid function, Relu represents the Relu function, Maxpool represents the maximum pooling function, Avgpool represents the average pooling function, and A c Represents the channel attention coefficient, dim represents the dimension, sum represents the summation function, and Out represents the fusion result.
[0087] In some embodiments of the present application, the above-mentioned data compression network is mainly used to perform compression processing of the fusion result output by the first fusion network at multiple multiples to obtain compressed data at multiple multiples.
[0088] The above-mentioned bidirectional Mamba network includes multiple bidirectional Mamba layers, and the output end of each bidirectional Mamba layer is connected to the input end of the second fusion network; multiple bidirectional Mamba layers are used to receive compressed data at multiple multiples, and multiple bidirectional Mamba layers correspond to compressed data at multiple multiples one by one, that is, the data compression network will input compressed data at different multiples to different bidirectional Mamba layers. As a preferred example, the number of bidirectional Mamba layers is 4, and the data compression network compresses the fusion result output by the first fusion network by 2 times, 4 times, 8 times and 16 times, respectively, and inputs the compressed data after 2 times, 4 times, 8 times and 16 times compression processing into 4 bidirectional Mamba layers (each bidirectional Mamba layer only receives compressed data after 1 multiple compression processing) to obtain relatively complete and real data at the same time, and finally the second fusion network integrates the results of the 4 bidirectional Mamba layers as output. It is worth mentioning that the use of the bidirectional Mamba model (i.e., the bidirectional Mamba layer) can comprehensively capture the temporal features in both the positive and negative directions. By processing the forward and reverse input sequences in parallel and fusing and adding their output data, the comprehensive features of the temporal relationship can be extracted and output more accurately.
[0089] The calculation formula for data processing by the bidirectional Mamba layer is:
[0090]
[0091] Among them, Bm out is the output of the bidirectional Mamba layer, Mamba is the Mamba network model, and X in The data that the data compression network outputs to the bidirectional Mamba layer (such as compressed data after 2x compression).
[0092] The second fusion network is mainly used to realize intelligent fusion of features. Specifically, the second fusion network fuses the received data (i.e., the data output by multiple bidirectional Mamba layers) through the following formula and outputs the fusion result:
[0093] s = Stack(input)
[0094] A c =Sigmoid(Relu(Maxpool(s)+Avgpool(s)))
[0095] stack out =A c *s
[0096] Out = sum(stack out , dim=1)
[0097] Where, input represents the received data (i.e., the data output by multiple bidirectional Mamba layers), Stack represents the stacking operation, s represents the data obtained by stacking the input along the channel dimension, Sigmoid represents the sigmoid function, Relu represents the Relu function, Maxpool represents the maximum pooling function, Avgpool represents the average pooling function, and A c Represents the channel attention coefficient, dim represents the dimension, sum represents the summation function, and Out represents the fusion result.
[0098] Specifically, the process of data processing by the one-dimensional time series processing network, the first fusion network based on channel attention, the data compression network, the bidirectional Mamba network, and the second fusion network based on channel attention is as follows:
[0099] Step 2.1, dimensionality reduction is performed through multi-granularity one-dimensional convolution (3, 5, 7, 9), and then the information obtained from multiple granularities is merged using a channel-based attention mechanism;
[0100] Step 2.2, the first fusion network concatenates the four sets of one-dimensional data, then extracts attention from the stacked input at different times of the same channel, multiplies the result with the stacked input, and sums the stacked input according to the channel dimension to obtain the final output. In order to ensure the effect, a residual connection is used, and the sum of the four sets of data is input into the output as a residual connection;
[0101] Step 2.3, using the compression module to process, the data is initially compressed to 1 / 2 using the convolution layer and the maximum pooling layer, and then input into the multi-granularity bidirectional Mamba;
[0102] In step 2.4, the data is compressed by 2 times, 4 times, 8 times, and 16 times respectively, and then input into 4 bidirectional MAMBAs. The resulting 4 sets of one-dimensional data are input into the channel attention-based fusion network again to obtain the fusion result.
[0103] It is worth mentioning that the data is compressed by 2, 4, 8 and 16 times in parallel and input into four different bidirectional Mamba layers to obtain relatively complete and real data at the same time, and finally the results are integrated as output. In order to fully capture the time series features in both the forward and reverse directions, the bidirectional Mamba model is used to process the forward and reverse input sequences in parallel, and the output data of the two are fused and added, so as to more accurately extract and output the comprehensive features of the time series relationship, balance the accuracy of sequence analysis and the integrity of data, and enhance the adaptability and robustness of the model in practical applications.
[0104] In some embodiments of the present application, the output network includes an average pooling layer, a reshaping layer, and a fully connected layer connected in sequence. The average pooling layer is used to rotate the fusion result output by the second fusion network, the reshaping layer is used to adjust the dimension of the result output by the average pooling layer for processing by the fully connected layer, and the fully connected layer is mainly used to map the data output by the reshaping layer and output the fault diagnosis result.
[0105] It is worth mentioning that the present application extracts fault features by introducing a method combining a one-dimensional time series processing network and a two-dimensional CNN network, which can fully capture the information of channel sequences and time series; in terms of feature extraction between different channels, the temporal attention network is introduced to further strengthen the capture of channel sequence and time series information, ensuring sufficient feature extraction; at the same time, by introducing a bidirectional Transformer network, the KQV of the multi-head attention mechanism is extracted using convolution operations in the dual directions of channel-time, and the features of different channels and different times are fully extracted while ensuring that the channel information is not chaotic; by introducing a bidirectional Mamba network, the information loss filling and information accuracy of the selective structured state space model (SSM) during long-distance modeling are unified through parallel input data with different compression degrees, ensuring that the data in the original fault vibration signal can be fully extracted, so that the output network can greatly improve the accuracy of fault diagnosis when performing fault diagnosis based on the fully extracted data.
[0106] The bearing fault diagnosis method based on the bidirectional Mamba network of the present application is exemplified below with reference to specific examples.
[0107] This application is applied to the axle box bearing fault diagnosis experimental platform of the comprehensive test bench of the high-speed train bogie for verification. Specifically applied to rolling bearing fault diagnosis, this application uses the axle box bearing data collected by the comprehensive test bench of the high-speed train bogie for example verification, taking the NTN CRI-2692 double-row tapered roller bearing for high-speed trains as the research object, selecting the outer ring, roller and cage specimens containing normal, cracked and pitted corrosion and installing them on the bogie of the entire rolling test bench. The test bench model is as follows Figure 3 As shown, the sensor layout is as follows Figure 4a and Figure 4b As shown. Different types of bearing faults were installed in a test bench with sensors for data collection. Three measurement points were arranged on each bearing, with a total of three sensors. The motor speed was measured at intervals of 500RPM, and the brake output was measured at intervals of 20%. The vibration data of the single-sided transmission system of the test bench was collected at a speed of 1000-3000RPM and a torque load of 0-40%. All experiments used a sampling frequency of 12k and a sampling time of 2s, with three groups of samples collected for each working condition.
[0108] Based on the acquired bearing fault vibration signal dataset, the proposed method is used to diagnose bearing faults. First, the original fault signal of the bearing is obtained and preprocessed (such as intercepting it using a fixed-length sliding window). Secondly, a two-dimensional CNN based on the channel-time attention mechanism is used for feature extraction to capture the relationship between the signal in the channel sequence and the time series. Then, the bidirectional multi-granularity Transformer is used to re-extract the channel-time series features to further enhance the information captured in the channel sequence and the time series, ensuring that the feature extraction is sufficient. Finally, the signal is compressed at different times and input into the bidirectional multi-granularity Mamba network that integrates the channel attention for fault diagnosis. Then, an ablation experiment is carried out to verify the performance of each module by gradually adding each module to the basic model.
[0109] This experiment uses a one-dimensional basic model as the starting point. The model is mainly composed of a multimodal convolutional layer, a single compression process, and a subsequent pooling layer and a linear connection layer. On this basis, by gradually introducing different modules, the performance of the model is expanded and verified to highlight the superiority of each module. These new modules include: a bidirectional multi-granularity Transformer module (i.e., the bidirectional Transformer network mentioned above), a two-dimensional preprocessing (i.e., the two-dimensional CNN network mentioned above) and a channel-time attention mechanism module (i.e., the temporal attention network mentioned above), and a bidirectional multi-granularity Mamba module under channel fusion (i.e., the bidirectional Mamba network mentioned above). The experimental settings are unified to 100 iterations, and the parameters are configured as a batch size (batch_size) of 32 and a learning rate (learning_rate) of 1e-5.
[0110] The final output is the accuracy of fault diagnosis. The higher the accuracy of bearing fault diagnosis, the more accurate the diagnosis. The comparison results are shown in Table 1 and Figure 5a1 to Figure 5e2 As shown in Table 1, ACC is the accuracy rate, and F1 is the comprehensive evaluation index of precision and recall.
[0111] Table 1 Test results of bearing fault diagnosis in ablation experiment
[0112]
[0113] In summary, the bearing fault diagnosis method based on channel-time attention, bidirectional multi-granularity Transformer and combined with bidirectional multi-granularity Mamba model optimization proposed in this application, under the conditions of time-varying working conditions and limited data, fully mines the existing data by introducing the channel-time attention mechanism and bidirectional multi-granularity Transformer, and optimizes the time series and extraction capabilities through the bidirectional multi-granularity Mamba model, which can solve the problem of low fault diagnosis accuracy caused by insufficient data extraction in the domain generalization bearing fault diagnosis method.
[0114] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A bearing fault diagnosis method based on a bidirectional Mamba network, characterized in that: include: Collect fault vibration signals of high-speed train bearings; Inputting the fault vibration signal into a fault diagnosis model for analysis and processing to obtain a fault diagnosis result of the high-speed train bearing; The fault diagnosis model includes a two-dimensional CNN network, a temporal attention network, a bidirectional Transformer network, a one-dimensional time series processing network, a first fusion network based on channel attention, a data compression network, a bidirectional Mamba network, a second fusion network based on channel attention, and an output network connected in sequence; The two-dimensional CNN network is used to extract feature information of the fault vibration signal, the temporal attention network is used to process the feature information output by the two-dimensional CNN network to obtain attention feature information, the bidirectional Transformer network is used to extract features of the attention feature information from both time and sequence directions, the one-dimensional time series processing network is used to extract one-dimensional feature information of the features output by the bidirectional Transformer network, the first fusion network is used to fuse multiple one-dimensional feature information output by the one-dimensional time series processing network, the data compression network is used to compress the fusion result output by the first fusion network, the bidirectional Mamba network is used to extract features of the compressed data output by the data compression network, the second fusion network is used to fuse multiple features output by the bidirectional Mamba network, and the output network is used to map the fusion result output by the second fusion network to output the fault diagnosis result.
2. The bearing fault diagnosis method according to claim 1, characterized in that: The two-dimensional CNN network includes multiple pseudo-two-dimensional convolution kernels and a first addition module. The input end of each pseudo-two-dimensional convolution kernel receives the fault vibration signal, the output end of each pseudo-two-dimensional convolution kernel is connected to the input end of the first addition module, and the output end of the first addition module is connected to the input end of the temporal attention network. The sizes of the multiple pseudo-two-dimensional convolution kernels are different.
3. The bearing fault diagnosis method according to claim 1, characterized in that: The temporal attention network includes: a first branch network for extracting attention information of different sequences at the same time, a second branch network for extracting attention information of different times in the same sequence, a first multiplication module, a second multiplication module, and a second addition module; Among them, the input end of the first branch network and the input end of the second branch network both receive the feature information output by the two-dimensional CNN network, the output end of the first branch network and the output end of the second branch network are both connected to the input end of the first multiplication module, the output end of the first multiplication module is connected to the input end of the second multiplication module, the output end of the two-dimensional CNN network is connected to the input end of the second multiplication module, the output end of the second multiplication module and the output end of the two-dimensional CNN network are both connected to the input end of the second addition module, and the output end of the second addition module is connected to the input end of the bidirectional Transformer network.
4. The bearing fault diagnosis method according to claim 3, characterized in that: The first branch network includes a first pooling layer and a first attention information extraction module connected in sequence, and the second branch network includes a second pooling layer and a second attention information extraction module connected in sequence; The first attention information extraction module is used to extract the information through the formula A1=Sigmoid(Relu(W1*Data t )) Obtain attention information A1 of different sequences at the same time; The second attention information extraction module is used to extract the information by the formula A2=Sigmoid(Relu(W2*Data c )) Obtain attention information A2 at different times in the same sequence; Among them, Sigmoid represents the sigmoid function, Relu represents the Relu function, Data t The data output by the first pooling layer, W1 represents Data t The weight of Data c The data output by the second pooling layer, W2 represents Data c The weight of the first pooling layer is 1*none, the pooling window size of the second pooling layer is none*1, and none is a placeholder.
5. The bearing fault diagnosis method according to claim 1, characterized in that: The one-dimensional time series processing network includes multiple one-dimensional convolution kernels, the input end of each one-dimensional convolution kernel receives the features output by the bidirectional Transformer network, the output end of each one-dimensional convolution kernel is connected to the input end of the first fusion network, and the sizes of the multiple one-dimensional convolution kernels are different.
6. The bearing fault diagnosis method according to claim 1, characterized in that: The first fusion network and the second fusion network both fuse the received data using the following formula and output a fusion result: s=Stack(input) A c =Sigmoid(Relu(Maxpool(s)+Avgpool(s))) stack out =A c *s Out=sum(stack out ,dim=1) Among them, input represents the received data, Stack represents the stacking operation, s represents the data obtained by stacking the input along the channel dimension, Sigmoid represents the sigmoid function, Relu represents the Relu function, Maxpool represents the maximum pooling function, Avgpool represents the average pooling function, and A c Represents the channel attention coefficient, dim represents the dimension, sum represents the summation function, and Out represents the fusion result.
7. The bearing fault diagnosis method according to claim 1, characterized in that: The data compression network is used to perform compression processing of the fusion result output by the first fusion network at multiple times to obtain compressed data at multiple times.
8. The bearing fault diagnosis method according to claim 7, characterized in that: The bidirectional Mamba network includes a plurality of bidirectional Mamba layers, and an output end of each bidirectional Mamba layer is connected to an input end of the second fusion network; The multiple bidirectional Mamba layers are used to receive the compressed data at the multiple multiples, and the multiple bidirectional Mamba layers correspond one-to-one to the compressed data at the multiple multiples.
9. The bearing fault diagnosis method according to claim 1, characterized in that: The output network includes an average pooling layer, a reshaping layer, and a fully connected layer connected in sequence.
10. The bearing fault diagnosis method according to claim 1, characterized in that: The fault vibration signal is collected by a plurality of vibration sensors arranged on the bogie of the high-speed train.
Citation Information
Patent Citations
Train bearing fault diagnosis method, system and device based on data fusion and medium
CN116858540A
Lithium battery thermal early warning method based on multi-mode BiLSTM-Mama
CN118587159A
Epilepsy prediction method based on adaptive space-time attention and dynamic fusion network
CN119344751A
Remote sensing image crop classification method based on Mama
CN119418141A
Multi-scale feature fusion triple branch network method for multi-organ segmentation
CN119478404A
Cited By
Predictive maintenance method for precision degradation of machine tool feed shaft based on continuously optimized Mamba network
CN120634533A