Improved bidirectional Mama and empirical mode decomposition aero-engine bearing fault diagnosis method
Through improved bidirectional Mamba network and empirical modal decomposition technology, the problem of insufficient non-stationary signal processing and long-range timing dependency in aircraft engine bearing fault diagnosis is solved, efficient and accurate fault diagnosis is achieved, and the generalization ability and computing efficiency of the model are improved.
Patent Information
- Application Number
- CN202510142415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
In the fault diagnosis of aircraft engine bearings, the problems of limited non-stationary signal processing capabilities, insufficient long-range timing dependency capture, feature extraction and classification separation, insufficient model generalization capabilities and low computing efficiency.
The improved bidirectional Mamba network structure and empirical modal decomposition technology are adopted, combined with long and short-term memory units and dynamic weight decay strategies, an end-to-end fault diagnosis model is built, inherent modal functions are extracted through empirical modal decomposition, and a bidirectional long-range timing-dependent feature extraction module is built, and the improved bidirectional Mamba network is used for feature fusion and softmax classifier output diagnostic results.
It significantly improves the processing capability of non-stationary signals, enhances the capture of long-range timing dependencies, realizes end-to-end learning of feature extraction and classification, improves the generalization ability and computing efficiency of the model, and improves the accuracy and reliability of diagnosis.
Smart Images

Figure CN120063732A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and particularly to an improved method for diagnosing faults of an aero-engine bearing by using bidirectional Mamba and empirical mode decomposition. Background Art
[0002] As the core power system of an aircraft, the reliability and safety of an aero-engine are directly related to the safety of air transportation. Among them, as a key component of the aero-engine, the health state of the bearing has a decisive impact on the overall performance and service life of the engine. Therefore, accurately and timely diagnosing the fault state of the aero-engine bearing has always been a key issue concerned by the aviation industry and academia.
[0003] In recent years, with the rapid development of signal processing technology and artificial intelligence algorithms, significant progress has been made in the field of aero-engine bearing fault diagnosis. Traditional fault diagnosis methods mainly rely on expert experience and signal processing technology, such as time-domain analysis, frequency-domain analysis, and time-frequency analysis, etc. These methods perform well in dealing with simple faults, but when facing complex working environments and diverse fault modes, their diagnostic accuracy and reliability often fail to meet the actual requirements.
[0004] In order to improve the diagnostic accuracy, researchers have tried to introduce machine learning technology into bearing fault diagnosis. Shallow learning methods such as support vector machine (SVM) and random forest (RF) have been widely used, but these methods still have limitations in dealing with high-dimensional and non-linear bearing vibration signals.
[0005] Recently, the rise of deep learning technology has brought new opportunities for bearing fault diagnosis. Deep learning models such as convolutional neural network (CNN) and recurrent neural network (RNN) have demonstrated superior performance in bearing fault diagnosis tasks. These methods can automatically learn the feature representation of signals, reduce the workload of manual feature extraction, and improve the diagnostic accuracy at the same time. However, these methods still have some problems:
[0006] 1. Limited ability to process non-stationary signals: Aero-engine bearings often face complex and changing working conditions during actual operation, and the generated vibration signals are usually non-linear and non-stationary. Traditional deep learning methods perform poorly in dealing with such signals.
[0007] 2. Insufficient capture of long-range temporal dependence: The development of bearing faults is usually a gradual process, which contains long-term temporal dependence. Although existing RNN-based models can process sequence data, they still have deficiencies in capturing long-range dependence.
[0008] 3. Separation of Feature Extraction and Classification: Many methods treat feature extraction and fault classification as two independent steps, which may result in the extracted features not being well-matched to the final classification task.
[0009] 4. Insufficient Model Generalization Ability: When facing different working conditions and different types of bearing faults, the generalization ability of existing models is often not ideal, and it is difficult to adapt to the diversity and complexity of the actual industrial environment.
[0010] 5. Room for Improvement in Computational Efficiency: Although some complex deep learning models have high diagnostic accuracy, their high computational complexity makes it difficult to meet the requirements of real-time diagnosis. Summary of the Invention
[0011] In view of the above problems, the present invention proposes an improved bidirectional Mamba and empirical mode decomposition method for aeroengine bearing fault diagnosis. This method realizes the efficient analysis of aeroengine bearing vibration signals and accurate fault diagnosis by innovatively combining an improved bidirectional Mamba network structure and empirical mode decomposition technology.
[0012] The present invention proposes an improved bidirectional Mamba and empirical mode decomposition method for aeroengine bearing fault diagnosis, including:
[0013] An acquisition step, including:
[0014] Obtain the vibration signal data of the aeroengine bearing;
[0015] A processing step, including:
[0016] Perform empirical mode decomposition on the vibration signal data to extract intrinsic mode functions;
[0017] According to the intrinsic mode functions, construct a feature extraction module containing bidirectional long-range temporal dependence information;
[0018] Based on the output of the feature extraction module, construct a residual sub-network containing empirical mode decomposition technology;
[0019] According to the output of the residual sub-network, perform feature fusion using an improved bidirectional Mamba network;
[0020] An output step, including:
[0021] Based on the feature fusion result, use a softmax classifier to output the fault diagnosis result of the aeroengine bearing.
[0022] Preferably, the empirical mode decomposition specifically includes:
[0023] Iteratively screen the vibration signal data to obtain multiple intrinsic mode functions;
[0024] Calculate the adaptive weights of each of the intrinsic mode functions;
[0025] Reconstruct the vibration signal data based on the adaptive weights and the intrinsic mode functions.
[0026] Preferably, the construction of the feature extraction module including two-way long-range temporal dependence information specifically includes:
[0027] Construct a forward feature extraction sub-module to connect each input element with all the elements before it;
[0028] Construct a backward feature extraction sub-module to connect each input element with all the elements after it;
[0029] Perform a convolution operation on the outputs of the forward feature extraction sub-module and the backward feature extraction sub-module to obtain the final two-way features.
[0030] Preferably, both the forward feature extraction sub-module and the backward feature extraction sub-module are composed of multiple residual blocks, and the calculation formula for each residual block is:
[0031] F(x) = ReLU(W 2 ·ReLU(W 1 ·x + b 1 ) + b 2 ) + x,
[0032] where x is the input, W 1 , W 2 are weight matrices, b 1 , b 2 are bias vectors, and ReLU is the activation function.
[0033] Preferably, the construction of the residual sub-network including the empirical mode decomposition technique specifically includes:
[0034] Perform empirical mode decomposition on the input features to obtain multiple intrinsic mode functions;
[0035] Construct a residual connection to add the original input features to the result of the empirical mode decomposition;
[0036] Use a dense connection layer to further process the above result to obtain the output of the residual sub-network.
[0037] Preferably, the improved two-way Mamba network includes:
[0038] A forward Mamba module to process the temporal information from the past to the present;
[0039] A backward Mamba module to process the temporal information from the future to the present;
[0040] A feature fusion layer that fuses the outputs of the forward Mamba module and the backward Mamba module.
[0041] Preferably, the improved bidirectional Mamba network further includes a long short-term memory unit, and its calculation formula is:
[0042] h(t) = LSTM(Mamba(X(t)), h(t - 1)),
[0043] where X(t) is the input at time t, h(t) is the hidden state at time t, and Mamba and LSTM represent the Mamba operation and the long short-term memory operation respectively.
[0044] Preferably, it further includes a loss function calculation step:
[0045] Calculate the cross-entropy loss;
[0046] Calculate the regularization loss based on label smoothing;
[0047] Perform a weighted combination of the cross-entropy loss and the regularization loss to obtain the final loss function.
[0048] Preferably, the calculation formula of the loss function is:
[0049]
[0050] where CE is the cross-entropy function, KL is the KL divergence function, y is the true label, is the predicted label, u is the uniform distribution, and ∈ is the smoothing parameter.
[0051] Preferably, it further includes a dynamic weight decay step:
[0052] Initialize the weight decay parameter λ 0 ;
[0053] During the training process, dynamically adjust the weight decay parameter λ(t) according to the current iteration number t and the total iteration number T;
[0054] The calculation formula for the dynamic adjustment is:
[0055]
[0056] where λ(t) is the weight decay parameter at time t, λ 0 is the initial weight decay parameter, t is the current iteration number, and T is the total iteration number.
[0057] The method of the present invention has significant advantages and beneficial effects in the following aspects:
[0058] 1. Improved the processing ability for non-stationary signals: By introducing the empirical mode decomposition technique, this method can effectively process non-linear and non-stationary bearing vibration signals and extract more effective fault features.
[0059] 2. Enhanced the capture of long-range temporal dependencies: The improved bidirectional Mamba network structure can consider both forward and backward temporal information simultaneously, greatly enhancing the modeling ability for long-range dependencies and facilitating the capture of the progressive process of bearing faults.
[0060] 3. Achieved end-to-end learning for feature extraction and classification: This method integrates feature extraction and fault classification in a unified framework and realizes the optimal matching of feature and classification tasks through end-to-end learning.
[0061] 4. Improved the generalization ability of the model: By introducing innovative loss function design and dynamic weight decay strategy, this method significantly improves the generalization ability of the model in different working conditions and different types of bearing fault diagnosis tasks.
[0062] 5. Optimized the computational efficiency: Compared with traditional deep learning methods, the improved Mamba network structure has higher computational efficiency and is more suitable for the requirements of real-time fault diagnosis.
[0063] 6. Improved the accuracy and reliability of diagnosis: Through the synergistic effect of the above innovative points, this method has achieved significant improvement in the accuracy and reliability of aero-engine bearing fault diagnosis, especially performing well in dealing with complex working conditions and multiple types of faults.
[0064] In summary, the method proposed in the present invention not only solves the main problems existing in the prior art, but also has made significant progress in terms of diagnosis performance, generalization ability, and computational efficiency. This has important practical significance for improving the reliability and safety of aero-engines, reducing maintenance costs, and extending the engine life. At the same time, the idea and technology of this method can also be extended and applied to the fault diagnosis field of other industrial equipment, having broad application prospects. Brief Description of the Drawings
[0065] Figure 1 It is the overall flowchart of the present invention.
[0066] Figure 2 It is the empirical mode decomposition process of the present invention.
[0067] Figure 3 It is the bidirectional long-range temporal dependence feature extraction module of the present invention.
[0068] Figure 4 It is the improved bidirectional Mamba network feature fusion of the present invention.
[0069] Figure 5 This is the dynamic weight decay strategy of the present invention.
[0070] Figure 6 This is the fault diagnosis flowchart of the present invention. Detailed implementation manners
[0071] Please refer to the appended Figure 1-6 , the present invention provides an improved bidirectional Mamba and empirical mode decomposition method for aeroengine bearing fault diagnosis. This method realizes the efficient analysis of the vibration signals of aeroengine bearings and accurate fault diagnosis by innovatively combining an improved bidirectional Mamba network structure and empirical mode decomposition technology.
[0072] Specifically, the method of the present invention includes the following steps:
[0073] First, in the acquisition step, this method acquires the vibration signal data of the aeroengine bearing. These vibration signal data are usually collected by sensors installed at key parts of the aeroengine. Preferably, the sampling frequency can be set in the range of 20 kHz to 100 kHz to ensure capturing the high-frequency characteristics of bearing faults.
[0074] Next, in the processing step, this method first performs empirical mode decomposition (EMD) based on the acquired vibration signal data to extract the intrinsic mode functions (IMFs). EMD is an adaptive signal decomposition method, especially suitable for processing non-linear and non-stationary signals. In an embodiment of the present invention, the number of iterations of EMD can be set to 10 - 15 times to balance the calculation efficiency and decomposition accuracy.
[0075] The process of empirical mode decomposition can be expressed as:
[0076]
[0077] Among them, x(t) is the original signal, c i (t) is the i-th IMF, r n (t) is the residual term, and n is the total number of IMFs. After extracting the IMFs, this method constructs a feature extraction module containing bidirectional long-range temporal dependence information based on these MFs. The innovation of this module lies in its ability to capture both the forward and backward temporal dependence relationships of the signal. Specifically, this module contains two sub-modules, a forward sub-module and a backward sub-module, and each sub-module consists of multiple residual blocks.
[0078] The output of the forward sub-module can be expressed as:
[0079] F f (X) = Conv(Res 1 (Res 2 (...Resn (X))))
[0080] The output of the backward sub-module can be expressed as:
[0081] F b (X) = Conv(Res n (Res n-1 (...Res 1 (X))))
[0082] where X is the input sequence, Res i represents the i-th residual block, and Conv represents the convolution operation. The final bidirectional feature is obtained by fusing the forward and backward features:
[0083] F(X) = F f (X) + F b (X)
[0084] This bidirectional structure can effectively capture the long-range dependencies in the bearing vibration signal and improve the extraction effect of fault features.
[0085] After feature extraction, this method constructs a residual sub-network containing empirical mode decomposition technology based on the output of the feature extraction module. This innovative network structure combines EMD technology with deep learning and can better handle complex non-linear features. The output of the residual sub-network can be expressed as:
[0086] G(X) = Dense(EMD(X) + Res(X))
[0087] where EMD(X) represents performing empirical mode decomposition on the input X, Res(X) represents the residual connection, and Dense represents the fully connected layer.
[0088] Next, this method uses an improved bidirectional Mamba network for feature fusion. The Mamba network is a new type of sequence modeling architecture, and the present invention improves it to adapt to the bearing fault diagnosis task. The improved bidirectional Mamba network includes two Mamba modules, a forward one and a backward one, and a feature fusion layer.
[0089] The output of the forward Mamba module can be expressed as:
[0090] H f = Mamba f (G(X))
[0091] The output of the backward Mamba module can be expressed as:
[0092] H b = Mamba b (G(X))
[0093] The output of the feature fusion layer is:
[0094] H = Fusion(H f , H b ),
[0095] where Fusion represents the feature fusion operation, which can be a simple addition or a more complex attention mechanism. Finally, in the output step, based on the result of feature fusion, this method uses a softmax classifier to output the fault diagnosis result of the aero-engine bearing. The output of the softmax classifier can be expressed as:
[0096]
[0097] where y i represents the probability of the i-th type of fault, z i represents the original score of the i-th type, and K is the total number of fault categories.
[0098] Through the above steps, the method of the present invention can effectively analyze the vibration signal of the aero-engine bearing and accurately diagnose the possible faults. Compared with the traditional method, this method has significant advantages in dealing with non-linear and non-stationary signals and can better adapt to complex working environments and diverse fault modes.
[0099] In practical applications, this method can be integrated into the health monitoring system of the aero-engine to achieve real-time fault diagnosis and early warning. This is of great significance for improving aviation safety and reducing maintenance costs. In the future, with the further development of deep learning technology and signal processing technology, this method still has great room for optimization and expansion, and can further improve the accuracy and real-time performance of diagnosis.
[0100] In a preferred embodiment of the present invention, both the forward feature extraction sub-module and the backward feature extraction sub-module are composed of multiple residual blocks. This design of the residual structure helps to solve the problem of gradient disappearance in the training of deep networks, so that the network can more effectively learn complex feature representations. The calculation formula of each residual block is:
[0101] F(x) = ReLU(W 2 * ReLU(W 1 * x + b 1 ) + b 2 ) + x,
[0102] where x is the input, W 1 and W 2 are weight matrices, b 1 and b 2is the bias vector, and ReLU is the activation function. This structure allows information to flow directly through, which helps to maintain the stability of the gradient.
[0103] Preferably, W in this method 1 and W 2 can be initialized to a normal distribution with a mean of 0 and a standard deviation of where n is the number of input neurons. This initialization method, usually called He initialization, helps to maintain an appropriate gradient scale in the deep network.
[0104] In practical applications, the number of residual blocks can be adjusted according to the specific bearing fault diagnosis task. Generally, for complex fault patterns, the number of residual blocks can be increased to improve the model's expressive ability. For example, when dealing with the fault diagnosis of high-speed bearings, 8 - 12 residual blocks can be used; while for low-speed bearings, 4 - 6 residual blocks may be sufficient.
[0105] The method of the present invention adopts an innovative design when constructing the residual sub-network containing the empirical mode decomposition technique. Specifically, the construction of the network includes the following steps:
[0106] First, perform empirical mode decomposition on the input features to obtain multiple intrinsic mode functions (IMFs). In an embodiment of the present invention, the number of IMFs can be set to 4 - 6, which can usually fully capture the multi-scale features of the signal.
[0107] Next, construct a residual connection by adding the original input features to the result of the empirical mode decomposition. This design allows the network to utilize both the original features and the decomposed features simultaneously, improving the richness of feature representation. The expression of the residual connection can be written as:
[0108]
[0109] where X is the original input feature, IMF i is the i-th intrinsic mode function, and n is the total number of IMFs.
[0110] Finally, further process the above result using a dense connection layer to obtain the output of the residual sub-network. The use of the dense connection layer can enhance the interaction between features and improve the model's non-linear expressive ability. The output of the dense connection layer can be expressed as:
[0111] Y = σ(W * X res + b),
[0112] where σ is the activation function (such as ReLU), W is the weight matrix, and b is the bias vector.
[0113] In the method of the present invention, the improved bidirectional Mamba network is a key innovation. The network includes the following main components:
[0114] 1. Forward Mamba module: This module processes the temporal information from the past to the present. It can effectively capture the long-range dependencies in the bearing vibration signals, especially effective for the detection of progressive faults.
[0115] 2. Backward Mamba module: This module processes the temporal information from the future to the present. This backward processing can capture some predictive features in the signals, which is of great significance for early fault warning.
[0116] 3. Feature fusion layer: This layer fuses the outputs of the forward Mamba module and the backward Mamba module. The choice of the fusion strategy has an important impact on the final diagnostic effect.
[0117] In the preferred embodiment of the present invention, the feature fusion adopts a weighted summation method:
[0118] H = α * H f +(1 - α) * H b ,
[0119] where H f and H b are the outputs of the forward and backward Mamba modules respectively, and α is the weight coefficient. The value of α can be automatically learned through network training or preset according to specific application scenarios. In addition, the method of the present invention also introduces long short-term memory units in the improved bidirectional Mamba network, further enhancing the network's ability to process temporal information. The calculation formula of the long short-term memory unit is:
[0120] h(t) = LSTM(Mamba(X(t)), h(t - 1)),
[0121] where X(t) is the input at time t, h(t) is the hidden state at time t, and Mamba and LSTM represent the Mamba operation and the long short-term memory operation respectively.
[0122] The design of this structure combines the advantages of Mamba in modeling long-range dependencies and the ability of LSTM in processing short-term dependencies, enabling the network to more comprehensively capture the temporal features in the bearing vibration signals. In practical applications, this structure shows good performance for detecting both instantaneous and progressive faults of bearings.
[0123] Through the above design, the method of the present invention can effectively extract and utilize the multi-scale time series features in the bearing vibration signal, significantly improving the accuracy and reliability of fault diagnosis. Especially when dealing with bearing faults under complex working conditions, this method demonstrates obvious advantages.
[0124] In the training process of the method of the present invention, an innovative loss function calculation step is introduced, aiming to improve the generalization ability and robustness of the model. Specifically, this step includes calculating the cross-entropy loss, the regularization loss based on label smoothing, and combining these two losses with weights to obtain the final loss function.
[0125] In practical applications, the cross-entropy loss is a commonly used metric to measure the difference between the model's prediction results and the true labels. However, simply using the cross-entropy loss may lead to overfitting of the model to the training data, especially in tasks such as aero-engine bearing fault diagnosis where the samples may be imbalanced. To address this issue, the method of the present invention introduces a regularization loss based on label smoothing.
[0126] Preferably, the calculation formula of the loss function of the present invention is as follows:
[0127]
[0128] where CE is the cross-entropy function, KL is the KL divergence function, y is the true label, is the predicted label, u is the uniform distribution, and ε is the smoothing parameter.
[0129] In this formula, the cross-entropy term ensures that the model can learn the correct label information, while the KL divergence term introduces a certain regularization effect to prevent the model from being overconfident in certain categories. The smoothing parameter ε is used to adjust the relative importance of these two terms.
[0130] In an embodiment of the present invention, ε can be set to 0.1. This value usually achieves a good balance between model performance and regularization effect. However, the specific value of ε may need to be fine-tuned according to the actual bearing fault diagnosis task. For example, for cases with fewer samples or severely imbalanced categories, the value of ε can be appropriately increased (such as 0.2) to enhance the regularization effect.
[0131] In addition, the method of the present invention also introduces a dynamic weight decay step to further improve the generalization ability of the model. The core idea of this step is to dynamically adjust the weight decay parameter during the training process, enabling the model to learn quickly in the initial stage of training and pay more attention to fine-tuning in the later stage of training.
[0132] Specifically, the dynamic weight decay step includes the following:
[0133] First, initialize the weight decay parameter λ 0 . The choice of this initial value has an important impact on the training effect of the model. In the preferred embodiment of the present invention, λ 0 can be set to 0.01. This value can usually achieve a good balance between the training speed and the model stability. Next, during the training process, according to the current iteration number t and the total iteration number T, dynamically adjust the weight decay parameter λ(t). The adjustment calculation formula is:
[0134]
[0135] where λ(t) is the weight decay parameter at time t, λ 0 is the initial weight decay parameter, t is the current iteration number, and T is the total iteration number.
[0136] This dynamic adjustment strategy makes the weight decay parameter maintain a relatively large value in the initial stage of training, which is beneficial to the rapid convergence of the model; while in the later stage of training, the weight decay parameter gradually decreases, allowing the model to perform more refined parameter adjustment. This strategy is particularly suitable for tasks such as aeroengine bearing fault diagnosis that require considering both large-scale features and subtle differences at the same time.
[0137] In practical applications, the choice of the total iteration number T needs to be determined according to the specific dataset size and model complexity. For example, for a bearing fault dataset containing 10,000 samples, when using a batch size of 64, T can be set to 5000 - 10000 epochs. This can ensure that the model has enough time to learn the complex patterns in the data and at the same time avoid overfitting.
[0138] By introducing this dynamic weight decay strategy, the method of the present invention can adaptively adjust the model parameters during the training process, effectively improving the generalization ability of the model in different types of bearing fault diagnosis tasks. Especially for those cases with complex fault patterns and unbalanced sample distributions, this strategy shows obvious advantages.
[0139] In summary, the improved bidirectional Mamba and empirical mode decomposition aeroengine bearing fault diagnosis method proposed by the present invention realizes the efficient analysis and accurate diagnosis of bearing vibration signals by innovatively combining advanced deep learning technologies and signal processing technologies. The method has made innovations in multiple aspects such as feature extraction, network structure design, and loss function optimization, significantly improving the accuracy and reliability of fault diagnosis. This has important practical significance for improving aviation safety and reducing maintenance costs.
[0140] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Improved bidirectional Mamba and empirical mode decomposition aero-engine bearing fault diagnosis method, characterized in that: include: The acquisition steps include: Obtain vibration signal data of aircraft engine bearings; Processing steps include: Based on the vibration signal data, performing empirical mode decomposition to extract intrinsic mode functions; According to the intrinsic mode function, constructing a feature extraction module containing bidirectional long-range temporal dependency information; Based on the output of the feature extraction module, construct a residual sub-network including empirical mode decomposition technology; According to the output of the residual sub-network, using an improved bidirectional Mamba network to perform feature fusion; Output steps include: Based on the feature fusion result, a softmax classifier is used to output the fault diagnosis result of the aircraft engine bearing.
2. The method according to claim 1, characterized in that The empirical mode decomposition specifically includes: Iteratively screening the vibration signal data to obtain a plurality of intrinsic mode functions; Calculating an adaptive weight for each of the intrinsic mode functions; The vibration signal data is reconstructed based on the adaptive weights and the natural mode functions.
3. The method according to claim 1, characterized in that The construction of the feature extraction module containing bidirectional long-range temporal dependency information specifically includes: Construct a forward feature extraction submodule that concatenates each input element with all its previous elements; Construct a backward feature extraction submodule to connect each input element to all elements after it; The outputs of the forward feature extraction submodule and the backward feature extraction submodule are convolved to obtain the final bidirectional features.
4. The method according to claim 3, characterized in that The forward feature extraction submodule and the backward feature extraction submodule are both composed of a plurality of residual blocks, and the calculation formula of each residual block is: F(x)=ReLU(W2·ReLU(W1·x+b1)+b2)+x, Among them, x is the input, W1 and W2 are weight matrices, b1 and b2 are bias vectors, and ReLU is the activation function.
5. The method according to claim 1, characterized in that The construction of the residual subnetwork including the empirical mode decomposition technology specifically includes: Perform empirical mode decomposition on the input features to obtain multiple intrinsic mode functions; Construct a residual connection to add the original input features to the results of empirical mode decomposition; The above results are further processed using a densely connected layer to obtain the output of the residual sub-network.
6. The method according to claim 1, characterized in that The improved two-way Mamba network includes: The forward Mamba module processes temporal information from the past to the present; The backward Mamba module processes the temporal information from the future to the present; The feature fusion layer fuses the outputs of the forward Mamba module and the backward Mamba module.
7. The method according to claim 6, characterized in that The improved bidirectional Mamba network also includes a long short-term memory unit, and its calculation formula is: h(t)=LSTM(Mamba(X(t)),h(t-1)), Among them, X(t) is the input at time t, h(t) is the hidden state at time t, Mamba and LSTM represent Mamba operation and long short-term memory operation respectively.
8. The method according to claim 1, characterized in that It also includes the loss function calculation step: Calculate cross entropy loss; Calculate the regularization loss based on label smoothing; The cross entropy loss and the regularization loss are weightedly combined to obtain the final loss function.
9. The method according to claim 8, characterized in that The calculation formula of the loss function is: Among them, CE is the cross entropy function, KL is the KL divergence function, y is the true label, is the predicted label, u is a uniform distribution, and ∈ is a smoothing parameter.
10. The method according to claim 1, characterized in that A dynamic weight decay step is also included: Initialize weight decay parameter λ0; During the training process, the weight decay parameter λ(t) is dynamically adjusted according to the current number of iterations t and the total number of iterations T; The calculation formula for the dynamic adjustment is: Among them, λ(t) is the weight decay parameter at time t, λ0 is the initial weight decay parameter, t is the current number of iterations, and T is the total number of iterations.
Citation Information
Cited By
Aero-engine fault diagnosis method based on improved bidirectional Mama
CN120804951A