Fault diagnosis method based on adaptive time-frequency fusion gated attention network

CN122333110BActive Publication Date: 2026-08-07JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGNAN UNIV
Filing Date
2026-06-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]为了解决目前采用多传感器融合在轴承故障诊断中,主要依赖固定的融合结构或预定义融合策略,对多源信号间深层耦合关系的刻画仍不充分,及单纯依赖时域分析或固定频带划分的时频分析,难以完整提取故障敏感信息,导致关键信息损失,最终影响诊断精度的问题,本发明提供了一种基于自适应时频融合门控注意力网络的故障诊断方法,所述技术方案如下:

Benefits of technology

本申请提出一种基于自适应时频融合门控注意力网络的故障诊断方法,通过构建自适应时频域融合网络模型(ATFFNet),用于多传感器轴承故障诊断,该模型构建时域分支与时频分支协同特征提取框架,实现多源信息的联合建模;同时在时频分支中引入自适应频带划分与多尺度特征增强机制,以提高对复杂频率成分和局部退化特征的表征能力,提升模型对非平稳信号中时频特征及弱故障特征的表征能力;并且设计门控注意力融合模块,实现跨分支、跨频带信息的有效交互与自适应聚合,从而提升模型在多传感器场景下的诊断精度、鲁棒性与泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333110B_ABST
    Figure CN122333110B_ABST
Patent Text Reader

Abstract

The application discloses a fault diagnosis method based on an adaptive time-frequency fusion gated attention network and belongs to the technical field of equipment fault diagnosis. The method comprises the following steps: an adaptive time-frequency domain fusion network model (ATFFNet) is constructed, the model mainly comprises three modules of double-branch feature extraction, time-frequency information fusion and pooling fusion, a time-domain branch and a time-frequency branch cooperative feature extraction framework is constructed to realize joint modeling of multi-source information; an adaptive frequency band division and multi-scale feature enhancement mechanism is introduced in the time-frequency branch to improve the representation ability of complex frequency components and local degradation characteristics, and the representation ability of the model to time-frequency characteristics and weak fault characteristics in a non-stationary signal is improved; meanwhile, a gated attention fusion module is designed to realize effective interaction and adaptive aggregation of cross-branch and cross-band information, so that the diagnosis precision, robustness and generalization ability of the model in a multi-sensor scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a fault diagnosis method based on an adaptive time-frequency fusion gated attention network, belonging to the field of equipment fault diagnosis technology. Background Technology

[0002] Rotating machinery plays a vital role in modern industrial production, with widespread applications in machinery manufacturing, aerospace, mining, and medical equipment. Bearings, as a crucial component of rotating machinery, face increasingly stringent requirements under harsh conditions such as high speed and high load as mechanical equipment and parts continue to evolve. After prolonged use, bearings are prone to failure due to wear, fatigue, or insufficient lubrication, potentially leading to equipment downtime or even serious safety accidents, resulting in significant economic losses. Therefore, condition monitoring and fault diagnosis of bearings are of paramount importance.

[0003] Existing bearing fault diagnosis methods can be broadly categorized into model-based methods and data-driven methods. Model-based methods identify faults by establishing physical or mathematical models of the bearing and rotating machinery system, offering strong physical interpretability and advantages under conditions of data scarcity or relatively controllable operating conditions. However, due to the complexity of actual mechanical systems and significant coupling between components, these methods still face considerable challenges in high-precision modeling and engineering applications. In contrast, data-driven methods do not rely on explicit mechanistic models but directly learn the mapping relationship between fault characteristics and operating states from historical monitoring data. Traditional data-driven methods typically utilize signal processing techniques such as Fourier transform, wavelet transform, variational mode decomposition, and eigenmode decomposition to extract fault-sensitive features, combining them with traditional machine learning algorithms to complete state identification. With the development of deep learning, models such as convolutional neural networks, long short-term memory networks, autoencoders, and Transformers are widely used in single-sensor bearing fault diagnosis, significantly improving the representation ability of complex nonlinear modes through end-to-end feature learning. Meanwhile, researchers also converted one-dimensional vibration signals into time-frequency images or structured images, and combined them with 2D-CNN, attention mechanisms and Transformer models to further improve diagnostic performance.

[0004] Meanwhile, in data-driven methods, single-sensor fault diagnosis methods often struggle to comprehensively characterize bearing operating status information in real-world industrial scenarios due to factors such as sensor placement, spatial coverage, and environmental noise interference, leading to misjudgments or missed detections. In contrast, multi-sensor fusion technology can integrate complementary information from different sensing modes and measurement points, effectively improving the completeness of fault feature characterization and the accuracy and reliability of diagnostic results. For example, Xu et al. (see the published content in "Xu Z, Chen X, Li Y, et al. Hybrid Multimodal Feature Fusion with Multi-Sensor for Bearing Fault Diagnosis") performed data-level fusion of multi-sensor information using principal component analysis and combined residual neural networks and support vector machines to construct a diagnostic model, demonstrating good recognition performance in high-noise environments. Wan et al. (see “S. Wan, T. Li, B. Fang, K. Yan, J. Hong and X. Li. Bearing Fault Diagnosis Based on Multisensor Information Coupling and Attentional Feature Fusion. IEEE Transactions on Instrumentation and Measurement”) proposed a novel feature-level information coupling model, MICN, which can independently extract deep features from multi-sensor signals and achieve synchronous fusion at different network layers. Lin et al. (see “Lin T, Ren Z, Zhu L, et al. Neural architecture search for multi-sensor informationfusion-based intelligent fault diagnosis”) constructed a continuous search space and jointly searched for the optimal feature extraction unit and fusion initiation layer to obtain a better multi-sensor fusion diagnostic structure, achieving effective extraction and collaborative fusion of multi-sensor information.

[0005] Although existing research has demonstrated the effectiveness of multi-sensor fusion in bearing fault diagnosis, further research is needed on how to explore the deep coupling relationships between multi-source signals, improve feature modeling and information fusion processing, and enhance the model's generalization ability and robustness in complex environments. On the one hand, most existing methods rely on fixed fusion structures or predefined fusion strategies, making it difficult to achieve adaptive information interaction based on the strength of each sensor's response and its correlation in different samples. Therefore, the characterization of the deep coupling relationships between multi-source signals is still insufficient. On the other hand, bearing fault signals usually have obvious non-stationarity and multi-scale characteristics. Simply relying on time-domain analysis or time-frequency analysis with fixed frequency band division often fails to fully extract fault-sensitive information, resulting in the loss of key information and ultimately affecting diagnostic accuracy. Summary of the Invention

[0006] To address the shortcomings of current multi-sensor fusion methods in bearing fault diagnosis, which rely primarily on fixed fusion structures or predefined fusion strategies, resulting in insufficient characterization of deep coupling relationships between multi-source signals, and the difficulty in fully extracting fault-sensitive information through simple time-domain analysis or time-frequency analysis with fixed frequency band divisions, leading to the loss of key information and ultimately affecting diagnostic accuracy, this invention provides a fault diagnosis method based on an adaptive time-frequency fusion gated attention network. The technical solution is as follows: A fault diagnosis method based on an adaptive time-frequency fusion gated attention network, the method comprising: Step 1: Initialize the original multi-source vibration input signal from multiple sensors to obtain a one-dimensional vibration signal, and then use continuous wavelet transform (CWT) to convert the one-dimensional vibration signal into a two-dimensional image signal containing time-frequency features. Step 2: Input the one-dimensional vibration signal and the two-dimensional image signal into a dual-branch feature extraction module that includes a time-domain branch and a time-frequency branch. The local feature patterns and fault features of the one-dimensional vibration signal are captured in a deep hierarchical representation through the time-domain branch. The time-frequency branch focuses on the time-frequency information and detailed texture structure of the two-dimensional image and further complements the information of the time-domain branch. Step 3: Combine the time-frequency information fusion module to construct an adaptive time-frequency fusion network model (ATFFNet). The frequency band of the time-frequency signal is adaptively segmented by the context information of the time-domain branch. Then, the segmented features are gated and fused with the time-domain signal features output by the dual-branch feature extraction module by combining the attention mechanism, so as to realize adaptive feature extraction and weighted fusion of data from different sensors.

[0007] Furthermore, in step 1, the two-dimensional image signal obtained after continuous wavelet transform is further processed by Synchrosqueezing Transform (SST) to enable the two-dimensional image to obtain a more concentrated time-frequency representation and facilitate component reconstruction.

[0008] Specifically, the input signal The continuous wavelet transform is expressed by the following formula (1): (1) in Represents the normalized mother wavelet, As a scale, For time shift, the upper horizontal line indicates complex conjugate. The energy normalization coefficient is used to ensure that the mother wavelet energy is consistent across different scales. The instantaneous frequency at point (a,b) is expressed by the following formula (2): (2) in This represents the partial derivative with respect to variable b, and Only Calculate in time, then according to Reassigning the scale axis to the frequency axis yields the Synchronous Compression Transform (SST) representation: (3) The integration is performed on the scale set A, and the weighting factor is... It is the normalization result under scale transformation.

[0009] Furthermore, the time-domain branch includes a MOGA module, which employs a dual-path parallel structure to perform multi-level modeling of the temporal features.

[0010] Specifically, one path of the MOGA module uses a 1×1 convolution to linearly map the input features; the other path first extracts the local context through a 3×1 convolution, then divides the intermediate features into two parts along the channel dimension, and performs multi-scale temporal modeling using 5×1 and 7×1 convolutions respectively, and finally combines the 1×1 convolution branch to complete channel recombination.

[0011] Furthermore, the time-frequency branch includes an MSFE module, which sequentially models the input features using 7×7, 5×5, and 3×3 convolutions to simultaneously cover coarse-grained contours and fine-grained local textures. The MSFE module also introduces a direction-aware gating mechanism to obtain two directional attention weights in the frequency and time dimensions, which are used to modulate the multi-scale convolutional features.

[0012] Furthermore, in step 3, the adaptive frequency band segmentation includes using the context information provided by the time branch to predict the frequency band boundaries of the sub-band division of the time-frequency branch on a sample-by-sample and time-step basis, so that the division of the low-frequency, mid-frequency and high-frequency sub-bands of the time-frequency branch can be adaptively adjusted according to the changes in the local time sequence pattern.

[0013] Furthermore, in step 3, the attention mechanism includes a multi-branch, multi-head attention mechanism, which enables the three sub-bands of the adaptive frequency band segmentation to obtain attention outputs for the three frequency bands respectively. Combined with the time query features of the time domain branch, a dynamic fusion result is obtained. The dynamic fusion result is then subjected to global weighted fusion to generate a global fusion result. The dynamic fusion result and the global fusion result generate a final fusion result. The final fusion result is then connected with the original time branch features to establish residual connections for feature enhancement. The enhanced features are fed into a feedforward neural network (FFN) for nonlinear mapping to obtain the final module output.

[0014] Furthermore, the Adaptive Time-Frequency Domain Fusion Network Model (ATFFNet) also includes a pooling fusion module, which includes Global Average Pooling (GAP) and Global Max Pooling (GMP). Global Average Pooling and Global Max Pooling operations are applied to the final module output in step 3 in the time-series dimension to obtain two complementary channel-level global descriptive features. The two global descriptive features are then fused to obtain the final output.

[0015] A storage medium storing a computer program for a fault diagnosis method based on an adaptive time-frequency fusion gated attention network, wherein the computer program causes a computer to execute the fault diagnosis method described above.

[0016] The beneficial effects of this invention are: This application proposes a fault diagnosis method based on an adaptive time-frequency fusion gated attention network. An adaptive time-frequency domain fusion network model (ATFFNet) is constructed for multi-sensor bearing fault diagnosis. This model establishes a collaborative feature extraction framework between the time-domain branch and the time-frequency branch to achieve joint modeling of multi-source information. Simultaneously, adaptive frequency band division and multi-scale feature enhancement mechanisms are introduced into the time-frequency branch to improve the representation ability of complex frequency components and local degradation features, thereby enhancing the model's ability to represent time-frequency features and weak fault features in non-stationary signals. Furthermore, a gated attention fusion module is designed to achieve effective interaction and adaptive aggregation of cross-branch and cross-frequency band information, thereby improving the model's diagnostic accuracy, robustness, and generalization ability in multi-sensor scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the overall framework of the ATFFNet model according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the AFS_GA module according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the Embedding module and the Downsampling module in an embodiment of the present invention; Figure 4 This is a schematic diagram of the multi-head cross-attention mechanism according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the hardware configuration structure of the Southeast University bearing test bench according to an embodiment of the present invention; Figure 6 This is a confusion matrix of the model in this application and four prior art models under the SEU dataset of this invention embodiment; Figure 7 This invention relates to the model in this application and four existing technology models (tsne) under the SEU dataset of this embodiment. Figure 8 This is the adaptive frequency band segmentation diagram (C0~C2) of the ATFFNet model in this embodiment of the invention. Figure 9 These are adaptive frequency band segmentation diagrams (C3, C4) of the ATFFNet model in this embodiment of the invention. Figure 10 This is a schematic diagram of the hardware configuration structure of the self-built multi-sensor test bench according to an embodiment of the present invention; Figure 11 This is a confusion matrix between the model of this application and four existing technology models under the self-built test bench dataset of this invention embodiment; Figure 12 The self-built test bench dataset of this invention includes the model of this application and four existing technology models tsne. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0020] Example 1 This embodiment provides a fault diagnosis method based on an adaptive time-frequency fusion gated attention network. See [link to relevant documentation]. Figure 1 The method includes: Step 1: Initialize the original multi-source vibration input signals from multiple sensors to obtain a one-dimensional vibration signal, and then use continuous wavelet transform (CWT) to transform the one-dimensional vibration signal into a two-dimensional image signal containing rich time-frequency features.

[0021] Furthermore, in step 1, the two-dimensional image signal obtained after continuous wavelet transform is further processed by Synchrosqueezing Transform (SST) to enable the two-dimensional image to obtain a more concentrated time-frequency representation and facilitate component reconstruction.

[0022] Synchronous compression transform is an energy redistribution method proposed to address the shortcomings of conventional time-frequency representations such as short-time Fourier transform and continuous wavelet transform in terms of frequency resolution or energy diffusion. Its main principle is to first obtain the coefficients in the time-scale domain using continuous wavelet transform, and then estimate the instantaneous frequency based on the phase changes of the coefficients, thus compressing (reassigning) the coefficients from the scale axis to a more accurate frequency axis, thereby obtaining a more concentrated time-frequency representation and facilitating component reconstruction.

[0023] Specifically, the input signal The continuous wavelet transform is defined as shown in the following formula (1): (1) in Represents the normalized mother wavelet, As a scale, For time shift, the upper horizontal line indicates complex conjugate. The energy normalization coefficient is used to ensure that the mother wavelet energy is consistent across different scales. The instantaneous frequency at point (a,b) is expressed by the following formula (2): (2) in This represents the partial derivative with respect to variable b, and Only (ε is usually taken as a small positive number, such as 10) -6 10 -8 To avoid calculating instantaneous frequency when the wavelet coefficient amplitude is too small (thus preventing numerical instability and noise amplification), the calculation is performed, and then... according to Reassigning the scale axis to the frequency axis yields the Synchronous Compression Transform (SST) representation, as shown in the following formula (3): (3) The integration is performed on the scale set A, and the weighting factor is... It is the normalization result under scale transformation.

[0024] Synchronous Compressed Transform (SST) achieves energy compression in the frequency domain, resulting in a sharper frequency trajectory, which facilitates subsequent component extraction and classification. To ensure numerical stability, instantaneous frequency calculation employs derivative wavelet or frequency domain multiplication. The derivative operation is performed in this way, and the derivative is performed in this way. Set a threshold for points that are too small and then... SST performs small-scale smoothing. While satisfying wavelet invertibility, SST maintains signal reconfigurability and has significant advantages in transient and weak impulse detection.

[0025] Step 2: Input the one-dimensional vibration signal and the two-dimensional image signal into a dual-branch feature extraction module that includes a time-domain branch and a time-frequency branch. The local feature patterns and fault features of the one-dimensional vibration signal are captured in a deep hierarchical representation through the time-domain branch. The time-frequency branch focuses on the time-frequency information and detailed texture structure of the two-dimensional image and further complements the information of the time-domain branch.

[0026] Specifically, the MOGA (Multi-Order Gated Aggregation) module and the MSFE (Multi-Scale Feature Extraction) module are introduced in the time-domain branch and the time-frequency branch, respectively. The former focuses on the joint modeling of temporal local patterns and long receptive field information, while the latter is used to enhance the key structural responses and salient region representations in the time-frequency plot. Together, they provide a more robust feature foundation for subsequent cross-branch interactions.

[0027] The MOGA module employs a dual-path parallel structure to perform multi-level modeling of temporal features. Specifically, one path uses a 1×1 convolution to linearly map the input features, preserving the basic semantics in the original temporal representation; the other path first extracts the local context through a 3×1 convolution, then divides the intermediate features into two parts along the channel dimension, performing multi-scale temporal modeling using 5×1 and 7×1 convolutions respectively, and finally combines the aforementioned 1×1 convolution branch to complete channel recombination.

[0028] Let the input features of the time-domain branch be... The process is described as follows: (6) (7) in, , Indicates channel splicing. This represents the SiLU activation function.

[0029] The final output is: (8) Therefore, the MOGA module can preserve local transient details while introducing contextual information under a larger receptive field, thereby improving the model's ability to perceive impact features, modulation patterns, and periodic evolution patterns.

[0030] The MSFE module first utilizes cascaded convolutional pyramids to extract multi-scale time-frequency structure information. Specifically, let the input features of the time-frequency branch be... That is, the input features are modeled stepwise using 7×7, 5×5, and 3×3 convolutions in sequence to simultaneously cover coarse-grained contours and fine-grained local textures. The output of this multi-scale convolution is denoted as... Fs Building upon this, to highlight key response regions in the time-frequency plot, this paper further introduces a direction-aware gating mechanism. Specifically, gating features are first obtained through 1×1 projection. G Then, maximum aggregation is performed along the frequency dimension and the time dimension respectively to obtain two directional attention weights. The process can be described as follows: (9) in, This represents the Sigmoid function. and These represent the maximum aggregation operations along the frequency and time dimensions, respectively. Subsequently, the multi-scale convolutional features are modulated using weights in two directions: (10) Here, ⊙ represents element-wise multiplication. This approach enables the MSFE module to simultaneously focus on transient salient regions in the time direction and energy concentration regions in the frequency direction, thereby improving the separability of weak fault textures and local anomalous structures.

[0031] In summary, the MOGA and MSFE modules enhance feature representation capabilities from the perspectives of temporal dynamic modeling and temporal saliency enhancement, respectively. The MOGA module emphasizes the aggregation and complementation of temporal patterns across multiple receptive fields, while the MSFE module focuses on multi-scale structure extraction and orientation-sensitive enhancement. The two modules effectively complement each other, providing more discriminative input representations for subsequent cross-branch attention modeling and adaptive fusion.

[0032] Step 3: To improve the model's adaptability to local time-frequency pattern changes in non-stationary signals, an adaptive time-frequency fusion network model (ATFFNet) is constructed by combining the time-frequency information fusion module. The frequency band of the time-frequency signal is adaptively segmented by the context information of the time-domain branch. Then, the segmented features are gated and fused with the time-domain signal features output by the dual-branch feature extraction module by combining the attention mechanism, so as to realize adaptive feature extraction and weighted fusion of data from different sensors.

[0033] In step 3, the adaptive frequency band segmentation includes using the context information provided by the time branch to predict the frequency band boundaries of the sub-band division of the time-frequency branch on a sample-by-sample and time-step basis, so that the division of the low-frequency, mid-frequency and high-frequency sub-bands of the time-frequency branch can be adaptively adjusted according to the changes in the local time sequence pattern.

[0034] For details, see Figure 2 Let the time-frequency branching feature be represented as The time branch feature is represented as ,in B , C , F and T These represent the batch size, number of channels, frequency dimension length, and time dimension length, respectively. When the time lengths of the output features of the time branch and the time-frequency branch in step 2 are inconsistent, the time dimension is first aligned using interpolation to obtain... Subsequently, the aligned temporal branch features are input into the context encoder to extract local temporal context representations: (11) in, This represents a context encoding function consisting of two one-dimensional convolutional layers. Based on contextual features, a boundary prediction head generates dynamic offsets corresponding to two boundary parameters. (12) in, This represents a 1×1 convolution mapping. To avoid boundary instability caused by excessively large dynamic offset amplitudes, a tanh function is used to impose bounded constraints on the offset, and the DC bias in the time dimension is further eliminated. The process can be described as follows: , (13) (14) set up and As globally learnable fundamental boundary parameters, two dynamic frequency band boundaries are generated using an ordered parameterization method. The process can be described as follows: (15) in, Let be the Sigmoid function. From the above definition, we know that always holds. This ensures that the order of low-frequency, mid-frequency, and high-frequency sub-bands remains consistent. Therefore, the frequency band boundary is no longer a fixed constant, but a time-varying function dynamically determined by the time branch context.

[0035] After obtaining the dynamic boundary, a continuously differentiable soft mask is constructed based on the normalized frequency coordinates. Let the...f The normalized coordinates of each frequency position are The low-frequency, mid-frequency, and high-frequency masks are defined as follows: (16) in, , where is the temperature coefficient, used to control the smoothness of the boundary transition region. Furthermore, the characteristics of the three sub-bands can be expressed as: (17) Where ⊙ represents element-wise multiplication, and the three sub-band characteristics of low frequency, mid frequency, and high frequency are denoted as follows: X l , X m and X h (correspond Figure 3 (F3, F2, F1). Because a soft segmentation mechanism is used, the entire frequency band division process remains end-to-end differentiable, allowing the frequency band boundaries to be jointly optimized under the drive of the classification objective, while reducing the information mutation problem caused by hard boundary segmentation.

[0036] To further improve training stability, the model introduces constraints on the segmentation boundaries, including parameter magnitude constraints, minimum bandwidth constraints, prior boundary constraints, basic parameter anchor point constraints, and boundary dynamics constraints. Accordingly, the overall training objective can be expressed as: (18) in, For classifying losses, This is the regularization term for the band segmentation module. Through the above design, the adaptive band segmentation module can dynamically adjust the band segmentation position according to the temporal context of the time branch, and provide a clearer and more discriminative structured input for subsequent multi-branch attention modeling and cross-band fusion, thereby improving the model's ability to represent time-frequency features and weak fault features in non-stationary signals.

[0037] Furthermore, in step 3, the attention mechanism includes a multi-branch, multi-head attention mechanism, which enables the three sub-bands of the adaptive frequency band segmentation to obtain attention outputs for the three frequency bands respectively. Combined with the time query features of the time domain branch, a dynamic fusion result is obtained. The dynamic fusion result is then subjected to global weighted fusion to generate a global fusion result. The dynamic fusion result and the global fusion result generate a final fusion result. The final fusion result is then connected with the original time branch features to establish residual connections for feature enhancement. The enhanced features are fed into a feedforward neural network (FFN) for nonlinear mapping to obtain the final module output.

[0038] Specifically, after completing the adaptive frequency band segmentation, the model obtains three sub-band features: low frequency, mid frequency, and high frequency, denoted as follows:X l , X m and X h Since the discriminative information contained in different frequency bands varies significantly, direct splicing or simple weighting can easily overlook the imbalance in representational capabilities between frequency bands. To address this, this application constructs a multi-branch, multi-head attention mechanism after adaptive frequency band segmentation, models the three sub-bands separately, and uses temporal branch features as query signals to guide the model to extract the feature responses most relevant to the current temporal semantics from different frequency subspaces.

[0039] For each frequency band feature X b ,in b ∈{ l , m , h First, the effective bandwidth is calculated using the corresponding soft mask and then normalized to reduce amplitude bias caused by differences in bandwidth. Then, the bandwidth features are expanded along the frequency dimension and obtained through linear mapping, downsampling, and convolutional projection. For example... Figure 3 As shown, the frequency bands are represented as follows: , (19) in, This represents the frequency band projection operator. Thus, low-frequency, mid-frequency, and high-frequency features are mapped to a unified embedding space, serving as keys and values ​​in the attention mechanism, and the temporal branch is embedded into the query vector through a linear mapping.

[0040] Based on this, three independent multi-head cross-attention branches are constructed, and the multi-head cross-attention mechanism is as follows: Figure 4 As shown. With the first b Taking a frequency band as an example, its attention output can be expressed as: (20) Attention outputs for three frequency bands were obtained respectively. O l , O m and O h Then, by combining the time query features and the attention output of each frequency band, position-by-position fusion weights are dynamically generated, thus obtaining the dynamic fusion result: (twenty one) in, α、b This represents the adaptive weights for the corresponding frequency band, satisfying... Meanwhile, to preserve global frequency band bias, the model also introduces a global gating branch to perform global weighted fusion of the three frequency band outputs, denoted as... F glo The final fusion result is determined jointly by dynamic fusion and global fusion: (twenty two) in, λ These are learnable mixing coefficients. After obtaining the fused features, the model is further enhanced through convolutional transformation and query gating, and residual connections are established with the original temporal branch features to preserve the temporal backbone information and stabilize the training process, i.e.: (twenty three) in, This indicates post-processing convolution. This represents the query gating function. Finally, the enhanced features are fed into a feedforward network for nonlinear mapping to obtain the final output: (twenty four) Through the above design, this application employs a multi-branch, multi-head attention mechanism to model the correlation between frequency bands and temporal semantics in the low-frequency, mid-frequency, and high-frequency sub-band spaces, and integrates complementary information from different frequency bands through a subsequent adaptive fusion mechanism. Compared to single-path attention or simple concatenation methods, this method can more fully utilize the heterogeneity of features in each frequency band, thereby improving the model's ability to discriminate complex non-stationary signals.

[0041] Furthermore, the Adaptive Time-Frequency Domain Fusion Network Model (ATFFNet) also includes a pooling fusion module, which includes Global Average Pooling (GAP) and Global Max Pooling (GMP). Global Average Pooling and Global Max Pooling operations are applied to the final module output in step 3 in the time-series dimension to obtain two complementary channel-level global descriptive features. The two global descriptive features are then fused to obtain the final output.

[0042] Specifically, in existing technologies, common global pooling strategies mainly include Global Average Pooling (GAP) and Global Max Pooling (GMP). GAP, by averaging the feature map, can better reflect the overall distribution trend and global energy information, while GMP, by retaining the maximum response, is more conducive to highlighting local salient patterns or transient strong activation features. However, a single pooling method often struggles to simultaneously capture steady-state and abrupt changes, potentially leading to the loss of discriminative features. To address this, this application proposes a pooling fusion module. By introducing a learnable channel-level gating mechanism, it adaptively fuses the global features extracted by GAP and GMP through an adaptive fusion mechanism, thereby obtaining a more robust and discriminative compact representation.

[0043] The specific process is as follows, with the input feature mapping as follows: First, global average pooling and global max pooling operations are applied in the time dimension to obtain two complementary channel-level global descriptions: (25) in, This indicates that the mean is calculated over the time dimension. This indicates finding the maximum value over the time dimension.

[0044] To adaptively measure the importance of GAP and GMP on different channels, this application aggregates their features and inputs them into a lightweight gating network: (26) Where z represents the aggregated feature. Subsequently, a two-layer fully connected network with dimensionality reduction is used to generate channel-level gating weights: (27) in, Here, r is a learnable parameter, representing the channel compression ratio. and These represent the ReLU activation function and the Sigmoid function, respectively. The generated gate vectors... Each channel is assigned an independent fusion weight. After obtaining the channel-level gating weights, the two pooling features are fused in a channel-by-channel combination manner: (28) in, This represents element-wise multiplication. This fusion method ensures that the output features always lie between the GAP and GMP representation spaces, and can dynamically balance mean and extreme value information based on the statistical properties of the input features.

[0045] The adaptive time-frequency domain fusion network model (ATFFNet) proposed in this application mainly includes three modules: dual-branch feature extraction, time-frequency information fusion, and pooling fusion. First, the multi-source input signals are initialized. Then, CWT is used to transform the one-dimensional vibration signal into a two-dimensional image containing rich time-frequency features, providing an information-dense and highly separable input for subsequent deep learning models. The one-dimensional vibration signal is a one-dimensional time-series vibration signal. Next, the one-dimensional time-series vibration signal and the two-dimensional time-frequency image signal are input into a dual-branch feature extraction module. A multi-order 1DCNN is used to extract a deep hierarchical representation of the one-dimensional time-series vibration signal that captures local feature patterns and fault features. Max pooling layers are used to retain key feature representations while reducing feature resolution and computational cost. The feature extraction branch uses a shallow 2D CNN and a time- and frequency dual-dimensional structure attention mechanism to achieve multi-scale feature extraction and orientation sensitivity enhancement. Next, the features extracted by the dual-branch feature extraction module are input into the time-frequency information fusion module. The frequency band of the time-frequency signal is adaptively segmented by the context information of the time-domain branch. Combined with the attention mechanism, the segmented features are gated and fused with the time-domain signal features to achieve adaptive feature extraction and weighted fusion of data from different sensors. Finally, in the pooling fusion and classification output stage, the fused features are used for multi-class fault identification through a fully connected layer, thereby enhancing the model's generalization ability.

[0046] This application uses two datasets to experimentally verify the proposed method, demonstrating its superiority and robustness. One dataset is the Southeast University bearing dataset, and the other is the laboratory-built rotor test bench bearing dataset. These datasets cover various typical fault types of key rotating components of bearings, and multi-sensor technology is used for signal acquisition.

[0047] The experiment was conducted on a hardware configuration equipped with an NVIDIA RTX3090 24GB graphics card, using a Python 3.8 compilation environment and a PyTorch 2.4.1 platform.

[0048] On the one hand, taking the Southeast University bearing dataset as an example, its SEU bearing dataset was obtained through experiments using a dynamic simulator of a transmission system, such as... Figure 5 As shown, the data covers two speed-load configurations: 20Hz-0V and 30Hz-2V. Under each condition, it includes five states: healthy state, rolling element fault, inner race fault, outer race fault, and a combined inner and outer race fault. This dataset contains multi-sensor information from eight monitoring channels, including motor vibration, motor torque, and three-axis (x, y, z) vibration signals generated by the planetary gearbox and parallel gearbox. This application selected data collected under the 20Hz-0V condition, and segmented the original one-dimensional vibration sequence using a sliding window method. Each sample length was set to 1024 sampling points, resulting in 2000 samples for each fault type. Detailed information about the SEU dataset is shown in Table 1.

[0049] Table 1

[0050] The adaptive time-frequency fusion network model (ATFFNet) in this application, with detailed parameters for dual-branch feature extraction, adaptive frequency band segmentation, and gated attention mechanism, is shown in Table 2. Furthermore, the model was trained with a learning rate of 0.0003, a batch size of 32, a maximum number of iterations of 50, and a frequency band learning rate of 0.0002.

[0051] Table 2

[0052] To objectively evaluate the diagnostic performance of the ATFFNet model in this application, four representative existing bearing fault diagnosis methods—ETMD, MCMI-GCFN, MD-BiMamba, and MSTF—were selected as comparative models. These methods are all highly representative in terms of network structure design, particularly in multi-branch feature extraction, time-frequency information utilization, and image-based representation modeling.ETMD (see the published content in "M.-H. Vu, V.-Q. Nguyen, T.-T. Tran, V.-T. Pham and M.-T. Lo, Few-Shot Bearing Fault Diagnosis Via EnsemblingTransformer-Based Model With Mahalanobis Distance Metric Learning FromMultiscale Features") is a two-branch ensemble model that improves feature discrimination by combining Transformer global modeling with Mahalanobis distance metric; MCMI-GCFN (see the published content in "Wang Z, Nie P, Liu J, et al. Bearing fault diagnosis based on a multiple-constraint modal-invariant graph convolutional fusion network") adopts a multimodal, multi-branch structure, extracts features from different modalities separately, and uses graph convolution to achieve relation modeling and information fusion; MD-BiMamba (see the published content in "Wang P, Song Y, Wang X, et al. MD-BiMamba: An aero-engine") The inter-shaft bearing fault diagnosis method based on Mamba with modal decomposition and bidirectional feature fusion strategy first performs modal decomposition on the original signal, then uses bidirectional Mamba to model long-range temporal dependencies and complete feature fusion. MSTF (see the published content of "Zekun W, Zifei X, Chang C, et al. Rolling bearing fault diagnosis method using time-frequency information integration and multi-scale TransFusion network") converts the one-dimensional signal into a two-dimensional time-frequency image, combining multi-scale feature extraction and Transformer global modeling to achieve fault identification. To ensure fairness in the evaluation, all models are implemented using the original configuration.The diagnostic performance results of the ATFFNet model in this application compared with existing technologies such as ETMD, MCMI-GCFN, MD-BiMamba and MSTF are detailed in Table 3 below (comparison results of different models %). Table 3

[0053] As shown in Table 3, all the comparative models achieved high diagnostic performance on this dataset, with accuracy exceeding 98.00%, indicating that these methods can effectively identify bearing faults. Among them, MSTF performed best, with accuracy, precision, recall, and F1 score all reaching 99.00%, demonstrating that the modeling approach based on time-frequency images and multi-scale feature fusion has good feature representation capabilities. In contrast, the overall performance of ETMD, MCMI-GCFN, and MD-BiMamba models was slightly lower, indicating that while the synergy of local and global feature information, multimodal graph convolutional fusion, or temporal modeling can extract certain discriminative information, they still have limitations in characterizing complex fault features in the current task. However, the ATFFNet model proposed in this application achieved 99.70% in all four metrics (accuracy, precision, recall, and F1 score), making it the best result among all methods. Compared to the suboptimal MSTF model, ATFFNet improved all four metrics by 0.70 percentage points; compared to ETMD, accuracy, precision, recall, and F1 score improved by 1.70, 1.55, 1.70, and 1.69 percentage points, respectively. These results demonstrate that the ATFFNet model in this application can more fully integrate time-domain and time-frequency-domain information, and highlights key fault features through multi-branch feature extraction and adaptive fusion mechanisms, thus exhibiting significant advantages in overall recognition accuracy and comprehensive classification performance. Furthermore, the high consistency between accuracy, precision, recall, and F1 score of the ATFFNet model indicates that the model provides more balanced recognition across different categories, demonstrating good stability and robustness.

[0054] See Figure 6As shown in the diagram, the confusion matrix of the ATFFNet model in this application is compared with existing technologies such as ETMD, MCMI-GCFN, MD-BiMamba, and MSTF. While all models exhibit a strong diagonal distribution, indicating that each method can effectively classify bearing faults, significant differences remain in their discriminative abilities among adjacent categories. The misclassification of the ETMD model is mainly concentrated in classes C2 and C3 (corresponding to the label types in Table 1 above), with some samples incorrectly classified into classes C1 or C4, indicating that its local discriminative ability under similar fault modes is still limited. The MCMI-GCFN model has high overall recognition performance, but some confusion still exists between classes C1 and C2. Additionally, a small number of samples from classes C3 and C4 are also misclassified into class C2, suggesting that while its multimodal graph fusion strategy improves feature representation, there is still room for improvement in fine-grained category boundary modeling. The misclassification of the MD-BiMamba model is mainly manifested as a shift from class C3 to class C4, reflecting its advantage in modeling long-range temporal dependencies, but its ability to distinguish between locally similar fault features is still insufficient. In contrast, the MSTF model has significantly fewer off-diagonal elements, with only a small number of bidirectional misclassifications between classes C2 and C3, indicating that time-frequency image representation combined with a multi-scale Transformer structure can effectively enhance class separability. The ATFFNet model proposed in this application exhibits the most significant diagonal advantage in its confusion matrix, with the fewest off-diagonal elements, indicating that this method can achieve more accurate and balanced class recognition overall.

[0055] To further analyze the feature distribution characteristics of each model, Figure 7 The t-SNE visualization results of different models on the test set are presented. It can be observed that the ETMD model can form a basically separable cluster structure, but the intra-class distribution is relatively discrete, and the local cluster boundaries are not clear enough. The MCMI-GCFN and MD-BiMamba models improve the inter-class separation, but there are still a few outliers and the boundaries between neighboring classes are close together, indicating that their feature spaces still have some overlap. The MSTF model has a more compact cluster structure and a further increase in inter-class spacing, indicating that time-frequency information fusion helps to improve the discriminativeness of deep representations. In contrast, the features extracted by the ATFFNet model in this application show higher intra-class compactness and larger inter-class spacing, with samples of each class more concentrated in the embedding space and almost no obvious overlapping areas. This result shows that the ATFFNet model in this application can learn more discriminative fault representations, thereby effectively alleviating the feature aliasing problem between similar fault categories. Its superiority mainly comes from the collaborative modeling of time domain and time-frequency domain information, and the strengthening effect of multi-branch feature extraction and adaptive fusion mechanisms on key fault information.

[0056] See Figure 8 ,9 This paper presents the adaptive frequency band segmentation results of the ATFFNet model in this application on representative samples of different fault categories on the test set. Specifically, three samples were randomly selected from the test set data under five labels (C0-C4), and the corresponding frequency band segmentation results were plotted. It can be seen that the two dynamic frequency band boundary curves p1(t) and p2(t) (corresponding to the red and blue lines in the figure, respectively) change steadily across all types of samples, and there is high consistency among samples of the same type, indicating that the frequency band segmentation strategy learned by the model has good stability and robustness. From the overall distribution, the average p1 is about 0.133, the average p2 is about 0.575, and the average proportions of the low-frequency, mid-frequency, and high-frequency sub-bands are about 13.34%, 44.14%, and 42.52%, respectively. This indicates that while retaining basic information in the low-frequency range, the model focuses more on mining fault discrimination features in the mid- and high-frequency regions. Furthermore, there are still some differences in the boundary positions between different categories. For example, p1 is relatively higher and p2 is relatively lower for category C3, indicating a slightly wider low-frequency band and a slightly narrower mid-frequency band, suggesting that this type of fault feature is more sensitive to low-frequency information. C4, on the other hand, has a slightly higher p2, indicating that the model allocates more attention to its mid-frequency information. The model is able to maintain a stable global frequency band structure while fine-grainedly adjusting the segmentation boundaries according to different fault features, thereby enhancing feature representation ability and fault recognition performance.

[0057] To verify the effectiveness of each module in the ATFFNet model of this application, ablation experiments were conducted based on the proposed ATFFNet model architecture. The specific configuration is shown in Table 4. Here, `time` and `time-frequency` represent time-domain and time-frequency single-branch modules, respectively. `w / o adaptive(equal)` and `w / o adaptive(learned)` are both modules with no learnable frequency band segmentation. `equal` indicates that the frequency band is equally divided, `learned` indicates that the frequency band is in the optimal frequency band state of the ATFFNet training band, and `w / o query gate` is an ungated attention mechanism.

[0058] Table 4

[0059] The results of the ablation experiments are shown in Table 5 below. Compared with the complete ATFFNet model framework of this application, the classification performance of all ablation models decreased to varying degrees, indicating that each key component in the proposed framework contributes positively to the final performance. Specifically, when only the time branch is retained, the model's accuracy and macro-F1 score are 47.49% and 43.97%, respectively, showing the most significant decrease compared to the complete model. When only the time-frequency domain branch is retained, the accuracy and macro-F1 score improve to 62.28% and 60.15%, respectively, but are still significantly lower than the complete model. This result shows that the time-domain branch and the frequency-domain branch are significantly complementary in fault characterization, and the absence of either branch will lead to a significant weakening of the discrimination ability. Regarding the frequency band division strategy, when using a fixed-width band, the model's accuracy drops to 87.65%, and the Macro-F1 score is 87.44%, indicating that a simple uniform band division method is insufficient to adequately adapt to the distribution of effective frequency components in the vibration signal. In contrast, when using the learned fixed band boundaries, the model still achieves 98.51% accuracy and 98.51% Macro-F1, showing only a slight performance degradation, demonstrating that a reasonable frequency band boundary design is effective in maintaining model performance. Furthermore, after removing the gating mechanism, the model's accuracy and Macro-F1 score drop to 98.60% and 98.60%, respectively, further illustrating that the gating unit can effectively enhance the discriminative information representation during feature fusion. Overall, these results validate the important role of the dual-branch structure, adaptive frequency band division, and gating fusion mechanism in improving the classification performance of ATFFNet.

[0060] Table 5

[0061] On the other hand, let's take the bearing dataset from a self-built rotor test bench in the laboratory as an example. This dataset was collected by a self-built multi-sensor bearing test bench, such as... Figure 10 As shown in the figure, the data includes information from two accelerometers and one acoustic emission sensor, with a sampling frequency of 51.2 kHz. Specifically, it covers two loads: 800 N and 1000 N, at 1500 rpm. Each load includes four bearing states: healthy, inner race fault, outer race fault, and cage fault. The original one-dimensional vibration sequence was segmented using a sliding window method, with each sample length set to 4096 sampling points and a window overlap rate of 50%. Ultimately, 1000 samples were obtained for each type of fault. Detailed hardware information is shown in Table 6.

[0062] Table 6

[0063] As shown in Table 7 (Comparison Experiment Results %) of Different Models, four representative existing bearing fault diagnosis methods—ETMD, MCMI-GCFN, MD-BiMamba, and MSTF—were selected as comparison models for this laboratory dataset. The ATFFNet model from this application achieved the highest values ​​across all four evaluation metrics on this laboratory dataset: accuracy, precision, recall, and F1 score of 98.38%, 98.38%, 98.38%, and 98.37%, respectively. These results further validate the effectiveness of the ATFFNet model in this classification task. In particular, compared to the suboptimal model MCMI-GCFN, the ATFFNet model still shows an improvement of 1.43–1.63 percentage points across the four metrics, indicating that this method possesses good effectiveness and stability in this classification task.

[0064] Table 7

[0065] See Figure 11 As shown, the confusion matrices of different models exhibit significant differences. The ETMD model shows a more pronounced distribution of off-diagonal elements, with misclassifications mainly concentrated in classes C0, C3, C4, and C6 (corresponding to the categories in Table 6), indicating that its ability to distinguish between similar categories remains limited. The MSTF model shows some improvement in recognition results on some samples, but significant confusion still exists between classes C0, C3, and C4, especially the phenomenon of class C4 being misclassified as class C0. In contrast, the MCMI-GCFN and MD-BiMamba models show significantly improved diagonal clustering, indicating enhanced category recognition capabilities; however, the MCMI-GCFN model still exhibits a small number of inter-class misclassifications, while the MD-BiMamba model still shows some confusion between classes C0 and C3. A comprehensive comparison reveals that the ATFFNet model in this application has the most concentrated diagonal elements and the fewest off-diagonal elements in its confusion matrix, indicating that its recognition results across all categories are more accurate and balanced, consistent with the aforementioned quantitative evaluation results.

[0066] See Figure 12As shown, the t-SNE visualization results further validate the differences in feature representation capabilities among the various models. The ETMD model exhibits significant class overlap in its feature distribution, with some classes showing continuous connections or blurred boundaries in the embedding space, corresponding to its high misclassification rate. The MSTF model also shows a certain degree of class overlap, indicating that its learned feature representations are insufficient to sufficiently separate closely related classes. The MCMI-GCFN and MD-BiMamba models can form relatively clear clustering structures, with improved intra-class compactness and inter-class separation compared to the previous two methods, but some local areas still show class proximity or unclear boundaries. In contrast, the ATFFNet model in this application learns more compact feature clusters overall, and maintains better separation between different classes, with only slight proximity at a few boundary samples. This indicates that the ATFFNet model in this application can extract more discriminative deep feature representations, thereby effectively improving intra-class consistency and inter-class separability, and ultimately achieving better classification performance.

[0067] Example 2 A storage medium storing a computer program for a fault diagnosis method based on an adaptive time-frequency fusion gated attention network, wherein the computer program causes a computer to execute the fault diagnosis method as described in Embodiment 1.

[0068] In summary, this application addresses the problems of insufficient information representation, inadequate utilization of cross-sensor correlation, and difficulty in identifying weak fault features in multi-sensor bearing fault diagnosis, and verifies the effectiveness of the proposed method. Based on the content of this application, its main advantages can be summarized in the following three aspects: (1) In the face of problems such as information dispersion, insufficient utilization of cross-sensor correlation and easy submersion of weak fault features in multi-sensor bearing fault diagnosis, the ATFFNet model proposed in this application can better realize the joint representation of multi-source fault information, and provide an effective modeling idea for multi-sensor bearing fault diagnosis. (2) Experimental results show that it performs well in complex time-frequency feature modeling and fault-sensitive information extraction, which can enhance the model’s ability to identify key features in non-stationary signals and improve the robustness and generalization ability of the diagnostic process to a certain extent. (3) The validation results on the SEU bearing dataset and the self-built test bench data reached 99.70% and 98.38% respectively, indicating that the ATFFNet model of this application has good diagnostic effect and application potential under different sensor configuration conditions.

[0069] Some steps in the embodiments of the present invention can be implemented using software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.

[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A fault diagnosis method based on an adaptive time-frequency fusion gated attention network, characterized in that, The method includes: Step 1: Initialize the original multi-source vibration input signal from multiple sensors to obtain a one-dimensional vibration signal, and then use continuous wavelet transform to convert the one-dimensional vibration signal into a two-dimensional image signal containing time-frequency features. Step 2: Input the one-dimensional vibration signal and the two-dimensional image signal into a dual-branch feature extraction module that includes a time-domain branch and a time-frequency branch. The local feature patterns and fault features of the one-dimensional vibration signal are captured in a deep hierarchical representation through the time-domain branch. The time-frequency branch focuses on the time-frequency information and detailed texture structure of the two-dimensional image and further complements the information of the time-domain branch. Step 3: Combine the time-frequency information fusion module to construct an adaptive time-frequency domain fusion network model. The frequency band of the time-frequency signal is adaptively segmented by the context information of the time-domain branch. Then, the segmented features are gated and fused with the time-domain signal features output by the dual-branch feature extraction module by combining the attention mechanism, so as to realize adaptive feature extraction and weighted fusion of data from different sensors. The adaptive frequency band segmentation includes using the context information provided by the time branch to predict the frequency band boundaries of the sub-bands of the time-frequency branch on a sample-by-sample and time-step basis, so that the division of the low-frequency, mid-frequency and high-frequency sub-bands of the time-frequency branch can be adaptively adjusted according to the changes in the local time-series pattern. The attention mechanism includes a multi-branch, multi-head attention mechanism, which enables the three sub-bands of the adaptive frequency band segmentation to obtain attention outputs for the three frequency bands respectively. Combined with the time query features of the time domain branch, a dynamic fusion result is obtained. The dynamic fusion result is then subjected to global weighted fusion to generate a global fusion result. The dynamic fusion result and the global fusion result generate the final fusion result. The final fusion result is then connected with the original time branch features to establish residual connections for feature enhancement. The enhanced features are fed into a feedforward neural network for nonlinear mapping to obtain the final module output.

2. The method according to claim 1, characterized in that, In step 1, the two-dimensional image signal obtained after continuous wavelet transform is then subjected to synchronous compression transform processing to obtain a more concentrated time-frequency representation of the two-dimensional image and facilitate component reconstruction.

3. The method according to claim 2, characterized in that, The input signal The continuous wavelet transform is expressed by the following formula (1): (1) in Represents the normalized mother wavelet, As a scale, For time shift, the upper horizontal line indicates complex conjugate. The energy normalization coefficient is used to ensure that the mother wavelet energy is consistent across different scales. The instantaneous frequency at point (a,b) is expressed by the following formula (2): (2) in This represents the partial derivative with respect to variable b, and Only Calculate in time, then according to Reassigning the data from the scale axis to the frequency axis yields the synchronous compression transform representation: (3) The integration is performed on the scale set A, and the weighting factor is... It is the normalization result under scale transformation.

4. The method according to claim 1, characterized in that, The time-domain branch includes the MOGA module, which uses a dual-path parallel structure to perform multi-level modeling of time-series features.

5. The method according to claim 4, characterized in that, One path of the MOGA module uses a 1×1 convolution to linearly map the input features; the other path first extracts the local context through a 3×1 convolution, then divides the intermediate features into two parts along the channel dimension, and performs multi-scale temporal modeling using 5×1 and 7×1 convolutions respectively, and finally completes channel recombination by convolution with a 1×1 convolution of one path.

6. The method according to claim 1, characterized in that, The time-frequency branch includes an MSFE module, which sequentially models the input features using 7×7, 5×5, and 3×3 convolutions to simultaneously cover coarse-grained contours and fine-grained local textures. The MSFE module also introduces a direction-aware gating mechanism to obtain two directional attention weights in the frequency and time dimensions to modulate the multi-scale convolutional features.

7. The method according to claim 1, characterized in that, The adaptive time-frequency domain fusion network model also includes a pooling fusion module, which includes global average pooling and global max pooling. Global average pooling and global max pooling operations are applied to the final module output in step 3 in the time dimension to obtain two complementary channel-level global descriptive features. The two global descriptive features are then fused to obtain the final output.

8. A storage medium, characterized in that, It stores a computer program for a fault diagnosis method based on an adaptive time-frequency fusion gated attention network, wherein the computer program causes the computer to execute the fault diagnosis method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Photovoltaic module fault diagnosis system and method based on deep learning

    CN119474671A

  • Remote sensing image super-resolution reconstruction method and system based on residual hierarchical Transform

    CN119671851A