SEMG gesture recognition method based on lightweight deep separable residual attention network

The problem of individual differences and dynamic variability in sEMG gesture recognition is solved through the lightweight depth separable residual attention network (DSRANet). The accuracy and robustness of gesture recognition are improved through the depth separable residual block and the multi-axis channel attention module, and efficient sEMG gesture recognition is achieved.

CN120256824APending Publication Date: 2025-07-04CHONGQING UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510345131.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When faced with individual differences and dynamic variability, the accuracy and consistency of gesture classification are insufficient, and the traditional CNN model is susceptible to channel crosstalk problem, ignoring the timing evolution and frequency domain dynamic characteristics of muscle activity.

Method used

The lightweight depth separable residual attention network (DSRANet) is adopted, and the depth separable residual block (DSRBlock) and the multi-axis channel attention module (MACA) are used to combine the depth separable convolution and residual connections to adaptively extract the spatiotemporal characteristics of the sEMG signal, and improve the model performance through signal preprocessing, feature extraction and attention mechanisms.

Benefits of technology

While maintaining computing efficiency, the accuracy and robustness of gesture recognition are significantly improved, achieving 92.13%, 75.65% and 91.00% classification accuracy on the NinaPro DB2, DB3 and DB4 datasets, which are 4.1%, 5.07% and 1.57% higher than the existing methods, respectively, and reducing the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256824A_ABST
    Figure CN120256824A_ABST
Patent Text Reader

Abstract

The invention discloses an sEMG gesture recognition method based on a lightweight deep separable residual attention network, and relates to the technical field of intelligent recognition. According to the invention, DSRANet and a lightweight deep separable residual attention network are provided, so that the spatial-temporal characteristics of sEMG gestures can be effectively captured while the calculation efficiency (0.459 M parameter, 0.1 G FLOPs) is maintained; a DSR block is designed to serve as an innovative architecture component, through combination of depth separable convolution and residual connection and efficient extraction of discriminative features, an ablation experiment shows that the calculation complexity is reduced by 45% compared with that of standard convolution, meanwhile, an MACA mechanism is introduced, key features are enhanced in a self-adaptive cross-channel and time-frequency dimension mode, and a contrast experiment verifies that the model precision can be improved by 6.41%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent recognition, and specifically to an sEMG gesture recognition method based on a lightweight depthwise separable residual attention network. Background Technique

[0002] Gesture recognition technology based on surface electromyogram (sEMG) signals is at the forefront of modern human-machine interfaces (HMIs), enabling seamless human-machine interaction by analyzing the electrical signals generated by muscle contractions. This technology decodes user intentions by leveraging the subtle electrical impulses generated during muscle activity, providing crucial support for non-invasive intuitive control systems. Its importance is particularly prominent in transformative applications such as prosthetic control and rehabilitation medicine: in the field of prosthetics, sEMG technology endows amputees with the ability to precisely execute complex movements; in the field of rehabilitation, it helps patients with nerve injuries recover motor function. In addition, sEMG gesture recognition can enhance the naturalness and responsiveness of interactive systems such as virtual reality and robotics. However, the individual differences and dynamic variability of sEMG signals pose challenges to the accuracy and consistency of gesture classification. Overcoming these obstacles is the key to unlocking the practical application potential of sEMG technology and has made it a core area of research and innovation.

[0003] The rise of deep learning has significantly promoted the development of sEMG gesture recognition. By automatically extracting effective patterns from raw time-domain signals, it reduces the dependence on manual feature engineering. Convolutional neural networks (CNNs), as representative tools, are good at mining spatial correlations from multi-channel sEMG data. For example, a study achieved an accuracy of 83.23% on the NinaPro DB2 dataset using a five-layer parallel CNN, verifying the ability of deep architectures to capture gesture-specific features from time-domain data. However, traditional CNNs are often troubled by the problem of channel crosstalk - irrelevant signals from different muscle groups are inadvertently mixed, resulting in a decline in classification performance. To address this, innovative methods such as NKDFF-CNN use narrow-kernel convolutions to maintain channel independence and improve accuracy through dual-view feature fusion. Although these time-domain methods have pushed the performance boundaries, they often neglect the temporal evolution and frequency-domain dynamic characteristics of muscle activity, indicating the need to combine complementary strategies to comprehensively capture the complexity of sEMG signals.

[0004] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide an sEMG gesture recognition method based on a lightweight depthwise separable residual attention network to solve the technical problems proposed in the background technique.

[0006] To achieve the above purpose, the present invention provides the following technical solution: An sEMG gesture recognition method based on a lightweight depthwise separable residual attention network, at least including the following steps:

[0007] S1: Preprocess the signal. To remove noise and enhance the signal-to-noise ratio, preprocess the original sEMG signal. The preprocessing process at least includes removing noise from the original data, dataset partitioning and signal segmentation, signal conversion, and data augmentation;

[0008] S2: Construct the DSRANet model. The DSRANet model is composed of depthwise separable residual blocks and multi-axis channel attention modules to achieve efficient feature extraction and adaptive focusing. The DSRANet model takes time-frequency representations as input. The shape of the time-frequency representation input data of the DSRANet model is C×F×T, where C is the number of channels, F is the number of frequency points, and T is the number of time frames. The applications after the time-frequency representation obtained after signal preprocessing is input into the DSRANet model include an initial stage, a three-stage feature processing, and a classification stage;

[0009] S3: Feature extraction. Use depthwise separable residual blocks (DSRBlocks) to achieve efficient feature extraction. The depthwise separable residual blocks are DSR blocks. The DSR blocks integrate depthwise separable convolutions and residual connections. The depthwise separable convolution includes two types of convolutions: DWConv that independently processes each input channel and PConv that aggregates information across channels, which can significantly reduce the number of parameters and computational costs. The residual connection helps train deep networks and improve feature learning effects;

[0010] S4: Apply the attention mechanism. Use a multi-axis channel attention module, namely the MACA module. The MACA module aims to adaptively focus on relevant features in the spatial and channel dimensions.

[0011] Further, removing noise from the original data specifically means eliminating noise and artifacts in the original sEMG signal. The digital filter combination of a 50Hz notch filter and a 500Hz low-pass filter is used to remove noise from the original data, and the method used is wavelet denoising;

[0012] The 50Hz notch filter is used to eliminate power frequency interference, which is a common noise source in sEMG signals;

[0013] The 500Hz low-pass filter is used to filter out high-frequency noise, which can remove the high-frequency components in the signal and ensure the smoothness of the signal;

[0014] The wavelet denoising uses a wavelet transform method based on Daubechies mother wavelets to decompose and reconstruct the signal, which can effectively remove noise and retain the main features of the signal. The order of the wavelet transform method based on Daubechies mother wavelets is 4, and the decomposition level is 4.

[0015] Further, removing the noise in the original data includes at least the following steps:

[0016] First, perform wavelet decomposition to decompose the signal into approximation coefficients and detail coefficients, that is, low-frequency components and high-frequency components;

[0017] Then, perform soft thresholding to suppress the noise of the detail coefficients and retain the key signal components. The soft threshold function is defined as:

[0018]

[0019] where the threshold N is the signal length, c is the coefficient after wavelet transform, and the signal is denoised by setting the threshold T;

[0020] The noise standard deviation σ is estimated by the median absolute deviation:

[0021]

[0022] where the constant 0.6745 is the conversion factor between MAD and the standard deviation in the Gaussian distribution, d1 is the detail coefficient of the first layer in the wavelet transform, the threshold T is positively correlated with the noise level σ. The greater the noise, the higher the threshold, and the more aggressive the denoising. Conversely, more details are retained;

[0023] Finally, perform signal reconstruction to reconstruct the detail coefficients and approximation coefficients after thresholding into a denoised signal through inverse wavelet transform:

[0024]

[0025] where c0 is the approximation coefficient, c0, c1, …, c n are the detail coefficients after thresholding.

[0026] Further, the dataset division and signal segmentation adopt a subject-dependent division strategy, that is, the sEMG data of each subject is used to independently train and validate the model. The subject-dependent division strategy can generate standardized inputs while retaining the temporal continuity. Specifically:

[0027] In the data of each subject, the 1st, 3rd, 4th, and 6th repetitions of each gesture are used as the training set, and the 2nd and 5th repetitions are used as the validation set, enabling the model to learn from diverse gesture samples and test on unseen repeated data to ensure the fairness of evaluation;

[0028] Independent training and validation for each subject helps the model adapt to the individual sEMG characteristics and avoid cross-subject interference. In addition, the resting state data is removed to avoid class imbalance, enabling the model to focus more on distinguishing valid gestures. Finally, the classification accuracy of each subject is calculated separately, and the average value is taken to evaluate the overall performance;

[0029] To address the problem of the variable length of repeated signals, the sliding window method is adopted to segment the sEMG signals into fixed-length segments. The window length is 300 ms, and the adjacent windows overlap by 250 ms.

[0030] Furthermore, the signal conversion uses the short-time Fourier transform, that is, the short-time Fourier transform is applied to each data window of all channels to convert the time-domain signal into a time-frequency representation, which is crucial for capturing the dynamic characteristics of sEMG. Specifically as follows:

[0031] For each 300-ms window, the STFT is calculated using a Hamming window with a length of 128 / 2000 ms and a step size of 32 / 2000 ms, resulting in 12 time-frequency representations, one for each channel. The size of each time-frequency representation is 65×19, where 65 represents the number of frequency bins and 19 represents the number of time frames;

[0032] Then, the 12 time-frequency representations are concatenated along the channel dimension to obtain a representation with a final size of 12×65×19;

[0033] To ensure that the model focuses on the relative changes rather than the absolute values of the signals, the min-max normalization method for each channel is used to normalize the short-time Fourier transform representation to the range [0,1].

[0034] Furthermore, the data augmentation is to further enhance the generalization ability of the model by applying the masking strategy twice to enhance the training data. Specifically:

[0035] In each data augmentation process, 10% of the STFT representation is randomly masked, that is, the corresponding frequency bins are set to zero, and 15% of the time frames are randomly masked, that is, the corresponding time bins are set to zero.

[0036] Furthermore, the DSRANet model starts from the initial stage. At this stage, a 3×3 depth convolution is applied to the input data;

[0037] X DW = DWConv 3×3 (I) ∈ R C×F×T (4)

[0038] where I is the input time-frequency representation, and R represents the space where the data is located;

[0039] Subsequently, pointwise convolution is performed to convert the input data from C = 12 channels to C initial = 32 channels;

[0040] Next, batch normalization and ReLU activation are performed on the output to introduce non-linearity and stabilize the training;

[0041]

[0042] After the initial stage, the DSRANet module enters three consecutive stages, namely three-stage feature processing. In each stage, the number of channels is increased by the DSR module, the frequency resolution is reduced by max pooling along the frequency axis, and MACA is applied to adaptively focus on relevant features. The specific operations are expressed as:

[0043]

[0044] where \(i = \{1, 2, 3\}\) is the index of the stage, \(X_{stage(0)} = X_{initial}\), and \(C\) i is the number of channels in stage \(i\). Specifically, \(C_1 = 64\), \(C_2 = 128\), and \(C_3 = 256\);

[0045] The output of the final stage of the DSRANet module reduces the spatial dimension of each channel to a single value through global average pooling. The pooled output generates class predictions through a fully connected layer. The specific operations are expressed as:

[0046]

[0047] X DSRANet = FC(X GAP ) \(\in \mathbb{R}\) N (11)

[0048] where \(N\) is the number of classes in the dataset.

[0049] Furthermore, in the DSR block, the input first passes through a 3×3 DWConv, and then a PConv, as follows:

[0050]

[0051] where \(X\) in is the input of the DSR block, \(C\) in is the number of input channels, \(C\) out is the number of output channels, \(F\) in and \(T\) in are the number of input frequency bins and the number of time frames, respectively;

[0052] Next, batch normalization and ReLU activation are applied to the output to introduce non-linearity:

[0053]

[0054] The combination of 3×3 DWConv and PConv is applied again to the output of the ReLU activation, followed by batch normalization:

[0055]

[0056] The DSR block incorporates residual connections, which combine the improved features with the original input. The design of the residual connection helps to directly transmit gradients through the network, preventing the problem of gradient vanishing. The residual connection is expressed as:

[0057]

[0058]

[0059] Since the input and output channels are different, first, 3×3 DWConv and PConv are applied to the input to match the dimensions. After batch normalization of the result, it is added to the output of the convolutional layer. Finally, ReLU activation is applied to the combined output.

[0060] Furthermore, the MACA module includes spatial attention and channel attention;

[0061] The spatial attention is calculated by independently applying depthwise convolution and pointwise convolution along the frequency axis and the time axis, thereby obtaining frequency attention and time attention. The frequency attention is calculated by applying depthwise separable convolution with a kernel size of (7,1) along the frequency axis, while the time attention is calculated by applying convolution with a kernel size of (1,7) along the time axis;

[0062] The spatial attention mechanism can be expressed as:

[0063]

[0064] where σ represents the sigmoid activation function, represents element-wise multiplication, C in 、F in and T in represent the number of channels, the number of frequency bins, and the number of time frames respectively;

[0065] The channel attention mechanism generates a channel-level feature descriptor through global average pooling operation along the spatial dimension. Then, the feature descriptor passes through a fully connected layer with a sigmoid activation function to generate channel-level attention weights. The channel attention mechanism is expressed as:

[0066]

[0067] where FC1 and FC2 are fully connected layers, and σ represents the sigmoid activation function;

[0068] The final output of the MACA module is obtained by combining the element-wise multiplication of the spatial attention and channel attention mechanisms:

[0069]

[0070] Among them represents element-wise multiplication.

[0071] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0072] 1. The present invention proposes DSRANet, a lightweight depthwise separable residual attention network, which can effectively capture the spatio-temporal features of sEMG gestures while maintaining computational efficiency (0.459M parameters, 0.1G FLOPs).

[0073] 2. The present invention designs the DSR block, as an innovative architecture component, which efficiently extracts discriminative features by combining depthwise separable convolution and residual connection. Ablation experiments show that its computational complexity is reduced by 45% compared with standard convolution.

[0074] 3. The present invention introduces the MACA mechanism to adaptively enhance key features across channels and time-frequency dimensions. Comparative experiments verify that it can improve the model accuracy by 6.41%.

[0075] 4. The present invention verifies its effectiveness and achieves the current best results (DB2: 92.13%, DB3: 75.65%, DB4: 91.00%) on the NinaPro dataset, which are improved by 4.1%, 5.07% and 1.57% respectively compared with the existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0077] Figure 1 It is a schematic diagram of gesture categories in the NinaPro dataset of the present invention;

[0078] Figure 2 It is a schematic diagram of the signal preprocessing process of the present invention;

[0079] Figure 3 It is a schematic diagram of the comparison of STFT representations of sEMG signals (channel 5) of the present invention (300 ms time window);

[0080] Figure 4 It is a schematic diagram of the depthwise separable residual attention network DSRANet of the present invention;

[0081] Figure 5 It is a schematic diagram of the difference between depth convolution, pointwise convolution and standard convolution of the present invention;

[0082] Figure 6 Schematic diagram of the architecture of the DSR block of the present invention;

[0083] Figure 7 Architecture of the multi-axis channel attention (MACA) mechanism of the present invention;

[0084] Figure 8 Comparison of the number of parameters between the DSRANet of the present invention and other latest methods;

[0085] Figure 9 Box plot of the classification accuracy of the DSRANet of the present invention on the NinaPro DB2, DB3, and DB4 datasets;

[0086] Figure 10 Schematic diagram of the gradient-weighted class activation mapping (Grad-CAM) visualization of the DSRANet of the present invention;

[0087] Figure 11 Schematic diagram of the projected features of the present invention;

[0088] Figure 12 Schematic diagram of the results of the anti-noise experiment of the present invention. Detailed implementation manners

[0089] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0090] The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network at least includes the following steps:

[0091] Refer to Figure 2 , S1: Preprocess the signal. In order to remove noise and enhance the signal-to-noise ratio, preprocess the original sEMG signal. The preprocessing process at least includes removing noise in the original data, dataset division and signal segmentation, signal conversion, and data augmentation; Figure 3 Shows the STFT representations of the original sEMG signal, the denoised sEMG signal, and the enhanced STFT representation, where (a) is the original signal, (b) is the denoised signal, and (c) is the enhanced STFT representation.

[0092] S2: Construct the DSRANet model. The DSRANet model is composed of a depthwise separable residual block (DSRBlock) and a multi-axis channel attention (MACA) module to achieve efficient feature extraction and adaptive focusing. The overall architecture is as shown in Figure 4As shown, the DSRANet model takes time-frequency representations as input. The shape of the time-frequency representation input data of the DSRANet model is C×F×T, where C is the number of channels, F is the number of frequency points, and T is the number of time frames. The applications after the time-frequency representation obtained by signal preprocessing are input into the DSRANet model, including the initial stage, three-stage feature processing, and classification stage;

[0093] S3: Feature extraction is implemented using a depthwise separable residual block (DSRBlock) to achieve efficient feature extraction. The depthwise separable residual block is the DSR block (refer to Figure 6 ), and the DSR block integrates depthwise separable convolution and residual connection. The depthwise separable convolution includes two types of convolutions: DWConv that independently processes each input channel and PConv that aggregates information across channels, which can significantly reduce the number of parameters and computational costs. Figure 5 The difference between DWConv, PConv, and standard convolution is shown. The residual connection helps train deep networks and improve feature learning effects;

[0094] S4: Apply the attention mechanism. Use a multi-axis channel attention module such as Figure 7 shown, that is, the MACA module. The MACA module aims to adaptively focus on relevant features in the spatial and channel dimensions.

[0095] Removing the noise in the original data specifically means eliminating the noise and artifacts in the original sEMG signal. The digital filter combination of a 50Hz notch filter and a 500Hz low-pass filter is used to remove the noise in the original data, and the method used is wavelet denoising;

[0096] The 50Hz notch filter is used to eliminate power frequency interference, which is a common noise source in sEMG signals;

[0097] The 500Hz low-pass filter is used to filter out high-frequency noise, which can remove the high-frequency components in the signal and ensure the smoothness of the signal;

[0098] Wavelet denoising uses the wavelet transform method based on the Daubechies mother wavelet to decompose and reconstruct the signal, which can effectively remove the noise and retain the main features of the signal. The order of the wavelet transform method based on the Daubechies mother wavelet is 4, and the decomposition level is 4.

[0099] Removing the noise in the original data includes at least the following steps:

[0100] First, perform wave decomposition to decompose the signal into approximation coefficients and detail coefficients, that is, low-frequency components and high-frequency components;

[0101] Then, perform soft threshold processing to suppress the noise of the detail coefficients and retain the key signal components. The soft threshold function is defined as:

[0102]

[0103] Among them, the threshold N is the signal length, c is the coefficient after wavelet transform, and the signal is denoised by setting the threshold T;

[0104] The standard deviation σ of the noise is estimated by the median absolute deviation (MAD):

[0105]

[0106] Among them, the constant 0.6745 is the conversion factor between MAD and the standard deviation in the Gaussian distribution, d1 is the detail coefficient of the first layer in the wavelet transform, the threshold T is positively correlated with the noise level σ, the greater the noise, the higher the threshold, and the more radical the denoising. Conversely, more details are retained;

[0107] Finally, signal reconstruction is performed, and the detail coefficients and approximation coefficients after threshold processing are reconstructed into a denoised signal through inverse wavelet transform:

[0108]

[0109] Among them, c0 is the approximation coefficient, c0, c1, …, c n are the detail coefficients after thresholding. Experiments show that this method can effectively suppress noise and retain the physiological characteristics of the sEMG signal. The latest research shows that wavelet transform can effectively separate and remove persistent noise components. This denoising method significantly improves the signal-to-noise ratio in various biomedical applications including sEMG.

[0110] The dataset division and signal segmentation adopt a subject-dependent division strategy, that is, the sEMG data of each subject is used to independently train and validate the model. The subject-dependent division strategy can generate standardized inputs while retaining the temporal continuity. Specifically:

[0111] In the data of each subject, the 1st, 3rd, 4th, and 6th repetitions of each gesture are used as the training set, and the 2nd and 5th repetitions are used as the validation set, enabling the model to learn from diverse gesture samples and test on unseen repeated data to ensure evaluation fairness;

[0112] Independent training and validation for each subject helps the model adapt to the individual sEMG characteristics and avoid cross-subject interference. In addition, the resting state data is removed to avoid class imbalance, making the model more focused on differentiating effective gestures. Finally, the classification accuracy of each subject is calculated separately, and then the average value is taken to evaluate the overall performance;

[0113] To address the problem of variable lengths of repeated signals, the sliding window method is adopted to segment the sEMG signals into fixed-length segments. The window length is 300 ms, and adjacent windows overlap by 250 ms.

[0114] The signal conversion uses the short-time Fourier transform, that is, the short-time Fourier transform (STFT) is applied to each data window of all channels to convert the time-domain signal into a time-frequency representation, which is crucial for capturing the dynamic characteristics of sEMG. Specifically as follows:

[0115] For each 300-ms window, the STFT is calculated using a Hamming window with a length of 128 / 2000 ms and a step size of 32 / 2000 ms, resulting in 12 time-frequency representations, one for each channel. The size of each time-frequency representation is 65×19, where 65 represents the number of frequency bins and 19 represents the number of time frames;

[0116] Then, the 12 time-frequency representations are concatenated along the channel dimension to obtain a final representation with a size of 12×65×19;

[0117] To ensure that the model focuses on the relative changes rather than the absolute values of the signals, the min-max normalization method for each channel is used to normalize the short-time Fourier transform representation to the range [0,1].

[0118] Data augmentation is to further enhance the generalization ability of the model. The training data is enhanced by applying the masking strategy twice. Specifically:

[0119] In each data augmentation process, 10% of the STFT representation is randomly masked, that is, the corresponding frequency bins are set to zero, and 15% of the time frames are randomly masked, that is, the corresponding time bins are set to zero;

[0120] This method is inspired by the Dropout idea, and the robustness is enhanced by simulating noise and data loss. The two augmentations further suppress overfitting.

[0121] The DSRANet model starts from the initial stage. At this stage, a 3×3 depthwise convolution (DWConv) is applied to the input data;

[0122] X DW =DWConv 3×3 (I)∈R C×F×T (4)

[0123] where I is the input time-frequency representation and R represents the space where the data is located;

[0124] Subsequently, pointwise convolution (PConv) is performed to convert the input data from C = 12 channels to Cinitial = 32 channels;

[0125]

[0126] Next, batch normalization and ReLU activations are applied to the output to introduce non-linearity and stabilize the training;

[0127]

[0128] After the initial stage, the DSRANet module enters three consecutive stages, namely three-stage feature processing. In each stage, the number of channels is increased by the DSR module, the frequency resolution is reduced by max-pooling along the frequency axis, and MACA is applied to adaptively focus on relevant features. The specific operations are as follows:

[0129]

[0130] where i = {1, 2, 3} is the index of the stage, Xstage(0) = Xinitial, and C i is the number of channels in stage i. Specifically, C1 = 64, C2 = 128, and C3 = 256;

[0131] The output of the final stage of the DSRANet module is globally average-pooled to reduce the spatial dimension of each channel to a single value. The pooled output passes through a fully connected layer to generate class predictions. The specific operations are as follows:

[0132]

[0133] X DSRANet = FC(X GAP ) ∈ R N (11)

[0134] where N is the number of classes in the dataset;

[0135] To clearly show the shape changes and key components of each stage of the DSRANet model, Table 1 lists the input and output shapes of each stage and highlights the operations applied.

[0136] Table 1 Shape changes and key components of each stage in DSRANet

[0137]

[0138] In the DSR block, the input first passes through a 3×3 depthwise convolution (DWConv), followed by a pointwise convolution (PConv), as follows:

[0139]

[0140] Among them, X in is the input of the DSR block, C in is the number of input channels, C out is the number of output channels, F in and T in are the number of input frequency bins and the number of time frames respectively, and these values vary according to the stage of the DSRANet model;

[0141] Next, batch normalization and ReLU activation are applied to the output to introduce non-linearity:

[0142]

[0143] The combination of 3×3 DWConv and PConv is applied again to the output of the ReLU activation, followed by batch normalization:

[0144]

[0145] Residual connections are incorporated into the DSR block to combine the improved features with the original input. The design of the residual connections helps to directly transmit gradients through the network and prevent the vanishing gradient problem. The residual connection is expressed as:

[0146]

[0147] Since the input and output channels are different, 3×3 DWConv and PConv are first applied to the input to match the dimensions. After batch normalization of the result, it is added to the output of the convolutional layer. Finally, ReLU activation is applied to the combined output.

[0148] The MACA module includes spatial attention and channel attention;

[0149] Spatial attention is calculated by independently applying depthwise convolution and pointwise convolution along the frequency axis and the time axis, thereby obtaining frequency attention and time attention. Frequency attention is calculated by applying depthwise separable convolution with a kernel size of (7,1) along the frequency axis, while time attention is calculated by applying convolution with a kernel size of (1,7) along the time axis;

[0150] The spatial attention mechanism can be expressed as:

[0151]

[0152] Among them, σ represents the sigmoid activation function, Denotes element-wise multiplication, C in , F in and T in represent the number of channels, the number of frequency bins, and the number of time frames respectively. These values remain unchanged in the MACA module and vary according to the stage of DSRANet;

[0153] The channel attention mechanism generates a channel-level feature descriptor through global average pooling operation along the spatial dimension. Then, the feature descriptor passes through a fully connected layer with a sigmoid activation function to generate channel-level attention weights. The channel attention mechanism is expressed as:

[0154]

[0155] where FC1 and FC2 are fully connected layers, and σ represents the sigmoid activation function;

[0156] The final output of the MACA module is obtained by combining element-wise multiplication of the spatial attention and channel attention mechanisms:

[0157]

[0158] where denotes element-wise multiplication.

[0159] Based on the above embodiment content, the following verification process is proposed:

[0160] Experimental setup:

[0161] The proposed model DSRANet is implemented using PyTorch and trained on a single NVIDIA GeForce RTX4080 GPU. The hyperparameter settings are shown in Table 2 for details.

[0162] Table 2 Model hyperparameter settings

[0163]

[0164]

[0165] Experimental results:

[0166] To evaluate the performance of the proposed DSRANet, the present invention compares it with the state-of-the-art methods on the NinaPro DB2, DB3, and DB4 datasets. See Figure 9. The comparison is based on classification accuracy, which is a commonly used metric in sEMG-based gesture recognition research. Table 3 summarizes the comparison results, highlighting the excellent performance of DSRANet on all three datasets. The proposed model achieved an accuracy of 92.13% on the NinaPro DB2 dataset, 75.65% on the DB3 dataset, and 91.00% on the DB4 dataset, improving by 4.1%, 5.07%, and 1.57% respectively compared to the existing best methods on their respective datasets. This demonstrates the effectiveness of the proposed DSRBlock and MACA modules in capturing relevant features from sEMG signals and improving classification accuracy.

[0167] Table 3 Comparison of the proposed DSRANet with the state-of-the-art methods on the NinaPro DB2, DB3, and DB4 datasets

[0168]

[0169] Since the baseline methods did not provide open-source code or replicated experimental data, which hindered the calculation of baseline variance or the construction of data distribution, the metrics of the previous best models were obtained from their reported results. Since paired t-tests or Wilcoxon tests require two sets of paired data, these methods were not applicable. Therefore, a one-sample t-test was used to verify the statistical significance of the improvement compared to the reported metrics of the previous best models. Specifically, using the performance values of the best models as the baseline reference, the p-values for the DB2, DB3, and DB4 datasets were 0.0098, 0.0846, and 0.0251 respectively, indicating that the improvement of DSRANet is statistically significant.

[0170] Figure 8 And Table 3 shows the comparison of the number of parameters and classification accuracy between the proposed DSRANet and other state-of-the-art methods. Compared with the existing methods, the proposed model achieved excellent classification accuracy with fewer parameters, demonstrating its efficiency and effectiveness in sEMG-based gesture recognition. Specifically, DSRANet has 660,000 parameters, significantly lower than the number of parameters of other models such as MvCNN (4.03 million), LI-TFMNet (1.98 million), and NKDFFCNN (1.05 million), while achieving higher classification accuracy on the NinaPro DB2, DB3, and DB4 datasets.

[0171] Benchmark datasets:

[0172] The present invention uses the NinaPro dataset, which is one of the largest publicly available sEMG databases and is widely used to evaluate various human-machine interaction (HMI)-based methods. This dataset contains ten sub-datasets, labeled DB1 to DB10. In the present invention, the NinaPro DB2 sub-dataset is mainly used to verify the effectiveness of the proposed MSDS-FusionNet. This dataset contains 12-channel upper limb sEMG data of 40 healthy subjects performing 49 hand movements (including the resting state) at a sampling frequency of 2000 Hz. The 12 sEMG channels correspond to specific electrode placements: columns 1-8 are electrodes evenly distributed on the forearm around the height of the radiohumeral joint; columns 9 and 10 capture the main activity signals of the flexor digitorum superficialis and extensor digitorum superficialis; columns 11 and 12 record the main activity signals of the biceps brachii and triceps brachii.

[0173] In addition, the present invention also incorporates NinaPro DB3 (containing sEMG data of 11 forearm amputees) and NinaPro DB4 (containing data of 10 non-amputee individuals). Table 4 provides details of the three sub-datasets used in the present invention, Figure 1 showing the gesture categories in the NinaPro dataset.

[0174] Table 4 Overview of the NinaPro DB2, DB3, and DB4 Datasets

[0175]

[0176] Ablation Experiments

[0177] To verify the effectiveness of the proposed DSRBlock and MACA modules, ablation experiments were conducted on the NinaPro DB2 dataset. The experiments involved training and evaluating three variants of DSRANet: DSRANet without DSRBlocks (replaced with standard 3×3 convolutions), DSRANet without MACA, and the complete DSRANet model. The DSRBlock enhances feature extraction by integrating depthwise separable convolutions and residual connections, while the MACA module optimizes performance by adaptively emphasizing significant features in the spatial and channel dimensions. These modules are crucial for accurate gesture recognition.

[0178] Table 5 Results of the Ablation Experiments on the NinaPro DB2 Dataset

[0179]

[0180] Table 5 summarizes the results of the ablation experiments, highlighting the importance of DSRBlocks and MACA in improving the classification accuracy. The complete DSRANet model achieved an accuracy of 92.13%, outperforming variants without DSRBlocks (89.13%) and MACA (85.72%). These results demonstrate the effectiveness of the proposed DSRBlock and MACA modules in enhancing the model performance.

[0181] The observed improvements can be attributed to the following factors:

[0182] Enhanced feature representation ability: The DSRBlock effectively captures the spatial and channel dependencies in the sEMG signals through depthwise separable convolutions and residual connections, thus extracting more discriminative features. This design not only improves the feature representation ability but also significantly reduces the model complexity, decreasing the number of trainable parameters and computational cost.

[0183] Reduced interference from irrelevant information: The MACA module dynamically assigns importance to different time-frequency components through spatial and channel attention mechanisms, prioritizing informative channels and suppressing noise. This dual attention mechanism reduces the interference from irrelevant information, enabling the model to focus on more discriminative features, thereby improving the robustness and generalization ability of the features.

[0184] Improved model robustness: The introduction of residual connections and attention mechanisms enables the DSRANet to exhibit stronger robustness in the face of complex signals and noise. By enhancing the accuracy and stability of feature extraction, the model performs more consistently on different samples, further improving its generalization ability on unseen data.

[0185] The significant improvements observed in the experimental models validate the effectiveness of the DSRBlock and MACA modules. By integrating these modules, the DSRANet achieved superior classification results on the NinaPro DB2 dataset. The DSRBlock and MACA modules play a crucial role in enhancing the performance of the DSRANet, enabling the model to focus on important features while alleviating the interference from irrelevant information, thereby improving the accuracy and reliability of the gesture recognition task.

[0186] Grad-CAM Visualization

[0187] To gain an in-depth understanding of the feature extraction process of the DSRANet, Gradient-weighted Class Activation Mapping (Grad-CAM) is used to visualize the activations of the DSRBlock and MACA. Grad-CAM is a powerful technique that provides a visual representation of the regions in the input data that contribute the most to the model's predictions. By analyzing these visualization results, one can better understand the relevant features that the DSRANet focuses on during the classification process.

[0188] Figure 10 The Grad-CAM visualization results in [reference] provide valuable insights into the decision-making process of the model, highlighting the regions in the input time-frequency representation that contribute most to the classification result.

[0189] Figure 10 (a-c) show the activation maps of the DSRBlock at different stages, while Figure 10 (d-f) illustrate the attention maps generated by the MACA module at the corresponding stages.

[0190] The heatmaps generated for the DSRBlock and MACA module at different processing stages show how the network gradually refines its focus on relevant features. Initially, the DSRBlock and MACA module exhibit broad sensitivity, with the heatmaps showing multiple small, scattered regions of moderate activation. These early patterns suggest that the network is capturing basic, local features across the entire input - such as edges or specific frequency - time patterns. As processing progresses, the DSRBlock begins to refine these features, transitioning to a more targeted representation. Its activations gradually evolve into long strips emphasizing persistent frequency patterns, eventually converging to a highly concentrated region, which may represent a key gesture - specific feature.

[0191] Meanwhile, based on these initial detections, the MACA module narrows its focus from a diverse set of activations to more defined and discriminative regions. The shift in MACA is manifested as a concentration of attention, with the module transitioning from broad area responses to isolating a distinct, vertically oriented strip region. This specific pattern implies a targeted focus on critical moments across multiple frequencies - perhaps sudden muscle contractions - which is crucial for accurate gesture classification.

[0192] t-SNE Feature Space Visualization and KDE Analysis

[0193] To verify the discriminative ability of the features learned by DSRANet, t-SNE is used for dimensionality reduction and KDE for density estimation. As Figure 11 shown, the projected features reveal a well - structured manifold where 49 sEMG gesture classes are organized into distinct clusters. The KDE contours further quantify the density distribution, with compact high - probability regions indicating strong within - class cohesion. Three key observations are drawn from the visualization.

[0194] First, most of the classes form isolated clusters with minimal overlap, indicating that DSRANet has the ability to extract discriminative features that maximize the inter-class margin. This strong separability highlights the effectiveness of the model in differentiating different sEMG gesture patterns. Second, the tightly bounded KDE profiles for many classes reflect low intra-class variance, indicating strong feature invariance to sEMG signal fluctuations. This consistency highlights the model's ability to generalize variations within the same gesture class. Finally, the limited overlap between elongated profiles is consistent with gesture pairs sharing similar muscle activation patterns, such as finger flexion and extension. Importantly, since t-SNE preferentially preserves local structure, these overlaps may not persist in the original high-dimensional space where the model operates, suggesting that the observed ambiguity may not significantly affect the model's performance.

[0195] This analysis validates that DSRANet has learned a feature space that balances discriminative and generalization capabilities. Clear cluster separation combined with local density peaks confirms the effectiveness of the model in capturing sEMG gesture semantics. Through mapping potentially ambiguous regions, visualization also guides targeted model improvements, such as adding training data for overlapping classes.

[0196] Robustness Test

[0197] To evaluate the generalization and robustness of the proposed DSRANet, a noise-resistant experiment was conducted using sample subjects from the NinaPro DB2 dataset. The model was trained on clean data and tested on data with Gaussian noise added at different signal-to-noise ratio (SNR) levels ranging from 10 dB to 50 dB. The signal-to-noise ratio is defined as the ratio of signal power to noise power, with the formula:

[0198]

[0199] where P signal and P noise are the signal power and noise power, respectively.

[0200] The results of the noise-resistant experiment are as Figure 12 shown in the figure, which presents the classification accuracy of DSRANet at different SNR levels. The experimental results indicate that the model has high robustness, with the accuracy remaining above 85% even at a low SNR (20 dB). This shows that DSRANet can effectively handle noisy data and is suitable for practical application scenarios where sEMG signals may be interfered by various noise sources.

[0201] Experimental Results

[0202] The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network proposed a lightweight deep learning architecture named DSRANet for gesture recognition based on surface electromyogram (sEMG) signals. This architecture is centered around novel depthwise separable residual blocks (DSRBlocks) and a multi-axis channel attention (MACA) mechanism. DSRBlocks effectively address the computational bottleneck by decomposing traditional convolutions into depthwise operations (processing each channel independently) and then efficiently integrating information across channels through pointwise convolutions. The introduction of residual connections helps with gradient flow, enabling the network to be trained deeper without performance degradation. The MACA mechanism significantly enhances the discriminability of features by adaptively concentrating attention in both spatial and channel dimensions. It is achieved through a dual path: using dedicated convolutions to process the frequency and time axes respectively, while generating channel attention weights through global pooling, enabling the network to dynamically emphasize the most discriminative spectro-temporal patterns and channels.

[0203] The synergistic combination of these components enables DSRANet to efficiently capture the dynamic spectral patterns crucial for differentiating complex hand movements. Ablation experiments show that removing DSRBlocks leads to a 3% decrease in accuracy and a 45% increase in parameters, while removing MACA causes a 6.4% performance drop, highlighting its key role in feature selection. Therefore, DSRANet achieves state-of-the-art accuracies of 92.13%, 75.65%, and 91.00% on the NinaPro DB2, DB3, and DB4 datasets respectively, while maintaining a lightweight model size (0.459M parameters, 0.1 GFLOPs), suitable for real-time applications in prosthetics and rehabilitation systems. Ablation experiments confirm the necessity of each architecture component: the absence of DSRBlocks or MACA results in a 3% - 6.4% decrease in accuracy, emphasizing their role in efficient feature extraction and adaptive attention. Visualization through Grad-CAM shows that DSRANet gradually refines its attention to the spectro-temporal patterns of specific gestures, while robustness tests indicate that its performance can still remain above 85% under low signal-to-noise ratio (20dB) conditions, highlighting its resistance to noise. These findings validate the model's role as a bridge between laboratory-level accuracy and real-world practicality.

[0204] Future work will explore extending DSRANet to multimodal sensor fusion, optimizing it for embedded hardware, and validating its performance in clinical settings, including amputees and rehabilitation patients. Additionally, integrating interpretable artificial intelligence techniques can further clarify the relationship between learned features and neuromuscular dynamics, enhancing trust and adaptability in the human-machine interface. By balancing accuracy, efficiency, and interpretability, DSRANet drives the development of responsive and accessible sEMG-driven technologies with the potential to transform healthcare and human-machine interaction.

[0205] In summary:

[0206] DSRANet significantly improves the accuracy and efficiency of gesture recognition in sEMG signals through depthwise separable convolutions and a multi-axis channel attention mechanism, including signal preprocessing, model construction, feature extraction, and application of the attention mechanism. In the signal preprocessing stage, the original sEMG signal is mainly denoised, segmented, and subjected to short-time Fourier transform (STFT) to generate a time-frequency representation. In the model construction stage, the DSRANet model is proposed, which consists of depthwise separable residual blocks (DSRBlocks) and a multi-axis channel attention (MACA) module to achieve efficient feature extraction and adaptive focusing. In the feature extraction stage, spatio-temporal features are efficiently extracted through depthwise separable residual modules (DSRBlocks) to reduce computational complexity. In the application stage of the attention mechanism, the multi-axis channel attention (MACA) module dynamically focuses on key features to enhance feature discrimination ability. The present invention innovatively addresses the deficiencies of traditional sEMG gesture recognition models in processing complex signals by combining depthwise separable convolutions and a multi-axis channel attention mechanism, significantly improving the accuracy and efficiency of recognition, especially in dynamic and complex sEMG signal scenarios. On the NinaPro DB2, DB3, and DB4 datasets, DSRANet achieved accuracies of 92.13%, 75.65%, and 91.00% respectively, improving by 4.1%, 5.07%, and 1.57% compared to the previous best methods, demonstrating its excellent performance.

[0207] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A sEMG gesture recognition method based on a lightweight depthwise separable residual attention network, characterized in that: At least include the following steps: S1: Preprocess the signal. In order to remove noise and enhance the signal-to-noise ratio, preprocess the original sEMG signal. The preprocessing process at least includes removing noise in the original data, dataset division and signal segmentation, signal conversion, and data augmentation; S2: Construct the DSRANet model. The DSRANet model is composed of depthwise separable residual blocks and multi-axis channel attention modules to achieve efficient feature extraction and adaptive focusing. The DSRANet model takes the time-frequency representation as the input. The shape of the time-frequency representation input data of the DSRANet model is C×F×T, where C is the number of channels, F is the number of frequency points, and T is the number of time frames. The applications after the time-frequency representation obtained by signal preprocessing is input into the DSRANet model include the initial stage, three-stage feature processing, and classification stage; S3: Feature extraction. Use depthwise separable residual blocks (DSRBlocks) to achieve efficient feature extraction. The depthwise separable residual block is the DSR block. The DSR block integrates depthwise separable convolution and residual connection. The depthwise separable convolution includes DWConv that independently processes each input channel and PConv that aggregates information across channels, which can significantly reduce the number of parameters and computational cost. The residual connection helps train deep networks and improve the feature learning effect; S4: Apply the attention mechanism. Use the multi-axis channel attention module, that is, the MACA module. The MACA module aims to adaptively focus on relevant features in the spatial and channel dimensions.

2. The sEMG gesture recognition method based on the lightweight depth separable residual attention network according to claim 1, wherein: Removing the noise in the original data specifically means eliminating the noise and artifacts in the original sEMG signal. The digital filter combination of a 50Hz notch filter and a 500Hz low-pass filter is used to remove the noise in the original data. The method used is wavelet denoising; The 50Hz notch filter is used to eliminate power frequency interference, that is, a common noise source in sEMG signals; The 500Hz low-pass filter is used to filter out high-frequency noise, which can remove the high-frequency components in the signal and ensure the smoothness of the signal; The wavelet denoising uses the wavelet transform method based on Daubechies mother wavelet to decompose and reconstruct the signal, which can effectively remove noise and retain the main features of the signal. The order of the wavelet transform method based on Daubechies mother wavelet is 4, and the decomposition layer is 4.

3. The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network according to claim 2, wherein: Removing the noise in the original data at least includes the following steps: First, perform wave decomposition to decompose the signal into approximation coefficients and detail coefficients, that is, low-frequency components and high-frequency components; Then perform soft threshold processing to suppress noise in the detail coefficients and retain key signal components. The soft threshold function is defined as: Among them, the threshold value N is the signal length, c is the coefficient after wavelet transform, and the signal is denoised by setting the threshold value T; The noise standard deviation σ is estimated by the median absolute deviation: Among them, the constant 0.6745 is the conversion factor between MAD and the standard deviation in the Gaussian distribution. d1 is the detail coefficient of the first layer in the wavelet transform. The threshold T is positively correlated with the noise level σ. The greater the noise, the higher the threshold, and the more aggressive the denoising. Conversely, more details are retained; Finally, signal reconstruction is performed. The detail coefficients and approximation coefficients after threshold processing are reconstructed into a denoised signal through inverse wavelet transform: Among them, c0 is the approximation coefficient, and c0, c1, …, c n are the detail coefficients after thresholding.

4. The sEMG gesture recognition method based on the lightweight depth separable residual attention network according to claim 2, wherein: The dataset division and signal segmentation adopt a subject-dependent division strategy, that is, the sEMG data of each subject is used to independently train and validate the model. The subject-dependent division strategy can generate standardized inputs while retaining temporal continuity. Specifically: In the data of each subject, the 1st, 3rd, 4th, and 6th repetitions of each gesture are used as the training set, and the 2nd and 5th repetitions are used as the validation set, enabling the model to learn from diverse gesture samples and be tested on unseen repeated data to ensure evaluation fairness; Independent training and validation for each subject helps the model adapt to the individual sEMG characteristics and avoid cross-subject interference. In addition, the resting state data is removed to avoid class imbalance, making the model more focused on distinguishing valid gestures. Finally, the classification accuracy of each subject is calculated separately, and then the average value is taken to evaluate the overall performance; To address the problem of variable length of repeated signals, the sliding window method is used to segment the sEMG signal into fixed-length segments. The window length is 300 ms, and the adjacent windows overlap by 250 ms.

5. The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network according to claim 4, wherein: The signal transformation uses the short-time Fourier transform, that is, the short-time Fourier transform is applied to each data window of all channels to convert the time-domain signal into a time-frequency representation, which is crucial for capturing the dynamic characteristics of sEMG. Specifically as follows: For each 300-millisecond window, the STFT is calculated using a Hamming window with a length of 128 / 2000 milliseconds and a step size of 32 / 2000 milliseconds, resulting in 12 time-frequency representations, one for each channel. The size of each time-frequency representation is 65×19, where 65 represents the number of frequency bins and 19 represents the number of time frames; Then, the 12 time-frequency representations are concatenated along the channel dimension to obtain a final representation with a size of 12×65×19; To ensure that the model focuses on the relative changes rather than the absolute values of the signals, the min-max normalization method for each channel is used to normalize the short-time Fourier transform representation to the range [0,1].

6. The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network according to claim 5, wherein: The data augmentation is to further enhance the generalization ability of the model. The training data is enhanced by applying the masking strategy twice. Specifically: In each data augmentation process, 10% of the STFT representation is randomly masked, that is, the corresponding frequency bins are set to zero, and 15% of the time frames are randomly masked, that is, the corresponding time bins are set to zero.

7. The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network according to claim 1, characterized in that: The DSRANet model starts from the initial stage. In this stage, a 3×3 depth convolution is applied to the input data; X DW = DWConv 3×3 (I) ∈ R C×F×T (4) Among them, I is the input time-frequency representation, and R represents the space where the data is located; Subsequently, pointwise convolution is performed to convert the input data from C = 12 channels to C initial = 32 channels; Next, batch normalization (and ReLU activation are performed on the output to introduce nonlinearity and stabilize the training; After the initial stage, the DSRANet module enters three consecutive stages, namely three-stage feature processing. In each stage, the number of channels is increased by the DSR module, the frequency resolution is reduced by max pooling along the frequency axis, and MACA is applied to adaptively focus on relevant features. The specific operations are as follows: where \(i = \{1, 2, 3\}\) is the index of the stage, \(X_{stage(0)} = X_{initial}\), and \(C\) i is the number of channels in stage \(i\). Specifically, \(C_1 = 64\), \(C_2 = 128\), and \(C_3 = 256\); The output of the final stage of the DSRANet module is reduced to a single value for each channel's spatial dimension through global average pooling. The pooled output generates class predictions through a fully connected layer. The specific operations are as follows: X DSRANet = FC(X GAP ) ∈ R N (11) where N is the number of classes in the dataset.

8. The sEMG gesture recognition method based on the lightweight depth separable residual attention network according to claim 7, characterized in that: In the DSR block, the input first goes through a 3×3 DWConv, followed by a PConv, as follows: Among them, X in is the input of the DSR block, C in is the number of input channels, C out is the number of output channels, F in and T in are the number of input frequency bins and the number of time frames respectively; Next, batch normalization and ReLU activation are applied to the output to introduce non-linearity: The combination of 3×3 DWConv and PConv is applied again to the output of the ReLU activation, followed by batch normalization: Residual connections are incorporated into the DSR block to combine the improved features with the original input. The design of the residual connections helps to directly pass gradients through the network and prevent the vanishing gradient problem. The residual connections are expressed as: Since the input and output channels are different, a 3×3 DWConv and PConv are first applied to the input to match the dimensions. After batch normalization of the result, it is added to the output of the convolutional layer. Finally, ReLU activation is applied to the combined output.

9. The sEMG gesture recognition method based on the lightweight depthwise separable residual attention network according to claim 8, characterized in that: The MACA module includes spatial attention and channel attention; The spatial attention is calculated by independently applying depth convolution and pointwise convolution along the frequency axis and the time axis, thereby obtaining frequency attention and time attention. The frequency attention is calculated by applying a depthwise separable convolution with a kernel size of (7,1) along the frequency axis, while the time attention is calculated by applying a convolution with a kernel size of (1,7) along the time axis; The spatial attention mechanism can be expressed as: where, σ represents the sigmoid activation function, denotes element-wise multiplication, C in 、F in and T in represent the number of channels, the number of frequency bins, and the number of time frames, respectively; The channel attention mechanism generates a channel-level feature descriptor through a global average pooling operation along the spatial dimension. Then, the feature descriptor generates channel-level attention weights through a fully connected layer with a sigmoid activation function. The channel attention mechanism is expressed as: where FC1 and FC2 are fully connected layers, and σ represents the sigmoid activation function; The final output of the MACA module is obtained through the element-wise multiplication combination of the spatial attention and channel attention mechanisms: where represents element-wise multiplication.

Citation Information

Cited By

  • Method for monitoring hitching state of grounding wire equipment based on multi-source signal fusion

    CN120671057A

  • A ground wire equipment hanging state monitoring method based on multi-source signal fusion

    CN120671057B

  • Acoustic emission source positioning method based on lightweight convolution and attention mechanism

    CN120801528A

  • Gesture awakening method based on micro-Doppler spectrogram and lightweight convolutional neural network

    CN121934725A