Generative adversarial network-based surface electromyogram signal channel expansion method and gesture classification model thereof
By generating high-quality multi-channel sEMG signals from the adversarial network, the hardware complexity and data scarcity of the multi-channel acquisition system are solved, and the performance of the gesture recognition model and device portability are improved.
Patent Information
- Application Number
- CN202510416697.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
AI Technical Summary
The existing multi-channel surface electromyography signal acquisition system has complex hardware and high cost, and the data quality and consistency are difficult to guarantee. The traditional data expansion method has limited effect. GANs face channel mapping and signal authenticity challenges in sEMG signal generation.
Adversarial training of Generative Adversarial Network (GAN) generator and discriminator is adopted to generate high-quality multi-channel signals through Wasserstein distance optimization, and combined with wavelet decomposition and convolutional neural network to process sEMG signals to generate additional channel signals consistent with the real signal.
It reduces hardware requirements, improves data set diversity and classification performance, and generates signals close to real signals in time and frequency domain characteristics, improving the accuracy of gesture recognition models and device portability.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical signal processing, and particularly to signal processing of surface electromyography (sEMG) signals. Background Art
[0002] Surface electromyography (sEMG) signals are non-invasive biological signals recorded through the electrical activity changes generated during muscle contraction, and are widely used in fields such as gesture recognition, human-computer interaction, and rehabilitation training. sEMG signals contain rich muscle activity information, and traditional gesture recognition systems usually rely on multi-channel acquisition to improve recognition accuracy.
[0003] However, the current multi-channel sEMG signal acquisition has the following main problems:
[0004] First of all, the multi-channel acquisition system requires complex hardware layout and sensor arrays. A large number of sensors will increase the cost of the device, the complexity of the hardware layout, and the discomfort of users wearing. For example, commercial sEMG systems often use 8 channels or even more electrodes to cover the target muscle area, which poses higher requirements for the design of portable devices.
[0005] Secondly, the collected multi-channel sEMG data is limited by factors such as the distribution of sensors, individual physiological differences, and environmental noise, and it is difficult to guarantee the data quality and consistency. The existing sEMG databases are limited, and the annotation process is time-consuming and laborious, making it difficult to meet the needs of deep learning models for a large amount of high-quality data.
[0006] Traditional data augmentation methods, such as interpolation or random noise addition, although expand the dataset scale to a certain extent, fail to fully capture the complex dynamic characteristics of sEMG signals. The signals generated by interpolation often lack biological significance, and simple noise addition easily destroys the essential characteristics of the signals, resulting in limited effects of the augmented data in classification tasks.
[0007] In recent years, generative adversarial networks (GANs), as a powerful data generation tool, have achieved remarkable results in fields such as images and speech. However, applying GANs to sEMG signal generation faces the following challenges: (1) how to design a generator to efficiently learn the dynamic mapping relationship between different channels; (2) how to constrain the authenticity and diversity of the generated signals through a discriminator; (3) how to use the generated multi-channel signals to optimize the classification performance of the gesture recognition model.
[0008] The technical problem to be solved by this application is: how to use GAN to augment sEMG signal channels. Summary of the Invention
[0009] To overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for expanding surface electromyogram signal channels based on a generative adversarial network, aiming to generate multi-channel signals with less-channel data as input, reduce hardware requirements, and improve classification performance at the same time.
[0010] The technical solution adopted by the present invention is as follows: A method for expanding surface electromyogram signal channels based on a generative adversarial network, comprising the following steps:
[0011] S1. Perform data preprocessing on the real signals of the original less-channel sEMG. The data preprocessing divides the real signals into standard signals with a specific length and a specific step size;
[0012] S2. Input the standard signals into the generator. The generator generates target channel signals, and the target channel signals are highly consistent with the standard signals in terms of length and distribution characteristics;
[0013] S3. Input the target channel signals and the standard signals into the discriminator at the same time. The discriminator calculates the Wasserstein distance;
[0014] S4. According to the feedback of the discriminator, the generator gradually improves the quality of the generated target channel signals by optimizing the loss function based on the Wasserstein distance;
[0015] S5. Alternately optimize the generator and the discriminator until the target channel signals are highly similar to the standard signals in terms of distribution;
[0016] S6. Combine the generated target channel signals and the standard signals to form an expanded multi-channel data set.
[0017] In some embodiments, the step of performing data preprocessing on the real signals of the original less-channel sEMG includes:
[0018] S11. Use the wavelet decomposition method to remove high-frequency noise and eliminate power frequency interference of a specific frequency through a notch filter;
[0019] S12. Perform min-max normalization on the amplitudes of the real signals and map them to the range of [-1, 1];
[0020] S13. Divide the real signals into standard signals of sampling points with a specific length and a specific step size through the overlapping sliding window segmentation method.
[0021] In some embodiments, the step of the generator generating target channel signals includes:
[0022] S21. The input layer receives standard signals with a specific length and a specific step size from less channels;
[0023] S22. Extract the temporal features of the signal through multi-layer one-dimensional convolution and gradually generate a high-resolution target channel signal in the output layer in combination with the sampling operation.
[0024] In some embodiments, regarding step S22, the signal output from the output layer is processed by the activation functions LeakyReLU and tanh.
[0025] In some embodiments, the discriminator includes a multi-dimensional feature extraction pipeline, a mini-batch discrimination module, and a feature fusion module.
[0026] In some embodiments, the multi-dimensional feature extraction pipelines are respectively used to extract time-domain features, Fourier transform frequency-domain features, envelope features, and multi-resolution features of wavelet transform.
[0027] In some embodiments, the mini-batch discrimination module calculates the similarity between the multi-instance features of the target channel signal, and the feature fusion module splices all the features and inputs them into a fully connected layer for classification and outputs the probability P that the signal is discriminated as a standard signal through the sigmoid activation function.
[0028] A gesture classification model, applying the multi-channel data set obtained by the above method, the multi-channel data set includes a training set, and the classification process of the classification model includes the following steps:
[0029] D1. Temporal convolution: The frequency-domain convolution extracts local features of the signals in the training set.
[0030] D2. Frequency-domain convolution: Perform Fourier transform on the training set and obtain its spectral features, and perform convolution operation on the spectrum through one-dimensional convolution to extract global frequency-domain features.
[0031] D3. Feature fusion: Perform inverse Fourier transform on the global frequency-domain features extracted in the frequency-domain convolution step, convert the global frequency-domain features back to global time series features, and then fuse the global time series features with the local features obtained in the temporal convolution step to generate comprehensive temporal features.
[0032] D4. The temporal features are input into a fully connected layer for classification operations through depthwise separable convolution, average pooling operation, and flattening.
[0033] In some embodiments, the classification operation uses the softmax activation function.
[0034] In some embodiments, the multi-channel data set includes a test set, and the test set is used to evaluate the performance of the classification model.
[0035] The beneficial effects of the present invention are as follows:
[0036] The method for expanding surface electromyogram signal channels based on a generative adversarial network generates additional channel signals through a generative adversarial network (GAN), solves problems such as high hardware complexity and scarce data in multi-channel acquisition systems, reduces the number of sensors, thereby reducing hardware dependence, and then designs a more concise and portable acquisition device, improving the usability and practical application scope of the device. At the same time, using the high-quality data expansion scheme generated by GAN ensures that the generated signals are close to real signals in the time domain, frequency domain, and statistical characteristics, thereby effectively enhancing the diversity of the training data set of the model. In the gesture classification model, the generated signals exhibit good classification performance and ensure the biological consistency of the signals, providing a reliable solution for various intelligent devices and medical applications. Brief Description of the Drawings
[0037] None Detailed Embodiments
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] The present invention provides a technical solution: a method for expanding surface electromyogram signal channels based on a generative adversarial network, including the following steps, where the generative adversarial network includes a generator and a discriminator:
[0040] S1. Perform data preprocessing on the real signals of the original few-channel sEMG. The data preprocessing divides the real signals into standard signals with a specific length and a specific step size;
[0041] S2. Input the standard signals into the generator, and the generator generates target channel signals that are highly consistent with the standard signals in terms of length and distribution characteristics;
[0042] S3. Input the target channel signals and the standard signals into the discriminator at the same time, and the discriminator calculates the Wasserstein distance through the feature extraction and classification module;
[0043] S4. According to the feedback of the discriminator, the generator gradually improves the quality of the generated target channel signals by optimizing the loss function based on the Wasserstein distance;
[0044] S5. Alternately optimize the generator and the discriminator until the target channel signals are highly similar to the standard signals in distribution;
[0045] S6. Combine the generated target channel signals and the standard signals to form an expanded multi-channel data set.
[0046] Generate high-quality multi-channel sEMG signals through the adversarial training of a generator and a discriminator. The generator uses few-channel sEMG signals to learn the complex mapping relationship between different channels to generate target channel signals.
[0047] Perform data preprocessing on the real signals of the original few-channel sEMG. The data preprocessing includes the following steps:
[0048] S11. Use the wavelet decomposition method to remove high-frequency noise, and eliminate the power frequency interference of specific frequencies (such as 50 Hz) through a notch filter to ensure the purity of the signal.
[0049] S12. Perform min-max normalization on the amplitude of the real signal, map it to the range of [-1, 1], eliminate the amplitude difference between different channels, and improve the stability of subsequent adversarial training.
[0050] S13. Divide the real signal into standard signals of sampling points with specific lengths and specific step sizes through the overlapping sliding window segmentation method to ensure the consistency and representativeness of the signals input to the generator.
[0051] The design of the generator is based on a convolutional neural network (CNN). The steps for the generator to generate are as follows: S21. The input layer receives standard signals of a specific length from few channels; S22. Extract the temporal features of the signal through multi-layer one-dimensional convolution, and gradually generate high-resolution target channel signals in the output layer in combination with the sampling operation; among them, the signal output in the output layer is processed by an activation function. The activation functions selected are LeakyReLU and tanh. LeakyReLU can enhance the non-linear expression ability of the training model, and the tanh activation function ensures that the amplitude of the generated signal is consistent with the amplitude range of the standard signal; the finally generated target channel signal is highly consistent with the standard signal in terms of length and distribution characteristics.
[0052] The goal of the discriminator is to distinguish between target channel signals and standard signals. The discriminator specifically includes the following structures: 1. Four parallel feature extraction pipelines, which are respectively used to extract time-domain features (such as signal mean and peak value), Fourier transform frequency-domain features (such as main frequency components), envelope features, and multi-resolution features of wavelet transform; 2. A mini-batch discrimination module, which ensures the sufficient diversity of target channel signals by calculating the similarity between multi-instance features of target channel signals; 3. A feature fusion module, which splices all features and inputs them into a fully connected layer for classification, and outputs the probability P that the signal is judged as a standard signal through a sigmoid activation function.
[0053] The training process of the generator and the discriminator is based on the Wasserstein GAN (WGAN) framework and combined with the gradient penalty term (GP) to improve the stability of adversarial training. The target channel signal generated by the generator and the real standard signal are used as the input of the discriminator. The discriminator evaluates the authenticity of the generated signal through a multi-dimensional feature extraction pipeline and a small batch identification module; then, the discriminator calculates the Wasserstein distance of the generated target channel signal and feeds the result back to the generator; finally, the generator is optimized using a loss function based on the Wasserstein distance, and the quality of the generated target channel signal is gradually improved through adversarial training, making its distribution closer to the standard signal. The target channel signal is used to supplement the problem of insufficient few-channel data.
[0054] The generated multi-channel signal is further used for gesture recognition classification tasks. The experimental design uses multi-channel data expanded by GAN for training and testing to verify the effect of the target channel signal on the improvement of classification performance. Specifically, the target channel signal is combined with the real standard signal to construct a multi-channel data set. The training set in the multi-channel data set is used to optimize the classification model, and the test set is used to evaluate the performance of the classification model. The performance of the classification model is evaluated by the accuracy of the test set.
[0055] In the gesture recognition classification task, the classification process of the classification model specifically includes the following steps:
[0056] D1. Time domain convolution: The time domain convolution part directly extracts local features of the input sEMG signal through one-dimensional convolution. This part can capture the short-term dynamic changes of the signal and thus extract local characteristics in the time domain.
[0057] D2, frequency domain convolution: Perform Fourier transform on the sEMG signal to obtain its spectrum characteristics. Then, perform convolution operation on the spectrum through one-dimensional convolution to extract global frequency domain characteristics. This part of the design aims to mine the correlation of the signal in the frequency domain and obtain the global characteristics in the frequency domain.
[0058] D3, feature fusion: Perform inverse Fourier transform on the global frequency domain features extracted by the frequency domain convolution part, and convert the global frequency domain features back to global time series features. Then, the global time series features are fused with the local features obtained by the time domain convolution part to generate comprehensive time domain features.
[0059] D4. The fused time domain features are further enhanced in representation capability through deep separable convolution, and then are input into the fully connected layer for classification after average pooling and straightening.
[0060] Among them, the classification operation uses the softmax activation function to provide the probability distribution of each gesture category and complete the prediction of gesture categories. During the training process of the classification model, the sEMG signal is switched to the signal of the training set for model training of the classification model.
[0061] Through this design of the classification model that combines time-domain and frequency-domain features, the present invention effectively improves the feature expression ability of sEMG signals and shows significant performance improvement in the gesture recognition task after multi-channel data augmentation.
[0062] Experiments were carried out on the Ninapro DB2 public dataset. The experimental results show that when using the multi-channel sEMG signals generated by the present invention for gesture recognition, compared with the model trained only with real signals, the classification accuracy is improved by about 10%. Specifically, the generated signals are highly similar to the real signals in terms of time-domain and frequency-domain characteristics, and their correlation exceeds 90%. In the experiment, the training set and test set were constructed with the multi-channel signals augmented by GAN. The classification model trained with the augmented dataset performs significantly better than the model with the non-augmented dataset in various gesture recognition tasks. The model using the augmented dataset is more stable in various gesture recognition tasks, verifying the effectiveness and practicality of the target channel signals. Through this method, not only the hardware channel requirements of the acquisition device are reduced, the hardware complexity is lowered, but also an efficient signal augmentation scheme is provided, which is applicable to the application of portable sEMG devices and deep learning models in large-scale data scenarios.
[0063] Finally, it should be noted that the above are only preferred examples of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for expanding surface electromyogram signal channels based on a generative adversarial network, characterized in that, It includes the following steps: S1. Perform data preprocessing on the real signal of the original few-channel sEMG. The data preprocessing divides the real signal into standard signals with a specific length and a specific step size. S2. Input the standard signal into a generator. The generator generates a target channel signal, and the target channel signal is highly consistent with the standard signal in terms of length and distribution characteristics. S3. Input the target channel signal and the standard signal into a discriminator at the same time. The discriminator calculates the Wasserstein distance. S4. According to the feedback of the discriminator, the generator gradually improves the quality of the generated target channel signal by optimizing the loss function based on the Wasserstein distance. S5. The generator and the discriminator are alternately optimized until the target channel signal is highly close to the standard signal in distribution. S6. The generated target channel signal and the standard signal are combined to form an expanded multi-channel data set.
2. The method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 1, wherein The steps for performing data preprocessing on the real signal of the original few-channel sEMG include: S11. Use the wavelet decomposition method to remove high-frequency noise and eliminate power frequency interference of a specific frequency through a notch filter. S12. Perform min-max normalization on the amplitude of the real signal and map it to the range of [-1, 1]. S13. Divide the real signal into standard signals of sampling points with a specific length and a specific step size through the overlapping sliding window segmentation method.
3. A method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 1, wherein The steps for the generator to generate the target channel signal include: S21. The input layer receives the standard signal with a specific length and a specific step size from the few channels. S22. Extract the temporal features of the signal through multi-layer one-dimensional convolution and gradually generate a high-resolution target channel signal in the output layer in combination with the sampling operation.
4. A method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 3, characterized in that, Regarding step S22, the signal output in the output layer is processed by the activation functions LeakyReLU and tanh.
5. A method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 1, characterized in that The discriminator includes a multi-dimensional feature extraction pipeline, a mini-batch discrimination module, and a feature fusion module.
6. The method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 5, wherein The multi-dimensional feature extraction pipelines are respectively used to extract time-domain features, Fourier transform frequency-domain features, envelope features, and multi-resolution features of wavelet transform.
7. A method for expanding surface electromyogram signal channels based on a generative adversarial network according to claim 5, characterized in that The mini-batch discrimination module calculates the similarity between the multi-instance features of the target channel signal. The feature fusion module splices all the features and inputs them into a fully connected layer for classification and outputs the probability P that the signal is discriminated as a standard signal through the sigmoid activation function.
8. A gesture classification model, which uses a multi-channel data set obtained by using the method for expanding surface electromyogram signal channels based on a generative adversarial network according to any one of claims 1-8, and is characterized in that, The multi-channel data set includes a training set. The classification process of the classification model includes the following steps: D1. Temporal convolution: The frequency-domain convolution extracts local features from the signals in the training set. D2. Frequency-domain convolution: Perform Fourier transform on the training set and obtain its spectral features. Perform convolution operation on this spectrum through one-dimensional convolution to extract global frequency-domain features. D3. Feature fusion: Perform inverse Fourier transform on the global frequency-domain features extracted in the frequency-domain convolution step, convert the global frequency-domain features back to global time-series features, and then fuse the global time-series features with the local features obtained in the temporal convolution step to generate comprehensive temporal features. D4. The time-domain features are input into a fully connected layer for classification operations through depthwise separable convolution, average pooling operations, and flattening processing.
9. A gesture classification model, characterized in that, The classification operation uses a softmax activation function.
10. A gesture classification model, characterized in that, The multi-channel dataset includes a test set, which is used to evaluate the performance of the classification model.