A method, system, electronic device and storage medium for distinguishing gestures independent of myoelectricity

By using the feature reconstruction network EMG-FRNet based on the ganomaly network in electromyographic gesture recognition, combining channel cropping, cross-layer codec feature fusion and SE channel attention mechanism, the problem of low accuracy in electromyographic gesture recognition is solved, and efficient irrelevant action discrimination is achieved.

CN116682146BActive Publication Date: 2025-05-06BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310780925.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-05-06
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

The prior art has problems with low accuracy and poor interaction effect in the recognition of EMG gestures, especially in the difficulty of effectively identifying irrelevant movement interference.

Method used

The feature reconstruction network EMG-FRNet based on the ganomaly network architecture is adopted. By adding channel cropping, cross-layer codec feature fusion and SE channel attention mechanisms, the model's ability to reconstruct the target samples and improve the performance of irrelevant action discrimination.

Benefits of technology

The accuracy and stability of EMG gesture recognition is improved, and the recognition ability of target samples and irrelevant samples is significantly improved. The experimental results show that the AUC value on multiple data sets has reached a high level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116682146B_ABST
    Figure CN116682146B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for distinguishing electromyography-independent gestures, which is implemented based on a feature reconstruction network EMG-FRNet, including: obtaining a data set, dividing the data, and forming an electromyography sample after data preprocessing; establishing a feature reconstruction network EMG-FRNet, based on adding a channel clipping strategy, a cross-layer encoding and decoding feature fusion strategy, and an SE channel attention strategy on the basis of the original ganomaly network architecture, the original ganomaly network architecture includes a generator and a discriminator; inputting the electromyography sample into the feature reconstruction network EMG-FRNet, extracting the potential feature z of the input electromyography sample, and reconstructing the potential feature z of the electromyography sample to obtain the reconstructed potential feature and then outputting it; calculating the feature reconstruction error error between z and , and comparing it with a predefined threshold threshold to determine whether the electromyography sample belongs to a target gesture or an irrelevant gesture. The present invention also discloses a corresponding system, an electronic device, and a computer-readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of signal processing and gesture recognition, and in particular to a method, system, electronic device and storage medium for recognizing myoelectric gestures that are independent of gestures. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technology, human-computer interaction through gesture recognition technology has gradually become a research hotspot. Surface electromyography is bioelectric information obtained from the surface of the skin, which has the advantages of non-invasiveness, non-traumaticity, and simple operation. Surface electromyography directly reflects the muscle contraction state that causes limb movement, contains rich movement information, and can predict the intention of hand movement. Compared with visual gesture recognition, gesture recognition based on surface electromyography shows advantages such as being unaffected by changes in the external environment background, small amount of calculation, and high real-time performance. Therefore, with the development of deep learning technology, gesture recognition technology based on surface electromyography has shown broad application prospects in human-computer interaction fields such as intelligent prostheses, rehabilitation exoskeletons, rehabilitation therapy, and sign language translation. In the field of electromyography pattern recognition, most of the research focuses on improving the accuracy of electromyography gesture recognition. Therefore, most gesture recognition technologies can achieve high recognition accuracy in a variety of gesture actions. The main technical means include:

[0003] (1) By selectively extracting various time-domain and frequency-domain features of EMG signals, such as mean absolute value (MAV), root mean square (RMS), median frequency (MDF), power spectrum (PS), etc., and combining deep learning models such as recurrent neural network (RNN), convolutional neural network (CNN), and Transformer to capture local or global muscle activity information of EMG signals, a high accuracy rate can be achieved in a variety of gestures.

[0004] (2) The LST-EMG-Net model was proposed in a laboratory environment, which further improved the accuracy of electromyographic gesture recognition.

[0005] However, compared with the ideal laboratory environment, in the actual EMG interaction application process, there are often many differences or interferences, such as irrelevant movements, electrode offset, muscle fatigue, user differences, etc., which ultimately lead to low EMG recognition accuracy and poor interaction effect. Among them, irrelevant movement interference is a type of interference that is very easy to occur. Irrelevant movement interference means that during use, the subject will inadvertently make movements that do not belong to the predefined target category. At this time, the classifier is forced to select one of the trained movements, causing the system to produce incorrect recognition results, which damages the safety of the device and the user. Therefore, it is very important to design a method for distinguishing irrelevant movements.

[0006] The existing technology for irrelevant action discrimination is mainly divided into probability-based methods and one-to-many classification rule-based methods.

[0007] (1) Probability-based methods: The core idea of ​​this type of method is to effectively distinguish target samples from irrelevant samples by comparing the relationship between the predicted probability value of the classifier for the test sample and the preset probability threshold. Specifically, the classifier calculates a predicted probability value for the test sample. If the predicted probability value is higher than the preset threshold, it is classified as a target sample, otherwise it is classified as an irrelevant sample. The probability-based method is simple in principle and has low implementation cost, but the classification probability of many target samples may be low, while the classification probability of irrelevant samples is high, which ultimately affects the accuracy of irrelevant action discrimination.

[0008] (2) Methods based on one-to-many classification rules: The core idea of ​​this type of method is to effectively distinguish target samples from irrelevant samples by training a one-class classifier for each target category. Specifically, the test sample is input into all classifiers to obtain the binary classification result of each classifier to determine whether the test sample belongs to the target category corresponding to the classifier. If the test sample does not belong to any known target category, it is classified as an irrelevant sample, otherwise it is classified as a target sample. In the methods based on one-to-many classification rules, some simple machine learning methods such as GC and SVDD are used to distinguish irrelevant gestures, but they assume that there are obvious differences between target gestures and irrelevant gestures in the feature space, while gestures in practical applications are difficult to predict.

[0009] The detection of irrelevant actions is very similar to the problem solved by anomaly detection (also known as outlier detection or novelty detection, which refers to the process of detecting those data instances that are significantly deviated from the majority). Autoencoder AE is a relatively classic method in anomaly detection. By calculating the reconstruction error, the AE method can more fully explore the subtle differences between the target gesture and irrelevant gestures, thereby improving the discrimination performance and stability of the model. However, methods based on AE and AE variants are usually susceptible to data noise presented in the training data, resulting in severe overfitting and abnormal small reconstruction errors, as well as limited reconstruction performance of the model, and can only be applied in small quantities in high-density electromyography systems.

[0010] With the development of ganomaly networks, ganomaly-based anomaly detection has quickly become a popular deep anomaly detection method. Ganomaly has shown excellent ability in generating realistic instances, thus being able to detect anomaly instances that are poorly reconstructed from the latent space. However, the detection performance of ganomaly networks in the field of myoelectric-independent gesture recognition needs to be further explored and improved. Summary of the invention

[0011] In order to solve the problems existing in the prior art, the present invention provides a method and system for distinguishing electromyography-independent gestures. Based on the ganomaly network architecture, the present invention adds channel pruning, cross-layer encoding and decoding feature fusion, and SE channel attention mechanism corresponding structures to further improve the model's ability to reconstruct target samples and enhance the performance of distinguishing irrelevant actions.

[0012] On the one hand, the present invention proposes a method for distinguishing gestures independent of electromyography, which is implemented based on a feature reconstruction network EMG-FRNet, comprising:

[0013] S1, obtaining a data set, dividing the data set and performing data preprocessing on the data set to form an electromyographic sample;

[0014] S2, establishing a feature reconstruction network EMG-FRNet, wherein the feature reconstruction network EMG-FRNet is implemented by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and an SE channel attention strategy to the original ganomaly network architecture, wherein the original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encode-decode-encode" structure composed of a first encoder, a decoder, and a second encoder;

[0015] S3, inputting the EMG sample into the feature reconstruction network EMG-FRNet, extracting the potential features of the input EMG sample, reconstructing the potential features of the EMG sample to obtain the reconstructed potential features, and then outputting the potential features of the EMG sample and the reconstructed potential features;

[0016] S4, calculate the latent feature z and reconstruct the latent feature and comparing the feature reconstruction error with a predefined threshold, and determining whether the electromyographic sample belongs to a target gesture or an irrelevant gesture based on a result of the comparison.

[0017] Preferably, the data set includes an existing and / or self-collected electromyographic data set containing multiple gesture movements; the data division includes a first division and a second division of the data set; the first division is to divide the electromyographic signal data into a target class data set and an irrelevant class data set, and the second division is to divide the electromyographic signal data into a training set and a test set, wherein the training set only has data related to the target class movement, and the test set has both data related to the target class movement and data related to the irrelevant class movement; the data preprocessing includes data segmentation and dimensionality transformation; wherein the data segmentation is based on a sliding window method to segment multi-channel electromyographic signals to obtain electromyographic samples of gesture movements; the dimensionality transformation includes transforming the first dimension and the second dimension of a two-dimensional matrix formed by the multi-channel electromyographic samples after the data segmentation; the transformed two-dimensional matrix is ​​more adaptable to the input requirements of the feature reconstruction network EMG-FRNet than the two-dimensional matrix before the transformation.

[0018] Preferably, S2 includes:

[0019] S21, build ganomaly network infrastructure;

[0020] S22, based on the original ganomaly network infrastructure, adds channel pruning strategy, cross-layer encoding and decoding feature fusion strategy and SE channel attention strategy to form the initial feature reconstruction network EMG-FRNet;

[0021] S23, using the target class data set to train the initial feature reconstruction network EMG-FRNet, using all categories of data sets as the test set to test the initial feature reconstruction network EMG-FRNet, and continuously repeating the training and the testing until the test effect reaches the optimal, thereby forming the feature reconstruction network EMG-FRNet.

[0022] Preferably, the channel clipping strategy is used to implement channel clipping, including proportional channel clipping and power channel clipping, wherein the proportional channel clipping includes clipping the number of channels of the network feature layer according to the channel ratio of the network input sample, and the power channel clipping includes changing the number of channels on the basis of proportional channel clipping, so that the number of channels after power channel clipping meets the requirement of an integer power of 2;

[0023] The cross-layer codec feature fusion strategy is used to implement cross-layer codec feature fusion, adding multiple jump connections between the first encoder and decoder of the generator, establishing a direct path between the upsampling layer and the downsampling layer, and fusing the downsampling feature map and the upsampling feature map through the multiple jump connections; the feature fusion obtains the feature splicing map of the codec by performing feature splicing in the channel dimension to perform the feature fusion on the feature information in the encoding and decoding processes; wherein the multiple jump connections include performing multiple feature fusions on the feature maps of different scales of the first encoder and the decoder;

[0024] The SE channel attention strategy is used to implement the SE channel attention method, so as to obtain a feature splicing map containing weight information from the feature splicing map of the codec based on the SE channel attention mechanism; the SE channel attention method includes a compression operation, an excitation operation and a scaling operation; wherein the compression operation is used to obtain the feature representation z of each channel of the feature splicing map of the codec; the excitation operation is used to obtain the weight vector of each channel, and the scaling operation is used to obtain the feature splicing map containing weight information.

[0025] Preferably, the compression operation includes: reducing the dimension of the feature splicing map M of the codec by global average pooling so that each feature channel has a numerical representation, thereby obtaining the feature representation of each channel of the feature splicing map of the codec, as shown in formula (1):

[0026]

[0027] Wherein, z represents the feature representation of each channel of the feature splicing map M of the codec, M represents the feature splicing map M of the codec, H and W represent the length and width of the feature splicing map M of the codec, respectively, and i, j represent the channel numbers in the length and width directions of the feature splicing map M of the codec, respectively;

[0028] The excitation operation includes: performing nonlinear transformation on the feature representation of each channel of the feature concatenation graph of the codec, and mapping the result into a weight vector, where different values ​​in the weight vector represent weight information of different channels, as shown in formula (2):

[0029] s=Excitation(z)=sigmoid(W2 Relu(W1z)) (2);

[0030] Where s represents the weight vector, W1 represents the first fully connected layer parameter, Relu is the activation function of the first fully connected layer, W2 represents the second fully connected layer parameter, and sigmoid represents the activation function of the second fully connected layer;

[0031] The scaling operation includes: assigning weights to the feature splicing graph using a weight vector through a multiplication operation to obtain a feature splicing graph containing weight information, as shown in formula (3).

[0032] M-weight=Scale(M,s)=M×s (3);

[0033] Among them, M-weight is a feature concatenation graph containing weight information.

[0034] Preferably, the use of the target class data set to train the initial feature reconstruction network EMG-FRNet includes: cross-training the generator and the discriminator based on the loss of the model structure of the feature reconstruction network EMG-FRNet for electromyography-independent gesture recognition, the loss including the loss of the generator and the loss of the discriminator; the loss of the generator includes the electromyography sample reconstruction loss, the electromyography sample potential feature reconstruction loss and the adversarial loss, the electromyography sample reconstruction loss represents the difference between the input electromyography sample and the electromyography sample restored by the generator, the electromyography sample potential feature reconstruction loss represents the difference between the electromyography sample potential feature and the electromyography sample reconstructed potential feature, the adversarial loss represents the difference between the true and false electromyography samples in the intermediate layer features of the discriminator, the generator loss is the weighted sum of the electromyography sample reconstruction loss, the electromyography sample potential feature reconstruction loss and the adversarial loss; the loss of the discriminator is obtained based on the binary cross entropy function.

[0035] Preferably, the feature reconstruction error in S4 is the difference between the potential feature and the reconstructed potential feature; determining that the electromyographic sample belongs to the target gesture or the irrelevant gesture based on the result of the comparison includes: when the feature reconstruction error is less than a predefined threshold, determining that the electromyographic sample belongs to the target gesture; when the feature reconstruction error is greater than a predefined threshold, determining that the electromyographic sample belongs to an irrelevant gesture.

[0036] The second aspect of the present invention is to provide an electromyography-independent gesture recognition system, comprising: a data processing module, used to obtain a data set, and to form an electromyography sample after data division and data preprocessing of the data set; wherein the electromyography sample provides a data basis for training and testing a network model; a feature reconstruction network establishment module, used to establish a feature reconstruction network EMG-FRNet, wherein the feature reconstruction network EMG-FRNet is established based on the original ganomaly network architecture by adding a channel clipping strategy, a cross-layer encoding and decoding feature fusion strategy and an SE channel attention strategy, wherein the original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encoder-decoder-encoder" structure composed of a first encoder, a decoder and a second encoder; a feature reconstruction module, used to input the electromyography sample into the feature reconstruction network EMG-FRNet, extract the potential feature z of the input electromyography sample, and reconstruct the potential feature of the electromyography sample to obtain the reconstructed potential feature Output the latent features z of the EMG sample and reconstruct the latent features Unrelated gesture discrimination module, used to calculate the latent feature z and reconstruct the latent feature and compare the feature reconstruction error error with a predefined threshold threshold, and determine whether the electromyographic sample belongs to the target gesture or an irrelevant gesture based on the result of the comparison; wherein the feature reconstruction error error is a potential feature z and a reconstructed potential feature The difference between.

[0037] The electromyography-independent gesture recognition method, system, electronic device and computer-readable storage medium provided by the present invention have the following beneficial technical effects: linking the electromyography-independent gesture recognition field with the anomaly detection field, proposing a feature reconstruction network EMG-FRNet for electromyography-independent gesture recognition, applying ganomaly to electromyography-independent gesture recognition for the first time, and adding channel clipping, cross-layer encoding and decoding feature fusion and SE channel attention strategy on its basis, so that the network has a small feature reconstruction error for target category samples, and a large feature reconstruction error for irrelevant category samples, thereby improving the recognition ability of target samples and irrelevant samples. We verified the feasibility of the proposed method through experiments, and the experimental results show that the performance of irrelevant gesture recognition of the method proposed by the present invention on all electromyography data sets can be maintained at a high level. Specifically, the AUC values ​​of the method of the present invention on DB1, DB5 and self-collected data sets reached 0.940, 0.926 and 0.962 respectively, which are better than 0.882, 0.819 and 0.902 of AE and 0.744, 0.753 and 0.723 of SVDD. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of the method for distinguishing hand gestures independent of myoelectricity according to the present invention;

[0039] Figure 2 The gesture types of the data set used in the present invention are: Figure 2 (a) 8 gestures of the public dataset DB1 / public dataset DB5 Exercise B dataset; Figure 2 (b) 7 gestures from the self-collected dataset;

[0040] Figure 3 This is a schematic diagram of data segmentation according to the present invention;

[0041] Figure 4 (a) is a schematic diagram of the original ganomaly model structure; Figure 4 (b) is a schematic diagram of the EMG-FRNet model structure of the present invention;

[0042] Figure 5 Schematic diagram of the principle of the channel clipping method of the downsampling layer feature map of the present invention, which includes proportional channel clipping and power channel clipping, and the bold numbers represent the number of channels of the feature map;

[0043] Figure 6 Schematic diagram of cross-layer encoding and decoding feature fusion according to the present invention; wherein the down-sampled feature map and the up-sampled feature map are feature fused through a skip connection;

[0044] Figure 7 Schematic diagram of SE channel attention of the present invention; wherein the Squeeze operation obtains the feature representation z of each channel of the feature splicing graph M; the Excitation operation obtains the weight s of each channel; the Scale operation obtains the feature splicing graph M-weight containing weight information;

[0045] Figure 8 A schematic diagram of the confusion matrix of the present invention;

[0046] Fig. 9 Schematic diagram of the ROC curve of the present invention; wherein the ROC curve is the dark curve in the figure, the horizontal axis is the false positive rate FPR, and the vertical axis is the true positive rate TPR;

[0047] Fig.10 It is the AUC line graph when the first subject of each data set sets different target gestures in the comparative experiment described in the present invention; Fig.10 (a)- Fig.10 (c) AUC line graphs of the first subject of the three algorithms, SVDD, AE, and EMG-FRNet, on DB1, DB5, and self-collected datasets;

[0048] Fig.11 It is a line graph of AUC of different subjects in each data set in the comparative experiment described in the present invention; wherein Fig.11 (a)- Fig.11 (c) AUC line graphs of SVDD, AE, and EMG-FRNet algorithms on different subjects on DB1, DB5, and self-collected datasets;

[0049] Fig.12 It is the AUC line graph when the first subject of each data set sets different target gestures in the ablation experiment described in the present invention; Fig.12 (a)- Fig.12 (c) AUC line graphs of the first subject of the four algorithms, ganomaly, ganomaly+CC, ganomaly+CC+CLEDF, and EMG-FRNet, on DB1, DB5, and self-collected datasets;

[0050] Fig.13 is a line graph of the AUC of different subjects in each data set in the ablation experiment described in the present invention; wherein Fig.13 (a)- Fig.13 (c) AUC line graphs of ganomaly, ganomaly+CC, ganomaly+CC+CLEDF, and EMG-FRNet algorithms on different subjects on DB1, DB5, and self-collected datasets;

[0051] Fig.14 This is a structural diagram of the myoelectricity-independent gesture recognition system of the present invention;

[0052] Fig.15 It is a schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION

[0053] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0054] Embodiment 1

[0055] like Figure 1 As shown, this embodiment provides a method for distinguishing gestures independent of electromyography, which is implemented based on a feature reconstruction network EMG-FRNet and includes:

[0056] S1, obtaining a data set, and performing data division and data preprocessing on the data set to form electromyographic samples; wherein the electromyographic samples provide a data basis for training and testing a network model;

[0057] As a preferred embodiment, the data set includes an existing and / or self-collected electromyographic data set containing multiple gestures. In this embodiment, the existing electromyographic data set includes: a public data set NinaproDB1 and a public data set NinaproDB5; the self-collected electromyographic data set is the self-collected data set.

[0058] 1) Public dataset NinaproDB1: The eight basic hand postures in Exercise B in the public dataset DB1 are used (such as Figure 2 (a)), each gesture is collected 10 times, with a total of 10 healthy subjects. The acquisition device is a 10-channel Otto Bock13E200 with a sampling frequency of 100Hz.

[0059] 2) Public dataset NinaproDB5: We use the eight basic hand gestures from Exercise B in the public dataset NinaproDB5 (e.g. Figure 2 (a) Each gesture was collected 6 times, with a total of 10 healthy subjects. The collection equipment is two Myo electromyography bracelets, each of which has 8 channels and a sampling frequency of 200Hz.

[0060] 3) Self-collected dataset: 7 hand gestures commonly used in life (such as Figure 2 (b) Each gesture was collected 6 times, with a total of 6 healthy subjects. The collection device is a Myo electromyography bracelet with a total of 8 channels and a sampling frequency of 200Hz.

[0061] As a preferred embodiment, the data division includes a first division and a second division of the data set; the first division is to divide the electromyographic signal data into a target class data set and an irrelevant class data set, and the second division is to divide the electromyographic signal data into a training set and a test set.

[0062] In this embodiment, the division of the target class data set and the irrelevant class data set is as follows: the experiment sets one gesture as the target class action and the other gestures as irrelevant class actions, and traverses all possible situations. Specifically, in the public data set NinaproDB1 and the public data set NinaproDB5, each subject conducts 8 experiments in total, and in the self-collected data set, each subject conducts 7 experiments in total.

[0063] Division of training set and test set: The training set only has data related to the target class action, and the test set has data related to both the target class action and the irrelevant class action. Specifically, in the public dataset NinaproDB1 dataset experiment, the 1st, 3rd, 4th, 6th, 8th, 9th, and 10th gesture repetition data of the target class are used to construct the training set, and the 2nd, 5th, and 7th gesture repetition data of the target class and all gesture repetition data of the irrelevant class are used to construct the test set. In the public dataset NinaproDB5 and self-collected dataset experiments, the 1st, 3rd, 4th, and 6th gesture repetition data of the target class are used to construct the training set, and the 2nd and 5th gesture repetition data of the target class and all gesture repetition data of the irrelevant class are used to construct the test set.

[0064] As a preferred embodiment, the data preprocessing includes data segmentation and dimensionality transformation; wherein the data segmentation is based on a sliding window method to segment multi-channel electromyographic signals to obtain electromyographic samples of gesture movements; the dimensionality transformation includes transforming the first dimension and the second dimension of a two-dimensional matrix formed by the multi-channel electromyographic samples after the data segmentation; the transformed two-dimensional matrix is ​​more adaptable to the input requirements of the feature reconstruction network EMG-FRNet than the two-dimensional matrix before the transformation.

[0065] In this embodiment, only 8 channels of myoelectric data are used for all data sets. The reason is that the most important muscle group for myoelectric gesture recognition is concentrated around the brachioradialis muscle of the forearm below the elbow. The eight-channel myoelectric bracelet provided by the prior art can cover this part of the muscle, and the configuration of this bracelet is convenient to carry and has a wide range of practical application prospects. Therefore, this data acquisition scheme is the preferred method for implementing the myoelectric gesture recognition method of the present invention.

[0066] (1) Data segmentation: Figure 3 As shown, the preferred embodiment of the present invention uses a sliding window method to segment multi-channel electromyographic signals to obtain electromyographic samples of gesture actions and use them for model training and testing. Specifically, the length of the sliding window is set to 128 sampling points ( Figure 3 In the figure, windows1 and windows2 represent EMG windows with a length of 128 sampling points, and a sliding step of 16 sampling points ( Figure 3 The stride represents the step length of 16 sampling points). This method can effectively segment the electromyographic signal, thereby extracting useful time domain features and providing a data basis for subsequent model training and testing.

[0067] (2) Data dimension transformation: In order to better learn feature representation in the basic GAN network (i.e., ganomaly network model) of the feature reconstruction network EMG-FRNet, the present invention transforms the data dimension. Specifically, the segmented multi-channel electromyographic samples are transformed from a two-dimensional matrix of 8×128 to a two-dimensional matrix of 32×32 to better meet the input requirements of the basic GAN network of the feature reconstruction network EMG-FRNet.

[0068] S2, establishing a feature reconstruction network EMG-FRNet, wherein the feature reconstruction network EMG-FRNet is implemented by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and an SE channel attention strategy to the original ganomaly network architecture, wherein the original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encode-decode-encode" structure composed of a first encoder, a decoder, and a second encoder;

[0069] As a preferred embodiment, the S2 includes:

[0070] S21, build ganomaly network infrastructure;

[0071] S22, based on the original ganomaly network infrastructure, adds channel pruning strategy, cross-layer encoding and decoding feature fusion strategy and SE channel attention strategy to form the initial feature reconstruction network EMG-FRNet;

[0072] S23, using the target class data set to train the initial feature reconstruction network EMG-FRNet, using all categories of data sets as the test set to test the initial feature reconstruction network EMG-FRNet, and continuously repeating the training and the testing until the test effect reaches the optimal, thereby forming the feature reconstruction network EMG-FRNet.

[0073] As a preferred implementation, the feature reconstruction network EMG-FRNet is implemented based on the original ganomaly network architecture by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and a SE channel attention strategy to improve the reconstruction performance of the network and achieve high-reliability recognition of irrelevant actions.

[0074] In this embodiment, the original ganomaly network architecture is as follows Figure 4 As shown in (a), it mainly consists of two parts, namely the generator G and the discriminator D. Among them, the generator network G consists of the encoder G E1 , decoder G D , encoder G E2 The structure of "encoding-decoding-encoding" is formed. First, the encoder G E1Downsampling learns the latent features z of the input data X, and then passes through the decoder G D Upsample the potential feature z to obtain the reconstructed input data Then, through the encoder G E2 Downsampling learning to reconstruct input data The feature representation That is, the reconstructed potential features of the input data X. Encoder G E2 Adopt and encoder G E1 The same network structure. The discriminator network D is used to distinguish the input data X and reconstruct the input data In the training phase, the parameters of the generator G and the discriminator D are updated alternately, and the discriminator D is discarded in the test phase. Finally, in the test phase, the network obtains the latent features z and reconstructed latent features of the EMG sample.

[0075] The original ganomaly network shows good performance in the field of image anomaly detection, but when it is applied to myoelectric gesture recognition, the network performance needs to be further improved. The model structure of the feature reconstruction network EMG-FRNet for myoelectric gesture recognition proposed in the embodiment of the present invention is as follows: Figure 4 As shown in (b), during testing, the established feature reconstruction network EMG-FRNet has the characteristics of small feature reconstruction error for target class samples in the target class dataset, but large feature reconstruction error for irrelevant class samples.

[0076] (1) Channel pruning strategy: used to implement channel pruning, such as Figure 4 As shown in (a), the original ganomaly network input is a three-channel RGB image, while the input in the preferred embodiment of the present invention is a single-channel electromyographic sample. The network has a large number of redundant feature channels, resulting in a large number of network parameters, increased training difficulty, and affected network performance. In order to reduce the redundant feature channels of the original ganomaly network and improve the performance of the network in electromyographic-independent gesture recognition, the channel clipping method clips the number of feature layer channels of the original ganomaly network according to a certain ratio, thereby reducing the redundant features of the network, improving the network accuracy and performance, and significantly reducing the scale of network parameters.

[0077] The present invention proposes a channel clipping method, which includes two parts: proportional channel clipping and power channel clipping.

[0078] Proportional channel cropping includes: cropping the number of channels of the network feature layer according to the channel ratio of the network input samples.

[0079] Power channel clipping includes ensuring that the number of clipped channels is still divisible by 2 based on proportional channel clipping to maximize the processing power of the computer. This is because when a computer processes data, the most efficient way is to divide the data into integer power of 2 units for processing.

[0080] The channel cutting method of this embodiment is as follows: Figure 5 As shown in the figure. Since the channel ratio of the three-channel RGB image and the single-channel EMG sample is 3:1, the proportional channel cropping part crops the feature map of the original ganomaly network according to the ratio of 3:1. The power channel cropping part ensures that the number of channels after cropping meets the requirement of an integer power of 2.

[0081] Figure 5 The channel pruning process of the downsampling layer feature map is shown in Figure 2. The number of channels of the downsampling feature layer is changed from (64, 128, 256) to (16, 32, 64). The number of channels of the upsampling feature layer is also modified similarly, from (256, 128, 64) to (64, 32, 16). In addition, the potential features z and The number of channels remains the same as 100. Figure 4 As shown in (b), this setting enables the latent feature z and the reconstructed latent feature Contains more information, so that Figure 1 The reconstruction error calculation in the irrelevant gesture discrimination module is shown to be more accurate.

[0082] (2) Cross-layer codec feature fusion strategy: used to implement cross-layer codec feature fusion. The original ganomaly network downsamples the original input data X to the latent feature z. E1 During the encoding process, due to the use of convolution and pooling operations, the original data is continuously compressed, and there will be information loss, resulting in the reconstruction of the input data from the potential feature z. The upsampling decoder G D During the decoding process, the available feature information is limited, which limits the reconstruction performance of the generator G and affects the recovery of the original data.

[0083] In order to compensate for the information loss in the downsampling process and improve the reconstruction performance of the generator, this embodiment adopts a cross-layer encoding and decoding feature fusion method. Figure 6 As shown, the encoder G in the figure E1 The feature blocks of G are visualized as dark parts, and the decoder G D The feature blocks are visualized as light-colored parts. E1 and decoder G DA series of skip connections are added between the upsampling layer and the downsampling layer to establish a direct path between the upsampling layer, so that the downsampling feature map and the upsampling feature map can be fused through the skip connection. The skip connection can directly pass the original high-resolution feature information to the decoder, thereby avoiding information loss in the downsampling process. At the same time, in order to further improve the richness of the features, the concat operation is used to concatenate features in the channel dimension, so that the feature information in the encoding and decoding processes can be effectively fused. Specifically, the implementation steps of this method are as follows: from the potential feature z to the reconstructed input data In the decoding recovery process, the potential feature z is first upsampled by 2 times to obtain the decoder G D 4×4×64 feature blocks. E1 and decoder G D The 4×4×64 feature blocks are concatenated into 4×4×128 feature blocks, and the decoder G is obtained by upsampling by 2 times D Then, the encoder G E1 and decoder G D The 8×8×32 feature blocks are concatenated into 8×8×64 feature blocks, and the decoder G is obtained by upsampling by 2 times D Finally, the encoder G E1 and decoder G D The 16×16×16 feature blocks are concatenated into 16×16×32 feature blocks, and the reconstructed input data is obtained by upsampling by 2 times

[0084] The cross-layer encoding and decoding feature fusion method can maximize the retention of feature information in the encoding and decoding process, make up for information loss, and make the input data X and reconstructed input data The degree of similarity is closer, which improves the reconstruction performance of the generator and ultimately improves the expressiveness of the model.

[0085] (3) SE channel attention strategy: used to implement the SE channel attention method. Since cross-layer codec feature fusion directly connects codec features of the same scale, this connection mechanism enables cross-level information transfer, thereby compensating for the problem of information loss. However, features of the same scale often contain similar but not identical information, so the feature information transmitted through cross-layer codec feature fusion may contain repeated redundant information. The presence of this redundant feature information will reduce the generalization ability of the network and may cause overfitting problems.

[0086] In order to avoid the problem of redundant information in cross-layer encoding and decoding feature fusion, this embodiment proposes to use the SE channel attention mechanism to apply different weights to each feature channel, strengthen the transmission of key feature information, and weaken the influence of redundant feature information. Figure 7 As shown, after obtaining the feature splicing graph M of the codec through cross-layer codec feature fusion, this embodiment obtains the feature splicing graph M-weight containing weight information through the SE channel attention mechanism, so that the decoder pays attention to important feature information during upsampling, thereby improving the performance of the network. The SE channel attention method includes a compression (Squeeze) operation, an excitation (Excitation) operation, and a scaling (Scale) operation; wherein the compression (Squeeze) operation is used to obtain the feature representation z of each channel of the feature splicing graph M of the codec; the excitation (Excitation) operation is used to obtain the weight vector s of each channel, and the scaling (Scale) operation is used to obtain the feature splicing graph M-weight containing weight information.

[0087] A. The squeeze operation includes: reducing the dimension of the feature concatenation map M of the codec by global average pooling so that each feature channel has a numerical representation, thereby obtaining a feature representation z, as shown in formula (1):

[0088]

[0089] Wherein, z represents the feature representation of each channel of the feature splicing map M of the codec, M represents the feature splicing map M of the codec, H and W represent the length and width of the feature splicing map M of the codec, respectively, and i, j represent the channel numbers in the length and width directions of the feature splicing map M of the codec, respectively;

[0090] B. The excitation operation includes: performing nonlinear transformation on the feature representation z of each channel of the feature concatenation map M of the codec, and mapping the result into a weight vector s. This process is completed through two fully connected layers. Different values ​​in the weight vector s represent the weight information of different channels, as shown in formula (2):

[0091] s=Excitation(z)=sigmoid(W2Relu(W1z)) (2);

[0092] Among them, s represents the weight vector, W1 represents the parameters of the first fully connected layer, Relu is the activation function of the first fully connected layer, W2 represents the parameters of the second fully connected layer, and sigmoid represents the activation function of the second fully connected layer.

[0093] C. The scaling operation includes: applying the weight vector s to the feature splicing map M of the codec to obtain a feature splicing map M-weight containing weight information. Specifically, by multiplication, the weight vector s is used to assign weights to the feature splicing map to obtain a feature splicing map containing weight information, as shown in formula (3), where M-weight is a feature splicing map containing weight information.

[0094] M-weight=Scale(M,s)=M×s (3)

[0095] The method of adding SE channel attention after cross-layer encoder-decoder feature fusion not only has advantages in retaining feature information in the encoding and decoding process, but also can further optimize the transmission of features and weaken the influence of redundant information, thereby improving the performance of the generator and the expressiveness of the model.

[0096] As a preferred embodiment, the generator and the discriminator are cross-trained based on the loss of the model structure of the feature reconstruction network EMG-FRNet for electromyography-independent gesture recognition, and the loss consists of two parts: the loss of the generator L G and the loss L of the discriminator D .

[0097] (1) Generator loss L G The loss is divided into three parts: the first part is the EMG sample reconstruction loss L1, which measures the input EMG sample X and the EMG sample restored by the generator through the L1 distance The difference between (specifically, the decoder in the generator performs the recovery action), the EMG sample reconstruction loss L1 is shown in formula (4); the second part of the loss is the EMG sample potential feature reconstruction loss L2, which is measured by the L2 distance between the EMG sample potential feature z and the EMG sample reconstruction potential feature The feature reconstruction loss L2 is shown in formula (5); the third part of the loss is the adversarial loss L3, which measures the difference between the true and false EMG samples in the middle layer features of the discriminator D through the L3 distance. The middle layer features are the features of size 4×4×64 in the discriminator D. The adversarial loss L3 is shown in formula (6), where D -2 (X) and They represent the middle layer features of true and false EMG samples in the discriminator D respectively.

[0098] The loss of the generator is L G It is composed of the above three parts of losses, as shown in formula (7), where w1, w2, and w3 respectively represent the proportion of the three parts of losses in the overall loss.

[0099] The loss ratio adopted in this embodiment is w1:w2:w3=50:1:1.

[0100]

[0101]

[0102]

[0103] L = w1L1 + w2L2 + w3L3 (7);

[0104] (2) The loss of the discriminator L D As shown in formula (8), the binary cross entropy BCE (Binary crossentropy) function enables the discriminator to better distinguish between generated EMG samples and real EMG samples.

[0105]

[0106] Where X represents the input electromyographic sample; represents the EMG sample restored by the generator; D(X) represents the characteristics of the input EMG sample in the discriminator, Represents the features of the EMG samples recovered by the generator in the discriminator;

[0107] S3, inputting the EMG sample into the feature reconstruction network EMG-FRNet, extracting the potential feature z of the input EMG sample, and reconstructing the potential feature of the EMG sample to obtain the reconstructed potential feature The final output is the latent feature z of the electromyographic sample and the reconstructed latent feature

[0108] S4, calculate the latent feature z and reconstruct the latent feature and compare the feature reconstruction error error with a predefined threshold threshold, and determine whether the electromyographic sample belongs to the target gesture or an irrelevant gesture based on the result of the comparison; wherein the feature reconstruction error error is a potential feature z and a reconstructed potential feature The difference between.

[0109] As a preferred embodiment, the calculation of the potential feature z and the reconstruction of the potential feature The feature reconstruction error between is shown in formula (5).

[0110] As a preferred implementation, determining that the electromyographic sample belongs to a target gesture or an irrelevant gesture based on the comparison result includes:

[0111] When the feature reconstruction error error is less than a predefined threshold threshold, it is determined that the electromyographic sample belongs to the target gesture;

[0112] When the feature reconstruction error error is greater than a predefined threshold threshold, it is determined that the electromyographic sample belongs to an irrelevant gesture.

[0113] In this embodiment, the reconstruction error error is compared with a predefined threshold threshold. If the reconstruction error is greater than the threshold, it is classified as an irrelevant gesture. If the reconstruction error is less than the threshold, it is classified as a target gesture. As shown in formula (9), 0 represents the target gesture and 1 represents the irrelevant gesture.

[0114]

[0115] Performance verification and evaluation of feature reconstruction network EMG-FRNet:

[0116] 1. Experimental environment and parameter settings of specific embodiments

[0117] The computer configuration used in the experiment of this embodiment is as follows: Intel Core i5-8250U CPU processor (8GB memory), NVIDIA GeForce 940MX graphics card (2GB video memory), Windows 10 operating system, and Python 3.7 programming language is used to implement network model training and testing under the Pytorch 1.2.0 deep learning framework.

[0118] The training parameter settings of the EMG-FRNet electromyography-independent gesture recognition feature reconstruction network are as follows: using the Adam optimizer, setting the smoothing constants (β1, β2) to (0.5, 0.999), the initial learning rate lr to 2e-3, the batch size to 64, and the number of training epochs to 200.

[0119] 2. Evaluation Indicators

[0120] This embodiment uses the area under the ROC (Receiver Operating Characteristic) curve AUC indicator to evaluate the performance of the model. The AUC value ranges from 0 to 1. The larger the AUC value, the better the performance of the model.

[0121] The confusion matrix is ​​the basis for drawing the ROC curve, such as Figure 8 As shown in the figure, each row of the confusion matrix represents the true category of the data, and each column represents the predicted category of the data. In the binary classification task, the final output of the model is 0 or 1, that is, positive or negative. In the confusion matrix, TP represents the true class, that is, the true class of the sample is the positive class, and the model recognition result is also the positive class. Similarly, FN represents the false negative class, FP represents the false positive class, and TN represents the true negative class.

[0122] According to formula (10) and formula (11), we can get the false positive rate FPR (False Positive Rate) and true positive rate TPR (True Positive Rate). FPR represents the ratio of samples that are incorrectly judged as positive among all samples that are actually negative. TPR represents the ratio of samples that are correctly judged as positive among all samples that are actually positive.

[0123]

[0124]

[0125] For the binary classification task, a fixed threshold is set to obtain a (FPR, TPR) pair. By plotting the (FPR, TPR) pairs corresponding to different thresholds on the coordinate system, the ROC curve can be obtained, as shown in Fig. 9 As shown in Figure 1, the ROC curve represents the recognition performance of the model under different thresholds. The ROC curve is the dark curve in the figure, the horizontal axis is the false positive rate FPR, and the vertical axis is the true positive rate TPR.

[0126] Furthermore, this embodiment uses the area under the ROC curve (AUC) value to quantitatively evaluate the irrelevant gesture recognition performance of the model.

[0127] 3. Comparative Experiment

[0128] The feature reconstruction network EMG-FRNet for myoelectric independent gesture recognition is compared with the more advanced methods in the existing field of independent gesture recognition: the method based on support vector machine SVDD and the method based on autoencoder AE.

[0129] Each subject needs to conduct multiple experiments, and each experiment sets a different target gesture category. The public dataset DB1 and the public dataset DB5 set 8 different target gestures respectively, and the self-collected dataset sets 7 different target gestures. There are 10 subjects in the public dataset DB1 and the public dataset DB5 respectively, and 6 subjects in the self-collected dataset. First, the experiment records the AUC values ​​when the subjects set different target gestures. Then, the AUC values ​​corresponding to different target gestures are averaged to obtain the AUC value of the subject. Finally, the AUC values ​​of all subjects in the dataset are averaged to obtain the AUC value of the dataset. Fig.10 The AUC line graphs for different target gestures set by the first subject in each dataset are shown. Fig.11 The AUC line graphs of different subjects in each dataset are shown. Table 1 shows the AUC values ​​of each dataset.

[0130] Table 1 AUC values ​​of each data set

[0131]

[0132] from Fig.10 It can be seen that for the same subject, when different target gestures are set, the AUC value of the EMG-FRNet model can always be maintained at a higher and more stable level compared with the comparison algorithms SVDD and AE. Fig.11 It can be seen that for different subjects, the performance of the EMG-FRNet model is still better than the comparison algorithms SVDD and AE. Specifically, as shown in Table 1, on different datasets DB1, DB5 and self-collected datasets, the AUC values ​​of EMG-FRNet are 0.940, 0.926, and 0.962, respectively, which are better than SVDD's 0.744, 0.753, and 0.723, and AE's 0.882, 0.819, and 0.902.

[0133] In the above comparative experiments, SVDD performed the worst, with an average AUC value of only 0.740. The reason may be that it uses the traditional single-class support vector machine method based on spherical hyperplanes, and the classification performance is poor for complex data distribution such as electromyographic samples. The average AUC value of AE reached 0.868, which was 0.128 higher than that of SVDD. The reason may be that AE uses the autoencoder method for feature learning and extraction, which can better mine the intrinsic characteristics of the data. The average AUC value of EMG-FRNet reached 0.943, which was 0.075 higher than that of AE. The reason may be that this method takes advantage of GAN in the field of anomaly detection, and combines strategies such as channel pruning, cross-layer encoding and decoding feature fusion, and SE channel attention to further improve the performance of the model in electromyographic-independent gesture recognition tasks.

[0134] The proposed EMG-FRNet achieves state-of-the-art (SOTA) performance in the task of myoelectricity-independent gesture recognition.

[0135] 4. Ablation Experiment

[0136] In order to verify the effectiveness of the method, the following ablation experiments were conducted: based on the original ganomaly network, channel lipping (CC), cross layer encoding and decoding feature fusion (CLEDF), and SE channel attention (SE) were added in sequence. In the following legend, ganomaly is represented as Baseline, ganomaly+CC is represented as Model 1, ganomaly+CC+CLEDF is represented as Model 2, and ganomaly+CC+CLEDF+SE (EMG-FRNet) is represented as Model 3.

[0137] Fig.12 The AUC line graphs for different target gestures set by the first subject in each dataset are shown. Fig.13 The AUC line graphs of different subjects in each dataset are shown. Table 2 shows the AUC values ​​of each dataset.

[0138] Table 2 AUC values ​​of each data set

[0139]

[0140] from Fig.12 It can be seen that for the same subject, when different target gestures are set, the performance of Baseline, Model 1, Model 2, and Model 3 improves in turn. Fig.13 It can be seen that the above results are still applicable to different subjects. Specifically, as shown in Table 2, on different datasets DB1, DB5 and self-collected datasets, the AUCs of Baeseline are 0.900, 0.739, and 0.870, respectively, the AUCs of model 1 are 0.902, 0.870, and 0.934, respectively, the AUCs of model 2 are 0.930, 0.919, and 0.960, respectively, and the AUCs of model 3 are 0.940, 0.926, and 0.962, respectively.

[0141] In the above ablation experiment, the average AUC value of ganomaly is 0.836. The average AUC value of model 1 is 0.902, which is 0.066 higher than the baseline. The reason is that channel pruning reduces the impact of redundant features in the original ganomaly network on model performance. The average AUC value of model 2 is 0.936, which is 0.034 higher than that of model 1. The reason is that cross-layer encoding and decoding feature fusion can make up for the information loss in the downsampling process. The average AUC value of model 3 is 0.943, which is 0.007 higher than that of model 2. The reason is that SE channel attention can avoid the redundant information problem caused by feature fusion, allowing the network to focus on more important features and improve the reconstruction performance of the model.

[0142] The effect of the model EMG-FRNet can achieve the best. Channel pruning, encoding and decoding feature fusion, and SE channel attention have improved the performance of the model in the electromyography-independent gesture recognition task to varying degrees.

[0143] Embodiment 2

[0144] See also Fig.14 This embodiment provides a myoelectricity-independent gesture recognition system, including:

[0145] The data processing module 101 is used to obtain a data set, and to form an electromyographic sample after data division and data preprocessing of the data set; wherein the electromyographic sample provides a data basis for training and testing of the network model;

[0146] A feature reconstruction network establishment module 102 is used to establish a feature reconstruction network EMG-FRNet, wherein the feature reconstruction network EMG-FRNet is established based on the original ganomaly network architecture by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and an SE channel attention strategy. The original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encode-decode-encode" structure composed of a first encoder, a decoder, and a second encoder;

[0147] The feature reconstruction module 103 is used to input the EMG sample into the feature reconstruction network EMG-FRNet, extract the potential feature z of the input EMG sample, and reconstruct the potential feature of the EMG sample to obtain the reconstructed potential feature z. Output the latent features z of the EMG sample and reconstruct the latent features

[0148] The irrelevant gesture identification module 104 is used to calculate the potential feature z and reconstruct the potential feature and compare the feature reconstruction error error with a predefined threshold value thresho1d, and determine whether the electromyographic sample belongs to the target gesture or an irrelevant gesture based on the result of the comparison; wherein the feature reconstruction error error is a potential feature z and a reconstructed potential feature The difference between.

[0149] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method described in the first embodiment.

[0150] like Fig.15 As shown, the present invention also provides an electronic device, including a processor 301 and a memory 302 connected to the processor 301, wherein the memory 302 stores a plurality of instructions, which can be loaded and executed by the processor so that the processor can execute the method described in Example 1.

[0151] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A myoelectric gesture recognition method based on feature reconstruction network EMG-FRNet, characterized in that: include: S1, obtaining a data set, dividing the data set and performing data preprocessing on the data set to form an electromyographic sample; S2, establish a feature reconstruction network EMG-FRNet, which is implemented by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and a SE channel attention strategy to the original ganomaly network architecture, wherein the original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encode-decode-encode" structure composed of a first encoder, a decoder, and a second encoder; S3, inputting the EMG sample into the feature reconstruction network EMG-FRNet, extracting the potential features of the input EMG sample, and reconstructing the potential features of the EMG sample to obtain the reconstructed potential features Then output the latent features of the EMG samples and reconstruct the latent features; S4, calculating a feature reconstruction error between the potential feature and the reconstructed potential feature, and comparing the feature reconstruction error with a predefined threshold, and determining whether the electromyographic sample belongs to a target gesture or an irrelevant gesture based on a result of the comparison; The channel clipping strategy is used to implement channel clipping, including proportional channel clipping and power channel clipping, wherein the proportional channel clipping includes clipping the number of channels of the network feature layer according to the channel ratio of the network input sample, and the power channel clipping includes changing the number of channels on the basis of proportional channel clipping, so that the number of channels after power channel clipping meets the requirement of an integer power of 2; The cross-layer codec feature fusion strategy is used to implement cross-layer codec feature fusion, adding multiple jump connections between the first encoder and decoder of the generator, establishing a direct path between the upsampling layer and the downsampling layer, and fusing the downsampling feature map and the upsampling feature map through the multiple jump connections, and the feature fusion obtains the feature splicing map of the codec by performing feature splicing in the channel dimension to perform the feature fusion on the feature information in the encoding and decoding processes; wherein the multiple jump connections include performing multiple feature fusions on the feature maps of different scales of the first encoder and the decoder; The SE channel attention strategy is used to implement the SE channel attention method, so that the feature splicing map containing weight information is obtained from the feature splicing map of the codec based on the SE channel attention mechanism; the SE channel attention method includes a compression operation, an excitation operation and a scaling operation; wherein the compression operation is used to obtain the feature representation of each channel of the feature splicing map of the codec z ; The excitation operation is used to obtain the weight vector of each channel, and the scaling operation is used to obtain the feature splicing map containing the weight information.

2. The method for distinguishing gestures independent of myoelectricity according to claim 1, characterized in that: The data set includes an existing and / or self-collected electromyographic data set containing multiple gesture movements; the data division includes a first division and a second division of the data set; the first division is to divide the electromyographic data into a target class data set and an irrelevant class data set, and the second division is to divide the electromyographic data into a training set and a test set, wherein the training set only has data related to the target class movement, and the test set has both data related to the target class movement and data related to the irrelevant class movement; the data preprocessing includes data segmentation and dimensionality transformation; wherein the data segmentation is based on a sliding window method to segment multi-channel electromyographic signals to obtain electromyographic samples of gesture movements; the dimensionality transformation includes transforming a first dimension and a second dimension of a two-dimensional matrix formed by the multi-channel electromyographic samples after the data segmentation; the transformed two-dimensional matrix is ​​more adaptable to the input requirements of the feature reconstruction network EMG-FRNet than the two-dimensional matrix before the transformation.

3. The method for distinguishing gestures independent of myoelectricity according to claim 2, characterized in that: The S2 includes: S21, build ganomaly network infrastructure; S22, based on the original ganomaly network infrastructure, adds channel pruning strategy, cross-layer encoding and decoding feature fusion strategy and SE channel attention strategy to form the initial feature reconstruction network EMG-FRNet; S23, using the target class data set to train the initial feature reconstruction network EMG-FRNet, using all categories of data sets as the test set to test the initial feature reconstruction network EMG-FRNet, and continuously repeating the training and the testing until the test effect reaches the optimal, thereby forming the feature reconstruction network EMG-FRNet.

4. The method for distinguishing gestures independent of myoelectricity according to claim 3, characterized in that: The compression operation includes: performing global average pooling on the feature concatenation map of the encoder and decoder M Dimensionality reduction is performed so that each feature channel has a numerical representation, thereby obtaining the feature representation of each channel of the feature concatenation graph of the codec, as shown in formula (1): (1); in, z Feature concatenation graph representing the encoder and decoder M The feature representation of each channel, M represents the feature concatenation graph of the encoder and decoder, H and W Represents the feature concatenation graph of the encoder and decoder respectively M The length and width of i, j Represents the feature concatenation graph of the encoder and decoder respectively M The channel numbers in the length and width directions; The excitation operation includes: performing nonlinear transformation on the feature representation of each channel of the feature concatenation graph of the codec, and mapping the result into a weight vector, where different values ​​in the weight vector represent weight information of different channels, as shown in formula (2): (2); in, s represents the weight vector, Represents the first fully connected layer parameters, Re lu is the activation function of the first fully connected layer, Represents the second fully connected layer parameters, sigmoid represents the activation function of the second fully connected layer; The scaling operation includes: assigning weights to the feature splicing map using a weight vector through a multiplication operation to obtain a feature splicing map containing weight information, as shown in formula (3). (3); in, M-weight It is a feature concatenation graph containing weight information.

5. The method for distinguishing gestures independent of myoelectricity according to claim 4, characterized in that: The training of the initial feature reconstruction network EMG-FRNet using the target class dataset includes: cross-training the generator and the discriminator based on the loss of the model structure of the feature reconstruction network EMG-FRNet for electromyography-independent gesture recognition, the loss includes the loss of the generator and the loss of the discriminator; the loss of the generator includes the electromyography sample reconstruction loss, the electromyography sample potential feature reconstruction loss and the adversarial loss, the electromyography sample reconstruction loss represents the difference between the input electromyography sample and the electromyography sample restored by the generator, the electromyography sample potential feature reconstruction loss represents the difference between the electromyography sample potential feature and the electromyography sample reconstructed potential feature, the adversarial loss represents the difference between the true and false electromyography samples in the intermediate layer features of the discriminator, the generator loss is the weighted sum of the electromyography sample reconstruction loss, the electromyography sample potential feature reconstruction loss and the adversarial loss; the loss of the discriminator is obtained based on the binary cross entropy function.

6. The method for distinguishing gestures independent of myoelectricity according to claim 5, characterized in that: The feature reconstruction error in S4 is the difference between the latent feature and the reconstructed latent feature; Determining whether the electromyographic sample belongs to a target gesture or an irrelevant gesture based on the comparison result includes: When the feature reconstruction error is less than a predefined threshold, determining that the electromyographic sample belongs to the target gesture; When the feature reconstruction error is greater than a predefined threshold, it is determined that the electromyographic sample belongs to an irrelevant gesture.

7. A myoelectricity-independent gesture recognition system, implemented based on a feature reconstruction network EMG-FRNet, for implementing the myoelectricity-independent gesture recognition method according to any one of claims 1 to 6, characterized in that: include: A data processing module (101) is used to obtain a data set, and to perform data division and data preprocessing on the data set to form an electromyographic sample; The electromyographic samples provide a data basis for the training and testing of the network model; A feature reconstruction network establishment module (102) is used to establish a feature reconstruction network EMG-FRNet, wherein the feature reconstruction network EMG-FRNet is established based on the original ganomaly network architecture by adding a channel pruning strategy, a cross-layer encoding and decoding feature fusion strategy, and an SE channel attention strategy, wherein the original ganomaly network architecture includes a generator and a discriminator, wherein the generator is an "encode-decode-encode" structure composed of a first encoder, a decoder, and a second encoder; A feature reconstruction module (103) is used to input the electromyographic sample into the feature reconstruction network EMG-FRNet to extract the potential features of the input electromyographic sample. z , and reconstruct the potential features of the electromyographic samples to obtain the reconstructed potential features , output the potential features of the EMG sample z and reconstructing latent features ; An irrelevant gesture discrimination module (104) for calculating potential features z and reconstructing latent features and comparing the feature reconstruction error error with a predefined threshold threshold, and determining whether the electromyographic sample belongs to a target gesture or an irrelevant gesture based on the comparison result; wherein the feature reconstruction error error is a potential feature z and reconstructing latent features The difference between.

8. An electronic device, comprising a processor and a memory, wherein the memory stores a plurality of instructions, characterized in that: The processor is used to read the instruction and execute the myoelectricity-independent gesture identification method as described in any one of claims 1-6.

9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The multiple instructions can be read by a processor and execute the myoelectricity-independent gesture identification method as described in any one of claims 1-6.