A lightweight electromyogram gesture recognition method based on multi-domain feature fusion
Through the lightweight EMG gesture recognition method and data cache strategy based on multi-domain feature fusion, the problems of low computational efficiency and high energy consumption of the EMG signal recognition model are solved, and efficient real-time control and accurate gesture recognition of intelligent bionic hands are realized.
Patent Information
- Application Number
- CN202411581992.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-07
AI Technical Summary
The existing electromyography signal recognition model has low computational efficiency and high energy consumption, resulting in the problem of delay in action decoding execution and misjudgment delay in gesture switching during real-time control of intelligent bionic hand.
A lightweight EMG gesture recognition method based on multi-domain feature fusion is proposed. By constructing input data sets, preprocessing data, training lightweight gesture recognition model, and using data caching method to solve the delay problem in real-time control.
The model is lightweight, the computing efficiency and recognition accuracy are improved, the delay and misjudgment problems in real-time control of intelligent bionic hands are solved, and the overall performance of the system is improved.
Smart Images

Figure CN119475033B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of myoelectric gesture recognition, and in particular to a lightweight myoelectric gesture recognition method based on multi-domain feature fusion. Background Art
[0002] Gesture recognition based on surface electromyography (sEMG) signals is one of the key technologies in intelligent human-computer interaction. By studying and understanding human movement intentions, sEMG-based gesture recognition technology can achieve intuitive and efficient control of prostheses, significantly improve the naturalness of operation, improve the quality of life of patients with upper limb amputations, and play a key role in promoting technological innovation.
[0003] In existing research on electromyographic gesture recognition, deep learning algorithms are usually used to improve the accuracy of gesture recognition, such as the TFN-FICFM network designed by HU et al. First, a temporal fusion network (Temporal Fusion Network) is used to obtain multi-level temporal feature representation using an attention-based recursive multi-scale convolution module, and deep fusion of temporal features is achieved through a feature pyramid module. The deeply fused temporal features are then used to generate multiple sets of gesture category prediction confidences through a feedback loop. Then, a fuzzy integral-based classifier fusion method (Fuzzy Integral-based Classifier Fusion) is used to fuzzy fuse the prediction confidences to improve the accuracy and robustness of gesture recognition. Karnam et al. proposed a hybrid CNN and Bi-LSTM architecture EMGHandNet, which effectively combines the advantages of CNN in extracting spatial features and the advantages of Bi-LSTM in encoding temporal dependencies to improve the accuracy of gesture recognition.
[0004] However, the above studies mostly focus on improving the accuracy of gesture recognition, but model computing efficiency is also one of the important factors affecting the commercial development of prostheses. In the research on electromyographic signal recognition, there is little research on the lightweight of deep learning models, and there are often problems of low model computing efficiency and high energy consumption.
[0005] Secondly, in order to improve the recognition accuracy, most researchers have focused on the continuous optimization of the network structure, ignoring the fact that myoelectric signals in different domains have different characteristic information. The fusion of different domain features can also improve the recognition accuracy, but there are many modal decomposition methods for extracting time-frequency domain signals, most of which have problems such as modal component aliasing and boundary effects. In addition, due to the continuity of myoelectric signals and the phenomenon of too fast model inference time, there are phenomena such as action decoding execution delay and gesture switching misjudgment delay in the real-time control process of the intelligent bionic hand. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention proposes a lightweight electromyographic gesture recognition method based on multi-domain feature fusion and a real-time control system based on the method, which solves the problems of low computational efficiency and high energy consumption of the current electromyographic signal recognition model. At the same time, the present invention proposes a data caching method, which solves the problems of action decoding execution delay and gesture switching misjudgment delay in the real-time control process of the intelligent bionic hand in the prior art.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] The first aspect of the present invention proposes a lightweight electromyographic gesture recognition method based on multi-domain feature fusion, which specifically includes the following steps:
[0009] S1: construct an input dataset, which includes a public dataset and a self-made dataset;
[0010] S2: Perform preprocessing operations on the data in the input data set and divide the processed data into a training set and a test set;
[0011] S3: input the data in the training set into the lightweight gesture recognition model to train the lightweight gesture recognition model;
[0012] The lightweight gesture recognition model includes two identical feature extraction networks, each of which consists of a point-by-point convolution layer, four improved Vanilla module layers and a global maximum pooling layer;
[0013] S4: Input the data in the test set into the trained lightweight gesture recognition model for evaluation to obtain the weight file and evaluation results.
[0014] Preferably, the lightweight gesture recognition model changes the channel dimension through a point-by-point convolution layer and performs cross-channel information fusion, generates high-dimensional features through an improved Vanilla module layer extraction, compresses the spatial dimension of the high-dimensional features to 1 through a global maximum pooling layer, and then serially fuses high-dimensional features from different domains, and finally uses two convolution layers to predict the results;
[0015] The improved Vanilla module consists of an improved convolution layer + a normalization layer, an activation function, and an improved convolution layer + a normalization layer.
[0016] Preferably, the step of generating high-dimensional features by extracting the improved Vanilla module layer includes:
[0017] 1): Use the activation function A(x) to train two improved convolutional layers at the same time. The improved convolutional layer uses Omni-Dimensional dynamic convolution to replace ordinary convolution. When the training reaches the identity mapping, the outputs of the two convolutional layers are the same, and then they are merged. The identity mapping calculation formula is as follows:
[0018] A′(x)=(1-λ)A(x)+λx
[0019] Among them, A′(x) represents the modified activation function, λ represents the nonlinear hyperparameter, e and E represent the current round number and the total number of training rounds, and λ=e / E. In the initial stage of training, e=0, A′(x)=A(x), at which time the network has strong nonlinearity. When the training converges, A′(x)=x, at which time there is no activation function between the two convolutional layers, and there is no nonlinearity.
[0020] 2): Combine the two improved convolutional layers + normalization layers and convert them into 1×1 convolution. Represented as having C in Input channels, C out The weight matrix and bias matrix of the output channels and kernel size k, the size, bias, mean and variance in the normalization layer are expressed as The weight matrix and bias matrix after the improved convolution layer + normalization layer are combined are as follows:
[0021]
[0022] where the subscript i∈{1,2…,C out} represents the value in the i-th output channel;
[0023] 3): Merge the two 1×1 convolutions obtained in step 2. and As input and output features, the convolution formula is as follows:
[0024] y=W*x=W·im2col(x)=W·X
[0025] Among them, * represents convolution operation, · represents matrix multiplication, It is derived from the im2col(x) operation. The weight matrices of the two convolutional layers are represented as W 1 and W 2 , for 1×1 convolution, the im2col(x) operation does not require overlapping sliding kernels, so the formula for merging two 1×1 convolutions is as follows:
[0026] y=W 1 *(W 2 *x)=W 1 ·W 2im2col(x)=(W 1 ·W 2 )*X;
[0027] 4): During the training process, the nonlinearity of each activation layer is increased by stacking activation functions concurrently. The stacked activation function can be expressed as follows:
[0028]
[0029] Where x represents the input, A(x) represents a single activation function, n represents the number of activation functions, and a i , b i is the size and bias of each activation. Given an input feature tensor Where H, W and C represent its width, height and number of channels respectively, and the activation function can be expressed as:
[0030]
[0031] Among them, h∈{1,2,…,H},w∈{1,2,…,W}, and c∈{1,2,…,C}. When n=0, the activation function A s (x) degenerates into a pure activation function A(x).
[0032] Preferably, an input data set is constructed in S1, including DB2, DB5 data sets in NinaPro and a self-made data set, and is transmitted to data preprocessing through an acquisition protocol.
[0033] Preferably, in S2, the preprocessing operation on the input data comprises the following steps:
[0034] 1): The input data is processed in sequence by high-frequency denoising, full-wave rectification, data smoothing, and low-frequency denoising;
[0035] 2): The data obtained in step 1 is first divided into actions using the sliding window method, and then the actions corresponding to each data label are segmented and reorganized. Finally, the corresponding actions within the same label are merged and stored separately to obtain time domain data;
[0036] 3): Perform variational mode decomposition on the time domain data obtained in step 2, and use the time-frequency domain data of the decomposed first modal signal component and the time domain data obtained in step 2 as multi-domain inputs of the network model;
[0037] 4): Divide the data obtained in step 4 into sEMG segments of equal size;
[0038] 5): Normalize the data obtained in step 5 and then adjust the data input format, and divide the time domain data and time-frequency domain data into training set and test set respectively.
[0039] Preferably, in S3, each training process traverses all samples in the training set, first performs forward propagation to calculate the prediction result, then calculates the error according to the loss function, and then performs back propagation to update the model weights to minimize the loss;
[0040] During training, the loss function, optimizer, initial optimization rate, training batch size, and number of training rounds are set, and a strategy to prevent overfitting is adopted.
[0041] Preferably, the loss function adopts the cross entropy loss function, the optimizer adopts SGD, the optimization rate is set to 0.001, the training batch size is set to 128, the number of training rounds is 200, and random inactivation, data enhancement and stratified K-fold cross validation are set to reduce overfitting.
[0042] Preferably, in S4, the evaluation indicators include calculation accuracy, model reasoning time, floating-point calculation amount, model parameter amount, weight file size, recall rate and average accuracy of all categories.
[0043] The second aspect of the present invention proposes a real-time control system for electromyographic gesture recognition, in which the weight file and model configuration file obtained by the above method are transplanted to the edge device for real-time decoding of the electromyographic signal, and the decoded prediction results are transmitted to the intelligent bionic hand. The intelligent bionic hand acts as an actuator and executes the received instructions through multiple degree-of-freedom motors to complete gesture movements.
[0044] The third aspect of the present invention proposes a data caching method. When using the above-mentioned electromyographic gesture recognition real-time control system to perform real-time control of an intelligent bionic hand, the cache amount of electromyographic signal data collected once during the real-time control process is set to a plurality of sample numbers, and a plurality of samples correspond to a plurality of detection results. By statistically analyzing the detection results, the result with the highest frequency of occurrence is selected as the final electromyographic gesture determination result.
[0045] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects.
[0046] 1. The present invention improves the VanillaNet network structure in data input and structural simplification so that it can be applied to electromyographic signals. Finally, while ensuring the recognition accuracy, the model is lightweight: on the DB2 data set, the model parameter volume is 1.43M, the floating point operation volume is 4.52G, and the reasoning time is 18.11ms; on the DB5 data set, the model parameter volume is 1.43M, the floating point operation volume is 8.05G, and the reasoning time is 45.69ms; on the self-built data set, the model parameter volume is 1.43M, the floating point operation volume is 0.01G, the weight volume is 51.9MB, and the reasoning time is 47.06ms; in the real-time control system of the intelligent bionic hand, the entire bionic hand completes an action in about 300ms, which improves the computing efficiency.
[0047] 2. In the data preprocessing stage, this application improves the modal decomposition method and multi-domain data fusion to avoid the problem of mode aliasing. On the DB2 data set, the precision rate is 95.5%, the recall rate is 95.78%, and the accuracy rate is 95.52%; on the DB5 data set, the precision rate is 92.0%, the recall rate is 92.77%, and the accuracy rate is 92.08%; on the self-built database, the accuracy rate is 99.57%; in the real-time control system of the intelligent bionic hand, the average success rate is 90.17%, which greatly improves the recognition accuracy.
[0048] 3. By adopting the data caching method of the present invention, the problem of misjudgment during gesture switching is effectively solved. Due to the existence of data caching, the bionic hand will not immediately execute the detected gesture, but will execute the corresponding gesture when the result frequency counted during the caching period is the highest, thereby reducing the misjudgment caused by gesture switching. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a schematic diagram of a lightweight network gesture recognition model of the present invention;
[0050] Figure 2 It is a hand gesture diagram of DB2 and DB5 data sets of the present invention;
[0051] Figure 3 This is a hand gesture diagram of the self-built database of the present invention;
[0052] Figure 4 It is a filtering comparison schematic diagram of the present invention;
[0053] Figure 5 This is a schematic diagram of segmenting and reassembling the same hand gesture signal of the present invention;
[0054] Figure 6 It is a schematic diagram of normalization of time domain data and time-frequency domain data of the present invention;
[0055] Figure 7It is a schematic diagram of the comparison of the Pearson correlation coefficients of the original data, filtered data and VMD modal decomposition data of the same gesture in the present invention;
[0056] Figure 8 This is a schematic diagram of the test accuracy and loss calculation diagram of the DB2 and DB5 data sets of the present invention;
[0057] Fig. 9 Schematic diagram of confusion matrix of DB2 (50 categories) and DB5 (53 categories) databases of the present invention;
[0058] Fig.10 This is a visualization diagram of the high-dimensional features of the DB2 and DB5 databases using the t-SNE technology of the present invention;
[0059] Fig.11 It is a schematic diagram of the myoelectric real-time control system of the present invention;
[0060] Fig.12 It is a real-time control schematic diagram of the present invention;
[0061] Fig.13 It is a schematic diagram of the accuracy of the bionic hand control of the present invention. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] The first aspect of the present invention proposes a lightweight electromyographic gesture recognition method based on multi-domain feature fusion, which specifically includes the following steps:
[0064] S1: construct an input dataset, which includes a public dataset and a self-made dataset;
[0065] S2: Perform preprocessing operations on the data in the input data set and divide the processed data into a training set and a test set;
[0066] S3: input the data in the training set into the lightweight gesture recognition model to train the lightweight gesture recognition model;
[0067] The lightweight gesture recognition model consists of two identical feature extraction networks, each of which consists of a point-by-point convolution layer, four improved Vanilla module layers and a global maximum pooling layer 1;
[0068] S4: Input the data in the test set into the trained lightweight gesture recognition model for evaluation to obtain the weight file and evaluation results.
[0069] Figure 1 This is a schematic diagram of a lightweight electromyographic gesture recognition method model based on multi-domain feature fusion. The model consists of two identical feature extraction networks. The two networks can extract features from different domains synchronously without interfering with each other. Each feature extraction network consists of a point-by-point convolution layer that changes the channel dimension and performs cross-channel information fusion, 4 improved Vanilla module layers that extract and generate high-dimensional features, and a global maximum pooling layer that compresses the spatial dimension of high-dimensional features to 1. The high-dimensional features of the two domains are then fused in series, and finally two layers of convolution layers are used to predict the results.
[0070] The following table compares the model accuracy, model parameter count, floating point number calculation, and average inference time when the feature layers are 5, 6, and 7. When the feature extraction network is designed to be 6 layers deep, the performance is optimal.
[0071] Table 1
[0072]
[0073] After the preprocessed time domain data and time-frequency domain data are adjusted to meet the network input, feature extraction is performed through the point-by-point convolution layer and the improved Vanilla module layer.
[0074] Step 1: The improved Vanilla module layer consists of an improved convolution layer + normalization layer, an activation function, an improved convolution layer + normalization layer. The improved convolution layer improves the ordinary convolution into an Omni-Dimensional dynamic convolution, and considers the dynamics in the spatial domain, input channel, output channel and other dimensions. The activation function A(x) is used to train two improved convolution layers at the same time. When the training reaches the identity mapping, the two convolution layers have the same output and are then merged. The identity mapping calculation formula is as follows:
[0075] A′(x)=(1-λ)A(x)+λx
[0076] Among them, A′(x) represents the modified activation function, λ represents the nonlinear hyperparameter, e and E represent the current round number and the total number of training rounds, and λ=e / E. In the initial stage of training, e=0, A′(x)=A(x), at which time the network has strong nonlinearity. When the training converges, A′(x)=x, at which time there is no activation function between the two convolutional layers, and there is no nonlinearity.
[0077] Step 2: Merge the two improved convolutional layers + normalization layers into a 1×1 convolution. Represented as having C in Input channels, Cout The weight matrix and bias matrix of the output channels and kernel size k, the size, bias, mean and variance in the normalization layer are expressed as The weight matrix and bias matrix after the improved convolution layer + normalization layer are combined are as follows:
[0078]
[0079] where the subscript i∈{1,2…,C out} represents the value in the i-th output channel;
[0080] Step 3: Merge the two 1×1 convolutions obtained in step 2. and As input and output features, the convolution formula is as follows:
[0081] y=W*x=W·im2col(x)=W·X
[0082] Among them, * represents convolution operation, · represents matrix multiplication, It is derived from the im2col(x) operation. The weight matrices of the two convolutional layers are represented as W 1 and W 2 , for 1×1 convolution, the im2col(x) operation does not require overlapping sliding kernels, so the formula for merging two 1×1 convolutions is as follows:
[0083] y=W 1 *(W 2 *x)=W 1 ·W 2 im2col(x)=(W 1 ·W 2 )*X
[0084] Step 4: Increase the nonlinearity of each activation layer by stacking activation functions concurrently during training. The stacked activation function can be expressed as follows:
[0085]
[0086] Where x represents the input, A(x) represents a single activation function, n represents the number of activation functions, and a i , b i is the size and bias of each activation. Given an input feature tensor Where H, W and C represent its width, height and number of channels respectively, and the activation function can be expressed as:
[0087]
[0088] Among them, h∈{1,2,...,H},, and c∈{1,2,…,C}. When n=0, the activation function A s (x) degenerates into a pure activation function A(x);
[0089] Technical means to improve accuracy:
[0090] NinaPro is a public and highly recognized electromyographic dataset that collects muscle locations including the superficial flexor and superficial extensor digitorum and below the elbow joint of the forearm. In this study, the data used are from the DB2 and DB5 databases in NinaPro, which include tasks related to upper limb movement.
[0091] See also Figure 2 , DB2 contains 40 complete subjects, consisting of 29 healthy males and 11 healthy females. The classified motor movements cover 49 movements, plus rest for a total of 50 categories (exercise B, C, D and rest gestures). The dataset uses a Delsys Trigno wireless system to sample 12 channels of sEMG signals at a rate of 2kHz. During the acquisition process, the subjects were asked to repeat these movements 10 times with their right hand. Each movement repetition lasted for 5 seconds, followed by a 3-second rest. DB5 contains 10 complete subjects, consisting of 8 healthy males and 2 healthy females. The classified motor movements cover 52 movements, plus rest for a total of 53 categories (exercise A, B, C and rest gestures). The dataset uses two Thalmic Myo armbands to sample 16 channels of Surface Electromyography signals at a rate of 200Hz. During the acquisition process, the subjects were asked to repeat these movements 6 times with their right hand. Each movement repetition lasted for 5 seconds, followed by a 3-second rest.
[0092] See also Figure 3 , the Self-Build Dataset (SBD) contains 5 complete subjects, consisting of 4 healthy men and 1 healthy woman. The classified movements include 5 movements, plus rest, a total of 6 categories (extending the little finger, extending the thumb, resting, extending the index finger, making a fist, and OK gesture). Data collection uses the armband of the BrainCo smart bionic hand to sample 8 channels of sEMG signals at a rate of 2000Hz. During the collection process, the subjects were required to repeat these movements 10 times with their right hands. Each movement repetition lasted for 5 seconds, followed by a 3-second rest.
[0093] Data preprocessing operations:
[0094] Step 1: Process the input data in sequence by high-frequency denoising, full-wave rectification, data smoothing, and low-frequency denoising to improve the signal-to-noise ratio of data in different databases, enhance amplitude characteristics, highlight signal trends, and improve signal quality. Figure 4 ;
[0095] Step 2: Use the sliding window method to divide the data obtained in step 1 into actions, then split and reorganize the actions corresponding to each data label, and finally merge the corresponding actions in the same label and store them separately to obtain time domain data. For the results, please refer to Figure 5 ;
[0096] Step 3: Perform Variational Mode Decomposition on the time domain data obtained in step 2. Use the time-frequency domain data of the first modal signal component after decomposition and the time domain data obtained in step 2 as the multi-domain input of the network model. Use the Pearson correlation coefficient to analyze the gesture sample data after decomposition. Figure 6 ;
[0097] Step 4: Perform data enhancement on the data obtained in step 3 to enhance the performance of the model for gesture recognition tasks;
[0098] Step 5. Use a sliding window to divide the EMG signal obtained in step 4 into sEMG analysis windows of the same size. In the DB2 dataset, the sampling rate is 2000Hz and the number of channels is 12, so the Hamming window size is set to 96ms, the step size is set to 25ms, and the input size is (192, 12). In the DB5 dataset, the sampling rate is 200Hz and the number of channels is 16, so the Hamming window size is set to 320ms, the step size is set to 50ms, and the input size is (64, 16). In the self-made dataset, the sampling rate is 2000Hz and the number of channels is 8, so the Hamming window size is set to 64ms, the step size is set to 10ms, and the input size is (128, 8);
[0099] Step 6: Normalize the data obtained in step 5 and then adjust the data input format. Figure 7 , and finally transmit the signal to the gesture recognition module;
[0100] In each training process, all samples in the training set are traversed, and forward propagation is first performed to calculate the prediction results, then the error is calculated according to the loss function, and then backpropagation is performed to update the model weights to minimize the loss. During training, random inactivation (inactivation rate 0.5), data enhancement, and stratified K-fold cross-validation are set to reduce overfitting. The loss function uses the cross entropy loss function, the optimizer uses SGD, the optimization rate is set to 0.001, the batch_size is set to 128, and the number of training rounds is 200 rounds;
[0101] After training is completed, the performance of the model is evaluated on the test set, and indicators such as accuracy, model inference time, floating-point calculation amount, model parameter amount, weight file size, recall rate, and average accuracy of all categories are calculated. During the entire training process, the model weight file with the best performance will be saved.
[0102] See also Figure 8 ,To detect the convergence speed and recognition performance, the proposed model was tested on Ninapro DB2 and Ninapro DB5. ,The test results of typical subjects (test accuracy and loss calculation graph).
[0103] Ninapro DB2 converges in about 50 rounds, with an accuracy rate of about 95%, and Ninapro DB5 converges in about 150 rounds, with an accuracy rate of about 93%. The difference in convergence time of network models is mainly affected by the complexity of input data and the range of sample variation. In terms of data sets, DB2 uses 12 channels and 50 types of gestures, while DB5 uses 16 channels and 53 types of gestures. Compared with DB5, DB2 has fewer channels, which reduces interference between channels. In addition, DB2 also has fewer gesture categories, which further simplifies the recognition task. Therefore, DB2's model usually exhibits faster convergence during training.
[0104] Comparing the performance of the model on the Pre, Re, and F1 evaluation indicators (see Table 2), on the DB2 dataset, the Pre, Re, and F1 scores are 95.5%, 95.78%, and 95.52%, respectively; and on DB5, the corresponding scores are 92.0%, 92.77%, and 92.08%. This shows that the performance on the two public datasets is similar, reflecting the overall superiority of the model. However, the data indicators of DB2 are slightly higher, mainly due to the lower data complexity of DB2 than DB5, and the input sample size of DB2 is 1*48*48, compared with the sample size of DB5 is 1*32*32. Larger input samples provide more feature information, resulting in better results on DB2.
[0105] Table 2
[0106]
[0107] See also Fig. 9 ,To demonstrate the capability of the present model in gesture classification, we calculated the confusion matrices of the 8th subject of the Ninapro DB2 database and the 4th subject of the Ninapro DB5 database, as shown Fig.10As shown. In NinaproDB2, the calculated results show that the average accuracy of gestures is 95.78%. The classification accuracy of most gestures exceeds 90%, and the accuracy of several gestures reaches 100%. However, the accuracy of the 30th, 35th and 47th gestures is less than 90%, showing recognition confusion. This shows that these gestures have certain similarities with other gestures in signal characteristics, which leads to classification difficulties. In Ninapro DB5, the calculated average gesture accuracy is 92.77%. Although the classification accuracy of most gestures exceeds 80%, and the accuracy of some gestures reaches 100%, the accuracy of the 25th, 39th, 42nd, 46th and 51st gestures is less than 80%, and recognition confusion also occurs. The lower accuracy is mainly due to the smaller input feature size and the large number of gesture categories in DB5, which increases the complexity of classification. Overall, the model performs well in fusing the time domain features and time-frequency domain features of the electromyographic signals, and shows good recognition performance in the multi-gesture classification task. See Fig.10, the spatial distribution of the 6 gesture features of the 8th subject in Ninapro DB2. These gestures include: the 1st gesture (label0, dark blue, accuracy: 90.91%), the 14th gesture (label1, dark purple, accuracy: 89.47%), the 26th gesture (label2, light purple, accuracy: 86.96%), the 35th gesture (label3, orange, accuracy: 70.59%), the 42nd gesture (label4, orange, accuracy: 95.24%) and the 50th gesture (label5, golden, accuracy: 100%). In the Ninapro DB2 dataset, the features of the 1st gesture (label0), the 42nd gesture (label4) and the 50th gesture (label5) are clustered more closely and have obvious intervals with other gestures, showing strong discrimination ability. However, the features of the 26th gesture (label2) and the 35th gesture (label3) are obviously scattered, indicating that their clustering ability is weak. This verifies the difference in feature clustering under different accuracy rates. Spatial distribution of the 6 gesture features of the 4th subject in Ninapro DB5. These gestures include: the 3rd gesture (label0, dark blue, accuracy: 100%), the 11th gesture (label1, dark purple, accuracy: 83.33%), the 25th gesture (label2, light purple, accuracy: 77.78%), the 46th gesture (label3, orange, accuracy: 65.22%), the 50th gesture (label4, orange, accuracy: 87.50%) and the 53rd gesture (label5, golden, accuracy: 90.20%). In the Ninapro DB5 dataset, there is a clear gap between the features of the 3rd gesture (label0) and the other gestures, and its clustering strength is the most significant, indicating that it has the strongest discriminative ability. In contrast, the features of the 46th gesture (label3) are more scattered, and there is a certain overlap with the 25th gesture (label2). This shows that the features of the 46th gesture are relatively vague, resulting in a low classification accuracy. The features of the remaining gestures are scattered to varying degrees, reflecting their relatively weak discrimination capabilities. These visualization results reveal the relationship between the degree of feature clustering and classification accuracy. Through the intuitive display of the feature space, we can clearly observe the separation between different categories and the clustering effect of each gesture.
[0108] Technical solutions to the delay problem of intelligent prosthetic hand control
[0109] See also Fig.11,The obtained model weights and model configuration files are transplanted to the edge device for real-time decoding of electromyographic signals, and the decoded prediction results are transmitted to the intelligent bionic hand. The intelligent bionic hand acts as an actuator and executes the received instructions through multiple degrees of freedom motors to complete the gesture. In the real-time control process, a data caching strategy is designed to solve the delay caused by the intelligent bionic hand device itself;
[0110] See also Fig.12 , the experimental setup of real-time control is presented, and the network model proposed on the BrainCo intelligent bionic hand of Qiangnao Technology is verified. In order to overcome the problem of packet loss that may occur in the real-time acquisition process of ordinary laptop Bluetooth devices, we use the LX1815 Bluetooth adapter to improve compatibility and stability. A data caching strategy is proposed. Specifically, we set the cache amount of the electromyographic signal data collected at one time to 125 milliseconds. Since the sampling rate is 2000Hz, the number of cached samples is 250. In this 500 millisecond period, a total of 7 detection results will be obtained. By counting the detection results, the result with the highest frequency is selected as the final electromyographic gesture judgment, which effectively solves the first delay phenomenon, that is, the problem that the bionic hand has not completed the action after the model decodes the gesture. In addition, the caching strategy also effectively solves the problem of misjudgment during gesture switching. Due to the existence of data caching, the bionic hand will not execute the detected gesture immediately, but will execute the corresponding gesture when the result frequency counted during the caching period is the highest, thereby reducing the misjudgment caused by gesture switching.
[0111] See also Fig.13 , showing the accuracy of bionic hand control during the experiment. It can be seen that the success rate of each action is around 90%, and the average success rate reaches 90.17%. The time for the entire bionic hand to complete an action is about 300ms, which can meet the needs of users in actual application scenarios.
Claims
1. A lightweight electromyographic gesture recognition method based on multi-domain feature fusion, characterized in that: The following steps are involved: S1: construct an input dataset, which includes a public dataset and a self-made dataset; S2: Perform preprocessing operations on the data in the input data set and divide the processed data into a training set and a test set; S3: input the data in the training set into the lightweight gesture recognition model to train the lightweight gesture recognition model; The lightweight gesture recognition model includes two identical feature extraction networks, each of which consists of a point-by-point convolution layer, four improved Vanilla module layers and a global maximum pooling layer; S4: Input the data in the test set into the trained lightweight gesture recognition model for evaluation, and obtain the weight file and evaluation results; The lightweight gesture recognition model described in step S3 changes the channel dimension through a point-by-point convolution layer and performs cross-channel information fusion, generates high-dimensional features through an improved Vanilla module layer extraction, compresses the spatial dimension of the high-dimensional features to 1 through a global maximum pooling layer, and then serially fuses high-dimensional features from different domains, and finally uses two convolution layers to predict the results; The improved Vanilla module layer consists of an improved convolution layer and a normalization layer, an activation function, an improved convolution layer and a normalization layer; In step S3, the steps of generating high-dimensional features through improved Vanilla module layer extraction include: 1): Using the activation function A( x ) Two improved convolutional layers are trained simultaneously. The improved convolutional layers use Omni-Dimensional dynamic convolution to replace ordinary convolution. When the training reaches the identity mapping, the outputs of the two convolutional layers are the same and then merged. The identity mapping calculation formula is as follows: A′(x)=(1-λ)A(x)+λx Among them, A′(x) represents the modified activation function, λ represents the nonlinear hyperparameter, e and E represent the current round number and the total number of training rounds, and λ=e / E. In the initial stage of training, e=0, A′(x)=A(x), at which time the network has strong nonlinearity. When the training converges, A′(x)=x, at which time there is no activation function between the two convolutional layers, and there is no nonlinearity. 2): Combine the two improved convolutional layers and the normalization layer into a 1×1 convolution. and Represented as having C in Input channels, C out The weight matrix and bias matrix with output channels and kernel size k, the size, bias, mean and variance in the normalization layer are expressed as The weight matrix and bias matrix after the improved convolution layer and normalization layer are combined are as follows: where the subscript i∈{1,2...,C out } represents the value in the i-th output channel; 3): Merge the two 1×1 convolutions obtained in step 2), and As input and output features, the convolution formula is as follows: y=Q*x=Q·im 2col(x)=Q·X Among them, * represents convolution operation, · represents matrix multiplication, It is derived from the im 2col(x) operation. The weight matrices of the two convolutional layers are represented as Q 1 and Q 2 , for 1×1 convolution, the im2col(x) operation does not require overlapping sliding kernels, so the formula for merging two 1×1 convolutions is as follows: y=Q 1 *(Q 2 *x)=Q 1 ·Q 2 ·im2col(x)=(Q 1 ·Q 2 )*X; 4): During the training process, the nonlinearity of each activation layer is increased by concurrently stacking activation functions. The stacked activation function is expressed as the following formula: Where x represents the input, A(x) represents a single activation function, n represents the number of activation functions, and a i , b i is the size and bias of each activation, given an input feature tensor Where H, W and C represent its width, height and number of channels respectively, and the activation function is expressed as: Among them, h∈{1,2,...,H},w∈{1,2,...,W}, and c∈{1,2,...,C}, when n=0, the activation function A s (x) degenerates into a pure activation function A(x).
2. According to the lightweight electromyography gesture recognition method based on multi-domain feature fusion according to claim 1, it is characterized in that: Input datasets are constructed in S1, including the DB2 and DB5 datasets in NinaPro and self-made datasets, and transmitted to data preprocessing through the acquisition protocol.
3. A lightweight electromyographic gesture recognition method based on multi-domain feature fusion according to claim 1 or 2, characterized in that: The preprocessing operation of the data in the input data set comprises the following steps: Step 1: The input data is processed in sequence by high-frequency denoising, full-wave rectification, data smoothing, and low-frequency denoising; Step 2: The data obtained in step 1 is first divided into actions using the sliding window method, and then the actions corresponding to each data label are segmented and reorganized. Finally, the corresponding actions within the same label are merged and stored separately to obtain time domain data; Step 3: Perform variational mode decomposition on the time domain data obtained in step 2, and use the decomposed time-frequency domain data of the first modal signal component and the time domain data obtained in step 2 as multi-domain inputs of the network model; Step 4: Divide the data obtained in step 3 into sEMG segments of equal size; Step 5: Normalize the data obtained in step 4 and then adjust the data input format, and divide the time domain data and time-frequency domain data into training set and test set respectively.
4. According to the lightweight electromyography gesture recognition method based on multi-domain feature fusion according to claim 1, it is characterized in that: In S3, each training process traverses all samples in the training set, first performing forward propagation to calculate the prediction results, then calculating the error according to the loss function, and then performing backpropagation to update the model weights to minimize the loss; During training, the loss function, optimizer, initial optimization rate, training batch size, and number of training rounds are set, and a strategy to prevent overfitting is adopted.
5. The lightweight electromyographic gesture recognition method based on multi-domain feature fusion according to claim 4 is characterized in that: The loss function adopts the cross entropy loss function, the optimizer adopts SGD, the optimization rate is set to 0.001, the training batch size is set to 128, the number of training rounds is 200, and random inactivation, data enhancement and stratified K-fold cross validation are set to reduce overfitting.
6. The lightweight electromyographic gesture recognition method based on multi-domain feature fusion according to claim 1 is characterized in that: In S4, the evaluation indicators include calculation accuracy, model inference time, floating-point calculation amount, model parameter amount, weight file size, recall rate and average accuracy of all categories.
7. A real-time control system for electromyographic gesture recognition, characterized in that: The weight file and model configuration file obtained by the method described in claim 1 are transplanted to the edge device for real-time decoding of electromyographic signals, and the decoded prediction results are transmitted to the intelligent bionic hand. The intelligent bionic hand acts as an actuator and executes the received instructions through multiple degree-of-freedom motors to complete gesture movements.
8. A data caching method, characterized in that: When using a myoelectric gesture recognition real-time control system as described in claim 7 to control the intelligent bionic hand in real time, the cache size of the electromyographic signal data collected once during the real-time control process is set to a plurality of sample numbers, and a plurality of detection results are obtained corresponding to the plurality of samples. By statistically analyzing the detection results, the result with the highest frequency of occurrence is selected as the final myoelectric gesture determination result.
Citation Information
Patent Citations
Myoelectric gesture recognition method based on multi-feature fusion CNN
CN111860410A
Multi-modal progressive hierarchical fusion method for natural gesture recognition
CN116028889A