A fast adaptation method for surface electromyography signal gesture recognition

By performing self-supervised pre-training and calibration of small data of target users on label-free data sets, the problem of high data acquisition and labeling costs in the prior art is solved, and the rapid adaptation and efficient application of surface electromyography signal gesture recognition is achieved.

CN114638258BActive Publication Date: 2025-05-09FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210183516.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-09
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

The existing surface electromyography signal gesture recognition technology requires a large amount of labeled signal data for pre-training, resulting in high data acquisition and labeling costs, hindering the promotion of technology in practical applications.

Method used

A fast adaptation method is adopted to perform offline self-supervised pre-training on the label-free data set, and a surface electromyography signal sample is generated using the sliding window method, and a comparison learning network is constructed for pre-training. Then, only a small amount of labeled signal data of the target user is needed to be collected for calibration, the feature extractor parameters are frozen, and the classification network is trained to obtain the trained gesture recognition network.

Benefits of technology

It realizes the rapid completion of calibration with only a small number of samples from the target user, reduces labor costs for designers and users, and improves the application efficiency and accuracy of surface electromyography signal gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638258B_ABST
    Figure CN114638258B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast adaptation method for surface electromyography signal gesture recognition; the method includes three stages: offline pre-training, online calibration and online application. In the offline pre-training stage, a large amount of unlabeled surface electromyography signal data is collected, and positive and negative sample pairs of surface electromyography signals are constructed to train a feature extractor network. After the pre-training is completed, a classification network is constructed using the trained feature extractor, and then a small amount of surface electromyography signals of a target user are collected online for calibration, and a gesture recognition network is obtained after the calibration is completed. In the online application stage, the gesture recognition network is used to perform gesture recognition on real-time samples to obtain corresponding gesture labels, and the recognized gesture results are output. The method of the present invention can make full use of existing data, and can quickly complete calibration when only a small amount of samples of the target user are collected, so as to realize the rapid application and adaptation of the recognition model on the target subject, and further promote the application of surface electromyography signals in practice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of biomedical engineering, artificial intelligence and computer application technology, and relates to surface electromyography signal gesture recognition, and more specifically to a rapid adaptation method based on surface electromyography signal gesture recognition. Background Art

[0002] Surface electromyography (EMG) is a bioelectric signal generated by muscle contraction during exercise. It is collected by electrodes placed on the surface of the skin. Due to its non-invasive and direct characteristics, it is currently widely used in the fields of interpersonal interaction, auxiliary diagnosis, medical rehabilitation, and prosthetic control. Gesture recognition is one of its important contents. Natural interaction with machines is achieved by identifying the gestures corresponding to surface electromyography (EMG). Gesture recognition requires the use of machine learning and deep learning technologies to decode the pattern corresponding to the signal, thereby establishing an association between the signal and the gesture. However, as a weak physiological signal, surface electromyography (EMG) is easily affected by the collection environment and the user's physiological factors, resulting in large signal changes in actual use. It is necessary to re-collect a large amount of user data for adaptation, which brings a huge burden to the user and hinders the application of surface electromyography (EMG) in practice.

[0003] To solve this problem, some methods simulate factors that cause signal changes, such as electrode displacement or detachment, muscle fatigue, etc., by performing data augmentation during offline training. Such methods require a good understanding of the factors that cause changes in surface electromyographic signals, and need to be able to accurately model or simulate the changes caused by these factors. They are highly dependent on the designer's professional knowledge and are difficult to cover all situations. There are also methods that collect data from the user to align the model, such as muscle alignment and muscle source selection. The adaptation process of these methods is not efficient enough and the scenarios they can cope with are also limited.

[0004] With the rapid development of deep learning methods in the field of gesture recognition, adaptive gesture recognition methods have gradually become the main method to solve the adaptation problem. However, most of the existing gesture recognition model adaptation technologies require a large amount of data with gesture labels for pre-training, and then collect a certain amount of data for fine-tuning and adaptation before users use it. Not only is data collection difficult, but data labeling also brings great labor costs. How to solve the problem of requiring a large amount of data standards and data collection is the key to the application of surface electromyography signal gesture recognition. However, there is currently no effective solution. Summary of the invention

[0005] In view of the actual needs of existing gesture recognition applications and the defects of existing adaptation technologies, the present invention provides a new rapid adaptation method for surface electromyography signal gesture recognition, which does not require a large amount of labeled signal data for pre-training, but only requires the collection of a small amount of labeled signal data from the user for rapid adaptation technology. The labor costs of designers and users are reduced respectively from the early and late data collection, thereby promoting the application of gesture recognition based on surface electromyography signals in practice.

[0006] The purpose of the present invention is mainly achieved through the following technical solutions.

[0007] A fast adaptation method for surface electromyography signal gesture recognition comprises the following steps:

[0008] (1) Offline self-supervised pre-training on unlabeled datasets

[0009] Firstly, the hand surface electromyography signal data segments for different purposes collected by the same device are delabeled and unified in format; then, the surface electromyography signal data segments are segmented using a sliding window to generate surface electromyography signal samples; then, surface electromyography signal positive and negative sample pairs are constructed according to the relationship between the surface electromyography signal samples; then, a contrastive learning network including a feature extractor and a projection mapping layer is constructed, and pre-training is performed using a contrastive learning method to update the network parameters;

[0010] (2) Calibration using some labeled data collected from the target user

[0011] First, a small amount of user labeled data is collected, and a surface electromyography signal sample data set for calibration is generated using the sliding window method; then a classification network including a feature extractor and a classification layer is constructed; then the feature extractor parameters are frozen, and the classification network is trained based on the calibrated surface electromyography signal sample data set, and the network parameters are updated to obtain a trained gesture recognition network.

[0012] (3) Real-time collection of surface electromyography data for gesture recognition

[0013] The surface electromyographic signals of the user when making gestures are collected in real time, and the trained gesture recognition network is used to classify and recognize the collected electromyographic signal samples, and finally the gesture recognition results are output.

[0014] In the present invention, in step (1), the surface electromyography signal data segments are unified into a T×V format, where T is the number of signal data segment frames and V is the number of device sampling channels.

[0015] In the present invention, in step (1), positive and negative sample pairs are constructed according to the adjacent position relationship of surface electromyography signal samples in the data segment; a neural network is used as a feature extractor in the contrastive learning network, followed by a projection mapping layer to project the features into a low-dimensional feature space, and the two samples of each sample pair are respectively sent to the contrastive learning network to calculate the similarity to obtain the loss of each sample pair, and then training is performed to update the feature extractor parameters; after the training is completed, the projection mapping layer is discarded and only the trained feature extractor is retained.

[0016] In the present invention, the feature extractor is implemented by a spatiotemporal convolutional neural network composed of 4 spatiotemporal convolution blocks and 1 global average pooling layer; the projection mapping layer is implemented by a two-layer fully connected network.

[0017] In the present invention, in step (2), in order to obtain the output probability value of each gesture, a classification layer is added after the feature extractor. The dimension of the classification layer is the total number of gesture categories, and the features output by the feature extractor are flattened and directly sent to the classification layer.

[0018] In the present invention, in step (3), a majority voting method is used to determine the final result of gesture recognition.

[0019] In the present invention, the training of the contrastive learning network in step (2) and the training of the classification network in step (3) both adopt the gradient descent method and the back propagation algorithm.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] The method of the present invention can make full use of existing data and can quickly complete calibration when only a small number of samples of target users are collected, thereby realizing rapid application and adaptation of the recognition model on the target subjects, and further promoting the application of surface electromyography signals in practice.

[0022] The pre-training data sources of the present invention are diversified and are not limited to the data collected by gesture recognition. Data collected by any task can also be used. Moreover, not only can data collected from the same subject for multiple days be used, but also data collected from different subjects can be used, which can greatly increase the scale of training data and train the model more effectively. In addition, no labels are required for pre-training, and real-time saved data can be used, which can reduce the cost of data annotation on the one hand, and further expand the data set on the other hand.

[0023] The present invention only needs to collect a small amount of labeled data during the calibration phase, and the labeling process can be further implemented in an automated manner, which reduces the threshold for use and greatly improves the user experience. When only a small amount of labeled data is used for calibration, the rapid migration of the model can be achieved to the greatest extent, the recognition ability of the original model is retained, and the recognition effect is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings used in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and ordinary technicians in this field can also obtain other drawings based on these drawings without creative work.

[0025] Figure 1 An overall training flow chart of a fast adaptation method for surface electromyography signal gesture recognition provided by the present invention.

[0026] Figure 2 A schematic diagram of a method for constructing positive and negative sample pairs provided in an embodiment of the present invention.

[0027] Figure 3 A schematic diagram of a pre-training network structure and training method provided in an embodiment of the present invention.

[0028] Figure 4 A schematic diagram of a calibration network structure and training method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0030] The present invention is based on the assumption that the data distribution of surface electromyography signals of the same action is similar and different. When different users make the same gesture or the same user makes the same gesture in different time periods, the movements of the motor units that make up the surface electromyography signal remain basically unchanged, but the physiological factors of the human body and the external environmental factors cause the collected signals to be different. Therefore, the present invention will use pre-training and fine-tuning methods to respectively learn the intrinsic representation of the surface electromyography signal and the personalized representation in specific scenarios. The overall process is as follows: Figure 1 As shown. Experiments show that when using the full amount of calibration data, the recognition accuracy on two public surface signal gesture recognition datasets exceeds the existing methods by up to 13.8%. When only 40% of the calibration data is used for adaptation, the recognition accuracy is still greater than 90%. The specific steps of pre-training and fine-tuning are as follows:

[0031] 1. Obtain hand surface electromyographic signals from different sources collected by the same device, remove the labels and unify them into the format of T×V, where T is the number of signal data segment frames and V is the number of device sampling channels.

[0032] Exemplarily, the data source may be not only gesture recognition data, but also hand posture estimation data and disease diagnosis data. T is the actual number of data segment frames, and V depends on the sampling channel of the device, such as V=8.

[0033] 2. Use the sliding window method with a window size of w1 and a window step size of s1 to segment the surface electromyography signal data segments and generate surface electromyography signal samples.

[0034] Exemplarily, a sliding window with a window size w1=150 ms and a step size s1=150 ms may be selected to segment the samples.

[0035] 3. Use the segmented data to construct positive and negative sample pairs of surface electromyography signals.

[0036] For example, Figure 2 As shown, a feasible construction method is:

[0037] Adjacent samples x in the same data segment i,m and x i,m+1 Constitute a positive sample pair like Figure 2 The sample x from the i-th data segment in i0 and x i1 The positive sample (x i0 , x i1 );

[0038] Non-adjacent samples x in the same data segment i,m and x i,n Constitute a negative sample pair, where n≠m-1,m,m+1, such as Figure 2 The sample x from the i-th data segment in i0 and x i2 The negative sample pairs (x i0 , x i2 );

[0039] Samples x of different data segments i,m and x j,n Constitute a negative sample pair, where i≠j, and data segment i and data segment j can come from different collection sessions or different collection objects, such as Figure 2 The sample x from the i-th data segment in i2 and the sample x from the jth data segment j2 The negative sample pairs (x j0 , x i2 ).

[0040] 4. Select a suitable neural network structure as the feature extractor to extract the surface electromyography signal features in the sample pair, and connect a projection mapping layer in series to project the features into a low-dimensional feature space.

[0041] For example, Figure 3 As shown in Figure 1, a spatiotemporal convolutional neural network is selected as the feature extractor. The network consists of four spatiotemporal convolutional layers, which can realize the joint extraction of temporal and spatial features. The projection mapping layer is implemented by a two-layer fully connected network with dimensions of 512 and 128 respectively. The sample pair (x i0 , x i1 ) will be projected into the 128-dimensional feature space to obtain the representation vector z i0 and z i1 .

[0042] 5. Based on the representation vector z in step 4 i0 and z i1 , calculate the similarity of samples within the positive and negative sample pairs, and calculate the loss based on this similarity.

[0043] Exemplarily, the similarity is calculated using cosine similarity, and the formula is as follows:

[0044] similarity(z i0 , z i1 )=z i0 ·z i1 / (||z i0 ||·||z i1 ||)

[0045] 6. Learn and update the parameters of the feature extractor and projection mapping layer according to the loss function. After training, the projection mapping layer is discarded and only the feature extractor is retained.

[0046] For example, the training process is as follows Figure 3 As shown, the parameter learning and updating method adopts gradient descent method and back propagation algorithm.

[0047] 7. After the user wears the device, guide the user to perform each target gesture p times, each lasting q seconds, and then use the sliding window method with a window size and step size of w2 and s2 respectively to generate samples for calibration.

[0048] Exemplarily, each target gesture may be collected p=1 times, each time lasting q=3 seconds, and the size and step length of the sliding window may be w2=150ms and s2=70ms, respectively.

[0049] 8. Connect the feature extractor trained in step 6 in series with a classification layer of dimension G and freeze the parameters of the feature extractor.

[0050] For example, the network structure is as follows Figure 4As shown in FIG. 1 , the network structure consists of a feature extractor, i.e., the spatiotemporal convolutional neural network in step 4, and a classification layer dimension composed of a fully connected layer with a dimension G in series, where G is the total number of gesture categories. For example, if the total number of gestures is 8, then G = 8.

[0051] 9. Use the surface electromyography signal samples obtained in step 7 to train the network constructed in step 8 and update the network to obtain the final network for gesture recognition.

[0052] For example, the training process is as follows Figure 4 As shown, the parameter learning and updating method adopts gradient descent method and back propagation algorithm.

[0053] 10. Collect L surface signal samples in real time, and use the trained network to recognize the gestures corresponding to the samples in turn. When a gesture recognition result meets certain conditions, output the gesture recognition result.

[0054] For example, the sample number L may be 5, the output condition may be majority voting, and the winning gesture must appear more than or equal to 2 times. When multiple gestures win, the gesture with the highest confidence is selected.

[0055] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by combining software with necessary hardware platforms. Based on such understanding, the technical solutions corresponding to the above embodiments can be embodied in the form of software products, which can be stored in a storage medium and execute the methods of the embodiments of the present invention through instructions or manually.

[0056] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A fast adaptation method for surface electromyography signal gesture recognition, characterized in that: The following steps are involved: (1) Offline self-supervised pre-training on unlabeled datasets Firstly, the hand surface electromyography signal data segments for different purposes collected by the same device are delabeled and formatted in a unified manner; then, the surface electromyography signal data segments are segmented using a sliding window to generate surface electromyography signal samples; then, surface electromyography signal positive and negative sample pairs are constructed based on the relationship between the surface electromyography signal samples; Reconstruct a contrastive learning network including a feature extractor and a projection mapping layer, use the contrastive learning method for pre-training, and update the network parameters; (2) Calibrate using some labeled data collected from the target user First, a small amount of user labeled data is collected, and a surface electromyography signal sample data set for calibration is generated using a sliding window method; then a classification network including a feature extractor and a classification layer is constructed; then the feature extractor parameters are frozen, and the classification network is trained based on the calibrated surface electromyography signal sample data set, and the network parameters are updated to obtain a trained gesture recognition network; (3) Real-time collection of surface electromyography data for gesture recognition The surface electromyographic signals of the user when making gestures are collected in real time, and the trained gesture recognition network is used to classify and recognize the collected electromyographic signal samples, and finally the gesture recognition results are output.

2. The rapid adaptation method according to claim 1, characterized in that: In step (1), the surface electromyography signal data segments are unified into The format, is the number of signal data segment frames, The number of sampling channels of the device.

3. The rapid adaptation method according to claim 1, characterized in that: In step (1), positive and negative sample pairs are constructed according to the adjacent position relationship of the surface electromyography signal samples in the data segment; a neural network is used as a feature extractor in the contrastive learning network, followed by a projection mapping layer to project the features into a low-dimensional feature space, and the two samples of each sample pair are respectively sent to the contrastive learning network to calculate the similarity to obtain the loss of each sample pair, and then training is performed to update the feature extractor parameters; After training is complete, the projection mapping layer is discarded and only the trained feature extractor is retained.

4. The rapid adaptation method according to claim 3, characterized in that: The feature extractor is implemented using a spatiotemporal convolutional neural network consisting of four spatiotemporal convolutional blocks and one global average pooling layer; the projection mapping layer is implemented using a two-layer fully connected network.

5. The rapid adaptation method according to claim 1, characterized in that: In step (2), the output dimension of the classification layer is the total number of gesture categories, and the features output by the feature extractor are flattened and directly sent to the classification layer.

6. The rapid adaptation method according to claim 1, characterized in that: In step (3), the majority voting method is used to determine the final result of gesture recognition.

7. The rapid adaptation method according to claim 1, characterized in that: The training of the contrastive learning network in step (1) and the training of the classification network in step (2) both use the gradient descent method and the back propagation algorithm.

Citation Information

Patent Citations

  • Method and device for detection of sleep apnea fragment based on unsupervised feature learning

    CN110801221A

  • Action mode recognition model updating method and device

    CN111310658A