A gesture action recognition method and device, electronic equipment and storage medium

By acquiring and processing electromyographic signal sequences and using convolutional neural networks and recurrent neural networks to generate probability matrices, the problem of the inability to recognize multi-user gestures in real time and continuously in existing technologies has been solved, and high-precision multi-user gesture recognition has been achieved.

CN116541775BActive Publication Date: 2026-04-10BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing electromyography (EMG) signal recognition solutions cannot achieve real-time continuous motion recognition and can only recognize the gestures of a single user, failing to quickly and accurately recognize the gestures of multiple users.

Method used

The electromyography (EMG) signal sequence of the target object is collected within a preset time period. The signal features are detected by convolutional neural network and recurrent neural network to generate a target probability matrix, determine the target gesture action corresponding to each moment, and generate a gesture action sequence.

Benefits of technology

It enables real-time detection of gestures from multiple objects, improving the accuracy of gesture recognition and allowing for continuous recognition of gestures from multiple users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116541775B_ABST
    Figure CN116541775B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a gesture action recognition method and device, electronic equipment and storage medium. Including: collecting at least one target object generated electromyographic signal sequence within a preset time; detecting the electromyographic signal sequence to obtain the target signal characteristics corresponding to the electromyographic signal sequence; calculating the target signal characteristics to obtain a target probability matrix, wherein the target probability matrix is used to represent the probability distribution of a plurality of gesture actions corresponding to each time within the preset time; determining the target gesture action corresponding to each time based on the target probability matrix, and generating the gesture action sequence of the target object within the preset time based on the target gesture action. The present disclosure realizes real-time detection of the gesture actions of multiple objects, and can identify continuous multiple gesture actions for each object, thereby improving the accuracy of gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of signal processing, and particularly relates to a gesture action recognition method and device, electronic equipment and storage medium. BACKGROUND

[0002] Electromyography signal is a muscle activity electrical signal collected from the surface of human skin, which contains rich limb behavior motion information. At present, the existing gesture action recognition scheme using electromyography signal can only recognize single action in a non-real-time manner, and cannot realize real-time continuous action recognition. Meanwhile, the existing scheme can only recognize the action of a single user, and cannot quickly and accurately recognize the gesture actions of multiple users. SUMMARY

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a gesture action recognition method, device, electronic equipment and storage medium.

[0004] According to an aspect of an embodiment of the present disclosure, a gesture action recognition method is provided, comprising:

[0005] collecting an electromyography signal sequence generated by at least one target object within a preset time;

[0006] detecting the electromyography signal sequence to obtain a target signal feature corresponding to the electromyography signal sequence;

[0007] calculating the target signal feature to obtain a target probability matrix, wherein the target probability matrix is used to represent a probability distribution of a plurality of gesture actions corresponding to each time within the preset time;

[0008] determining a target gesture action corresponding to each time based on the target probability matrix, and generating a gesture action sequence of the target object within the preset time based on the target gesture action.

[0009] According to another aspect of an embodiment of the present disclosure, a gesture action recognition device is also provided, comprising:

[0010] a collection module configured to collect an electromyography signal sequence generated by at least one target object within a preset time;

[0011] a detection module configured to detect the electromyography signal sequence to obtain a target signal feature corresponding to the electromyography signal sequence;

[0012] a calculation module configured to calculate the target signal feature to obtain a target probability matrix, wherein the target probability matrix is used to represent a probability distribution of a plurality of gesture actions corresponding to each time within the preset time;

[0013] The decoding module is configured to determine a target gesture action corresponding to each time based on the target probability matrix, and generate a gesture action sequence of the target object within a preset time based on the target gesture action.

[0014] According to another aspect of the embodiments of the present disclosure, a storage medium is also provided, which includes a stored program. When the program is run, the steps of the above method are performed.

[0015] According to another aspect of the embodiments of the present disclosure, an electronic device is also provided, which includes a processor, a communication interface, a memory and a communication bus. The processor, the communication interface and the memory can communicate with each other through the communication bus. The memory is configured to store a computer program. The processor is configured to perform the steps of the above method by running the program stored in the memory.

[0016] The embodiments of the present disclosure also provide a computer program product including instructions, which, when run on a computer, cause the computer to perform the steps of the above method.

[0017] The above technical solutions provided by the embodiments of the present disclosure have the following advantages. The method provided by the embodiments of the present disclosure first collects an electromyographic signal sequence of a target object within a preset time, and calculates signal features of the electromyographic signal sequence to obtain a target probability matrix. Then, the target probability matrix is used to determine a target gesture action of each time within the preset time. Finally, the target gesture action is used to generate a gesture action sequence of the target object. In this way, the gesture actions of multiple objects are detected in real time, and the continuous gesture actions of each object can be recognized, thereby improving the accuracy of gesture recognition. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, together with the description.

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.

[0020] Figure 1 A flowchart of a gesture action recognition method provided by the embodiments of the present disclosure;

[0021] Figure 2 A structural schematic diagram of a feature detection model provided by the embodiments of the present disclosure;

[0022] Figure 3A structural schematic diagram of a motion recognition model provided by an embodiment of the present disclosure;

[0023] Figure 4 A flowchart of a feature detection model training method provided by an embodiment of the present disclosure;

[0024] Figure 5 A flowchart of a motion recognition model training method provided by an embodiment of the present disclosure;

[0025] Figure 6 A block diagram of a gesture motion recognition device provided by an embodiment of the present disclosure;

[0026] Figure 7 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions serve to explain the present disclosure and do not constitute an improper limitation on the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.

[0028] It should be noted that, in this document, relational terms such as“first” and“second”, and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms“comprises”,“comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the phrase“comprising a……” does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0029] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, and the like should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0030] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application program, a server or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.

[0031] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0032] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0033] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0034] The embodiment of the present disclosure provides a gesture action recognition method, device, electronic device and storage medium. The method provided by the embodiment of the present disclosure can be applied to any required electronic device, for example, a server, a terminal or the like, which is not limited here. For the convenience of description, it is referred to as an electronic device hereinafter.

[0035] According to an aspect of the embodiment of the present disclosure, a method embodiment of a gesture action recognition method is provided, Figure 1 As shown in the flowchart of the gesture action recognition method provided by the embodiment of the present disclosure, Figure 1 The method comprises the following steps:

[0036] In step S11, an electromyographic signal sequence generated by at least one target object within a preset time is collected.

[0037] The method provided by the embodiment of the present disclosure is applied to a smart device capable of data processing and communication. The smart device can be a mobile phone, a computer, a smart watch or the like, and the target object can be a user currently wearing an electromyographic signal collection device. Specifically, the smart device can send a collection instruction to the electromyographic signal collection device. The electromyographic signal collection device can collect the electromyographic signal of the target object within the preset time based on the collection instruction, generate an electromyographic signal sequence, and then feed back the electromyographic signal sequence to the smart device.

[0038] It should be noted that in the spinal cord, the nerve center is connected with multiple motor neurons, when the motor center nerve generates a nerve impulse, the muscle fibers controlled by each neuron are excited synchronously, and the sum of the action potentials generated by all the muscle cells controlled by the neuron is called a motor unit action potential (MUAP). Since the nerve impulse generated by the central nerve is a pulse sequence, the motor unit action potential generated in the muscle group is also a pulse sequence in the time scale, and is also called a motor unit action potential time sequence (MUATP). The electrical pulse in the body is conducted to the skin surface through the soft tissue of the human body, and is collected by the electromyographic sensor, that is, the electromyographic signal is obtained.

[0039] When the motor center nervous system intends to realize different gesture postures, the activation of different motor neurons and the number of excitations of motor neurons per unit time are controlled, which is manifested in the surface electromyographic electrode as different action potential sequences. Therefore, the muscle contraction mode and muscle contraction strength information are implied in the electromyographic signal. By analyzing the signal characteristics of the electromyographic signal sequence, a series of gesture actions of the user within a preset time can be obtained.

[0040] Step S12, detecting the electromyographic signal sequence to obtain a target signal feature corresponding to the electromyographic signal sequence.

[0041] In the embodiment of the present disclosure, detecting the electromyographic signal sequence to obtain a target signal feature corresponding to the electromyographic signal sequence includes the following steps A1-A3:

[0042] Step A1, obtaining a pre-trained feature detection model, wherein the feature detection model includes a convolutional neural network and a recurrent neural network.

[0043] In the embodiment of the present disclosure, a structure diagram of the pre-trained feature detection model is as shown in Figure 2 The feature detection model includes a convolutional neural network and a recurrent neural network. The convolutional neural network can be a causal convolutional neural network (CCN). The causal convolutional neural network is mainly used to capture the dependency relationship between the electromyographic signals. The recurrent neural network can be a gated recurrent unit (GRU). By combining the causal convolutional neural network and the recurrent neural network, not only the dependency relationship between the electromyographic signals in the spatial dimension can be determined, but also the change trend of the electromyographic signals can be determined.

[0044] Step A2, extracting a plurality of first local signal features with a dependency relationship in the electromyographic signal sequence through the convolutional neural network, and inputting the first local signal features into the recurrent neural network.

[0045] In the embodiment of the present disclosure, the myoelectric signal sequence is input into a convolutional neural network, and the myoelectric signal sequence is subjected to convolution calculation by the convolutional neural network to obtain a plurality of signal features with low sampling frequency and dependent relationship, i.e., first local signal features, which include signal amplitude, signal frequency, etc. Then, the first local signal features are input into a recurrent neural network.

[0046] In step A3, the first local signal features are subjected to convolution calculation by the recurrent neural network to obtain target signal features.

[0047] In the embodiment of the present disclosure, the recurrent neural network performs convolution calculation on the first local signal features with dependent relationship to obtain signal features with high sampling frequency, i.e., target signal features. The target signal features are D-dimensional feature vectors.

[0048] It should be noted that the first local signal features can be obtained from at least one frame of myoelectric signal in the myoelectric signal sequence, and the target signal features are signal features corresponding to the entire myoelectric signal sequence.

[0049] In step S13, the target signal features are calculated to obtain a target probability matrix, wherein the target probability matrix is used to represent the probability distribution of a plurality of gesture actions at each time within a preset time.

[0050] In the embodiment of the present disclosure, the target signal features are calculated to obtain a probability matrix, including the following steps B1-B3:

[0051] In step B1, a pre-trained action recognition model is obtained, wherein the action recognition model includes a linear network and a normalization network.

[0052] In the embodiment of the present disclosure, a structure diagram of the pre-trained action recognition model is as shown in Figure 3 The feature detection model includes a linear network and a normalization network, the linear network includes a linear layer and a random inactivation layer, and the normalization network includes a linear layer and a normalization layer.

[0053] In step B2, the target signal features are subjected to linear calculation by the linear network to obtain a first linear feature vector, and the first linear feature vector is input into the normalization network.

[0054] In the embodiment of the present disclosure, the target signal features are input into the linear network, the target signal features are calculated by the linear layer in the linear network to obtain a first calculation result, then the first calculation result is calculated by using a random inactivation function to obtain a first linear feature vector, and the obtained first linear feature vector is input into the normalization network.

[0055] Step B3, normalizing the first linear feature vector through the normalization network to obtain a target probability matrix.

[0056] In the embodiments of the present disclosure, the first linear feature vector is calculated through the linear layer in the normalization network to obtain a second calculation result, and the second calculation result is normalized through the normalization layer to obtain a target matrix probability, the target matrix probability being used to represent a probability distribution of the plurality of gesture actions corresponding to each time within the preset time.

[0057] Step S14, determining the target gesture action corresponding to each time based on the target probability matrix, and generating a gesture action sequence of the target object within the preset time based on the target gesture action.

[0058] In the embodiments of the present disclosure, the target gesture action corresponding to each time is determined based on the target probability matrix, and the gesture action sequence of the target object within the preset time is generated based on the target gesture action, including the following steps C1-C3:

[0059] Step C1, determining the maximum probability corresponding to each time based on the target probability matrix, and determining the gesture action corresponding to the maximum probability as the target gesture action.

[0060] Step C2, detecting whether the target gesture actions corresponding to adjacent times within the preset time are the same.

[0061] Step C3, merging the adjacent times with the same target gesture action to obtain the gesture action sequence.

[0062] In the embodiments of the present disclosure, since the target probability matrix represents the probability distribution of the plurality of gesture actions corresponding to each time within the preset time, the maximum probability corresponding to each time can be obtained from the target probability matrix, and the gesture action corresponding to the maximum probability is determined as the target gesture action.

[0063] As an example: the target probability matrix currently represents that the gesture action corresponding to time T1 is wrist extension—0.85, pointing—0.12. The gesture action corresponding to time T2 has wrist extension—0.80, pointing—0.17. The gesture action corresponding to time T3 has fist—0.81, spherical grip—0.13. The gesture action corresponding to time T4 has wrist flexion—0.75, palm extension—0.22. At this time, it can be determined that the target gesture action corresponding to time T1 is wrist extension, the target gesture action corresponding to time T2 is wrist extension, the target gesture action corresponding to time T3 is fist, and the target gesture action corresponding to time T4 is wrist flexion.

[0064] In the embodiment of the present disclosure, whether the target gesture actions corresponding to adjacent time points in the preset time are the same is detected, and if the target gesture actions corresponding to adjacent time points are the same, the adjacent time points are merged. For example, the target gesture actions corresponding to time point T1 and time point T2 are both wrist stretching, and thus the time point T1 and the time point T2 can be merged to obtain the final gesture action sequence.

[0065] In addition, the electromyographic signals indicate that some time points are in a resting state (i.e., some time points do not generate gesture actions), resulting in discontinuous gesture actions in the preset time period. Based on this, the time points in the resting state can be discarded to obtain the final gesture action sequence. For example, the target gesture action corresponding to time point T1 is wrist stretching, the target gesture action corresponding to time point T2 is wrist stretching, time points T3 and T4 are in a resting state, and the target gesture action corresponding to time point T5 is wrist flexion. At this time, the time point T3 and the time point T4 are discarded.

[0066] The method provided by the embodiment of the present disclosure first collects the electromyographic signal sequence of the target object in the preset time, calculates the signal features of the electromyographic signal sequence, and obtains a target probability matrix. Second, the target gesture action of each time point in the preset time is determined by using the probability matrix. Finally, the gesture action sequence of the target object is generated by using the target gesture action. In this way, the gesture actions of multiple objects are detected at the same time, and continuous multiple gesture actions of each object can be recognized, thereby improving the accuracy of gesture recognition.

[0067] In another embodiment of the present disclosure, a training method of a feature detection model is also provided, Figure 4 The flowchart of the training method of the feature detection model provided by the embodiment of the present disclosure is shown in Figure 4 The method comprises the following steps:

[0068] In step S21, an electromyographic signal sequence sample and an initial feature detection model to be trained are obtained, wherein the initial feature detection model comprises a convolutional neural network and a recurrent neural network.

[0069] In the embodiment of the present disclosure, the initial feature detection model comprises a convolutional neural network and a recurrent neural network. The convolutional neural network can be a causal convolutional neural network (TCN). The causal convolutional neural network is mainly used to capture the dependency relationship between the electromyographic signals. The recurrent neural network can be a Gated Recurrent Unit (GRU). By combining the causal convolutional neural network and the recurrent neural network, the dependency relationship between the electromyographic signals in the spatial dimension can be determined, and the change trend of the electromyographic signals can also be determined.

[0070] Step S22, a plurality of second local signal features existing in the dependence relationship in the electromyographic signal sequence sample are extracted by using the convolutional neural network, and the third local signal feature is input into the recurrent neural network, wherein the third local signal feature is obtained by performing the masking processing on the second local signal feature.

[0071] In the embodiment of the present disclosure, the electromyographic signal sequence sample is input into the convolutional neural network, and the convolutional neural network is used to perform convolution calculation on the electromyographic signal sequence sample to obtain a plurality of signal features with low sampling frequency and existing dependence relationship, i.e., second local signal features, which include signal amplitude, signal frequency, etc. Then, the second local signal features are input into the recurrent neural network.

[0072] It should be noted that the purpose of the masking processing on the second local signal features is to train the model using the incomplete second local signal features, so that the model can still ensure the accuracy of the generated electromyographic signal sequence under the condition that the local signal features are incomplete.

[0073] Step S23, the initial signal feature is output by the recurrent neural network based on the third local signal feature.

[0074] In the embodiment of the present disclosure, the recurrent neural network performs convolution calculation on the second local signal features with dependence relationship to obtain signal features with high sampling frequency, i.e., initial signal features. The initial signal features are D-dimensional feature vectors.

[0075] Step S24, the fourth signal feature obtained by quantizing the second local signal feature is acquired.

[0076] In the embodiment of the present disclosure, the purpose of quantizing the second local signal feature is to discretize the continuous signal feature, and the loss of the feature detection model in the training process is calculated through the discretized feature (i.e., the fourth signal feature).

[0077] Step S25, the first loss between the initial signal feature and the fourth signal feature is determined, and the initial feature detection model is corrected by using the first loss until the trained feature detection model is obtained.

[0078] In the embodiment of the present disclosure, the first loss between the initial signal feature and the fourth signal feature is used to adjust the weight of the initial feature detection model by using the first loss, and then the initial feature detection model with the adjusted weight is continuously trained by using the above process until the first loss is less than the first preset threshold.

[0079] In another embodiment of the present disclosure, a training method of a motion recognition model is also provided, Figure 5 The flowchart of the training method of the motion recognition model provided in the embodiment of the present disclosure is as follows:Figure 5 The method comprises the following steps:

[0080] In step S31, an initial action recognition model to be trained is acquired, wherein the initial action recognition model comprises a linear network and a normalization network.

[0081] In the embodiment of the present disclosure, the initial feature detection model comprises a linear network and a normalization network, the linear network comprises a linear layer and a random inactivation layer, and the normalization network comprises a linear layer and a normalization layer.

[0082] In step S32, an electromyogram sequence sample and a sample label corresponding to the electromyogram sequence sample are acquired.

[0083] In the embodiment of the present disclosure, the electromyogram sequence sample and the sample label corresponding to the electromyogram sequence sample are acquired, comprising the following steps D1-D3:

[0084] In step D1, a plurality of first electromyogram samples are acquired from a first sample set, and a first sample label corresponding to the first electromyogram sample is acquired.

[0085] In the embodiment of the present disclosure, the first sample set is a pre-set sample set, a plurality of candidate electromyogram samples are included in the first sample set, N (N>1) candidate electromyogram samples are randomly selected from the first sample set as the first electromyogram samples, then the first electromyogram samples are detected to obtain a time period in which a gesture action exists in the first electromyogram samples, and the gesture action is added to the first electromyogram samples as the first sample label. It should be noted that one candidate electromyogram sample corresponds to one gesture action.

[0086] In step D2, a plurality of second electromyogram samples are acquired from a second sample set, and a second sample label corresponding to the second electromyogram sample is acquired, wherein a gesture action in the second electromyogram sample is different from a gesture action in the first electromyogram sample.

[0087] In the embodiment of the present disclosure, the second sample set is a pre-set false alarm sample set, a plurality of false alarm electromyogram samples are included in the second sample set, and it can be understood that the false alarm electromyogram sample is an electromyogram sample without detecting a gesture action. M (M>1) false alarm electromyogram samples are randomly selected from the second sample set as the second electromyogram samples. In addition, since the second electromyogram sample is a false alarm electromyogram sample, the corresponding second sample label is empty.

[0088] It should be noted that by setting the false alarm sample set, the second electromyogram samples selected from the false alarm sample set are spliced with the first electromyogram samples, so as to facilitate reducing the false alarm rate of the model in the subsequent training model process.

[0089] Step D3, splicing the first electromyogram signal sample and the second electromyogram signal sample to obtain an electromyogram signal sequence sample, and splicing the first sample label and the second sample label to obtain a sample label.

[0090] In the embodiments of the present disclosure, the purpose of splicing the plurality of first electromyogram signal samples and the second electromyogram signal samples is to obtain the electromyogram signal sequence sample, which is convenient for subsequent training using the electromyogram signal sequence sample. Meanwhile, the first sample label and the second sample label are spliced to obtain the sample label, which is convenient for subsequent calculation of the loss of the action recognition model in the training process.

[0091] It should be noted that the electromyogram signal sequence sample is obtained by splicing the first electromyogram signal sample and the second electromyogram signal sample, and the model is trained by using the electromyogram signal sequence sample. On the one hand, it can reduce the false positive rate of the model in identifying multiple gesture actions. On the other hand, through the splicing method, the training sample can be quickly modified in real time without the need to construct a new training sample.

[0092] Step S33, detecting the electromyogram signal sequence sample to obtain a signal feature corresponding to the electromyogram signal sequence sample, and performing linear calculation on the signal feature of the electromyogram signal sequence sample through a linear network to obtain a second linear feature vector, and inputting the second linear feature vector into a normalization network.

[0093] In the embodiments of the present disclosure, first, the electromyogram signal sequence sample is input into the trained feature detection model to obtain the signal feature of the electromyogram signal sequence sample. Second, the signal feature of the electromyogram signal sequence sample is input into the linear network, and the signal feature is calculated through the linear layer in the linear network to obtain a third calculation result. Then, the third calculation result is calculated by using a random deactivation function to obtain a second linear feature vector, and the obtained second linear feature vector is input into the normalization network.

[0094] Step S34, normalizing the second linear feature vector through the normalization network to obtain an initial probability matrix.

[0095] In the embodiments of the present disclosure, the second linear feature vector is calculated through the linear layer in the normalization network to obtain a fourth calculation result, and the fourth calculation result is normalized through the normalization layer to obtain an initial matrix probability. The initial matrix probability is used to represent the probability distribution of the plurality of gesture actions corresponding to each time.

[0096] Step S35, determining a second loss between the initial probability matrix and the sample label, and correcting the initial action recognition model by using the second loss until a trained action recognition model is obtained.

[0097] In the embodiment of the present disclosure, the second loss between the initial probability matrix and the sample label is utilized to adjust the weight of the initial action recognition model, and then the initial action recognition model with the adjusted weight is continuously trained by using the above process until the second loss is less than the second preset threshold.

[0098] The training method provided by the embodiment of the present disclosure adopts the splicing method to construct the electromyographic signal sequence sample in the face of the increasing demand for new gesture action recognition, realizes the rapid customization of the training sample, improves the training efficiency of the model, and does not need to retrain the model.

[0099] Figure 6 A block diagram of a gesture action recognition device provided by the embodiment of the present disclosure is provided. The device can be realized as part or all of an electronic device by software, hardware or a combination of both. As shown in the figure, the device comprises: Figure 6

[0100] The acquisition module 61 is configured to acquire an electromyographic signal sequence generated by at least one target object within a preset time.

[0101] The detection module 62 is configured to detect the electromyographic signal sequence to obtain a target signal feature corresponding to the electromyographic signal sequence.

[0102] The calculation module 63 is configured to calculate the target signal feature to obtain a target probability matrix, wherein the target probability matrix is used to represent a probability distribution of a plurality of gesture actions corresponding to each time within the preset time.

[0103] The decoding module 64 is configured to determine a target gesture action corresponding to each time based on the target probability matrix, and generate a gesture action sequence of the target object within the preset time based on the target gesture action.

[0104] In the embodiment of the present disclosure, the detection module 62 is configured to obtain a pre-trained feature detection model, wherein the feature detection model comprises a convolutional neural network and a recurrent neural network; the convolutional neural network is used to extract a plurality of first local signal features with a dependent relationship in the electromyographic signal sequence, and the first local signal features are input into the recurrent neural network; the recurrent neural network is used to perform convolution calculation based on the plurality of first local signal features to obtain the target signal feature.

[0105] In the embodiment of the present disclosure, the calculation module 63 is configured to obtain a pre-trained action recognition model, wherein the action recognition model comprises a linear network and a normalization network; the linear network is used to perform linear calculation on the target signal feature to obtain a first linear feature vector, and the first linear feature vector is input into the normalization network; the normalization network is used to perform normalization on the first linear feature vector to obtain the target probability matrix.

[0106] ​In the embodiment of the present disclosure, the decoding module 64 is configured to determine the maximum probability corresponding to each time based on the target probability matrix, and determine the gesture action corresponding to the maximum probability as the target gesture action; detect whether the target gesture actions corresponding to adjacent times within a preset time are the same; and merge the adjacent times with the same target gesture action to obtain the gesture action sequence.

[0107] In the embodiment of the present disclosure, the device further comprises a first training module configured to obtain an electromyographic signal sequence sample and an initial feature detection model to be trained, wherein the initial feature detection model comprises a convolutional neural network and a recurrent neural network; extract a plurality of second local signal features with a dependency relationship in the electromyographic signal sequence by using the convolutional neural network, and input a third local signal feature into the recurrent neural network, wherein the third local signal feature is obtained by performing a masking process on the second local signal feature; output an initial signal feature based on the third local signal feature by using the recurrent neural network; obtain a fourth signal feature obtained by quantizing the second local signal feature; determine a first loss between the initial signal feature and the fourth signal feature, and correct the initial feature detection model by using the first loss until a trained feature detection model is obtained.

[0108] In the embodiment of the present disclosure, the device further comprises a first training module configured to obtain an initial action recognition model to be trained, wherein the initial action recognition model comprises a linear network and a normalization network; obtain an electromyographic signal sequence sample and a sample label corresponding to the electromyographic signal sequence sample; detect the electromyographic signal sequence sample to obtain a signal feature corresponding to the electromyographic signal sequence sample, and perform linear calculation on the signal feature of the electromyographic signal sequence sample by using the linear network to obtain a second linear feature vector, and input the second linear feature vector into the normalization network; perform normalization on the second linear feature vector by using the normalization network to obtain an initial probability matrix; determine a second loss between the initial probability matrix and the sample label, and correct the initial action recognition model by using the second loss until a trained action recognition model is obtained.

[0109] In the embodiment of the present disclosure, the second training module is further configured to obtain a plurality of first electromyographic signal samples from the first sample set, and obtain a first sample label corresponding to the first electromyographic signal samples; obtain a plurality of second electromyographic signal samples from the second sample set, and obtain a second sample label corresponding to the second electromyographic signal samples, wherein the gesture actions in the second electromyographic signal samples are different from the gesture actions in the first electromyographic signal samples; splice the first electromyographic signal samples and the second electromyographic signal samples to obtain an electromyographic signal sequence sample, and splice the first sample label and the second sample label to obtain a sample label.

[0110] The embodiment of the present disclosure also provides an electronic device, such as Figure 7As shown, the electronic device can include a processor 1501, a communication interface 1502, a memory 1503 and a communication bus 1504, wherein the processor 1501, the communication interface 1502 and the memory 1503 complete communication with each other through the communication bus 1504.

[0111] The memory 1503 is used to store computer programs.

[0112] The processor 1501 is used to execute the computer programs stored in the memory 1503 to realize the steps of the above embodiments.

[0113] The communication bus mentioned in the above terminal can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0114] The communication interface is used for communication between the above terminal and other devices.

[0115] The memory can include a random access memory (RAM) and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.

[0116] The processor mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0117] In yet another embodiment provided by the present disclosure, a computer readable storage medium is also provided, which stores instructions, when executed on a computer, cause the computer to perform the gesture action recognition method according to any one of the above embodiments.

[0118] In yet another embodiment provided by the present disclosure, a computer program product containing instructions, when executed on a computer, cause the computer to perform the gesture action recognition method according to any one of the above embodiments.

[0119] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present disclosure are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk) and the like.

[0120] The above only describes preferred embodiments of the present disclosure and is not intended to limit the protection scope of the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

[0121] The above only describes specific embodiments of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied herein.

Claims

1. A method for recognizing hand gestures, characterized in that, include: Collect electromyographic signal sequences generated by at least one target object within a preset time period; The electromyographic signal sequence is detected to obtain the target signal features corresponding to the electromyographic signal sequence; The target probability matrix is ​​obtained by calculating the target signal features, wherein the target probability matrix is ​​used to represent the probability distribution of multiple gesture actions at each moment within the preset time period; The target gesture action corresponding to each moment is determined based on the target probability matrix, and the gesture action sequence of the target object within a preset time is generated based on the target gesture action. The step of determining the target gesture action corresponding to each moment based on the target probability matrix and generating a gesture action sequence of the target object within a preset time period based on the target gesture action includes: determining the maximum probability corresponding to each moment based on the target probability matrix and determining the gesture action corresponding to the maximum probability as the target gesture action; detecting whether the target gesture actions corresponding to adjacent moments within the preset time period are the same; merging adjacent moments with the same target gesture action and removing the moments corresponding to the resting state to obtain a gesture action sequence, wherein the moments corresponding to the resting state are the moments when no gesture action is generated.

2. The method according to claim 1, characterized in that, The step of detecting the electromyographic signal sequence to obtain the target signal features corresponding to the electromyographic signal sequence includes: Obtain a pre-trained feature detection model, wherein the feature detection model includes: a convolutional neural network and a recurrent neural network; The convolutional neural network extracts multiple first local signal features that are dependent on each other in the electromyographic signal sequence, and inputs the first local signal features into the recurrent neural network. The target signal features are obtained by performing convolution calculations based on multiple first local signal features using the recurrent neural network.

3. The method according to claim 1, characterized in that, The step of calculating the target probability matrix from the target signal features includes: Obtain a pre-trained action recognition model, wherein the action recognition model includes: a linear network and a normalized network; The target signal features are linearly calculated using the linear network to obtain a first linear feature vector, and the first linear feature vector is then input into the normalization network. The target probability matrix is ​​obtained by normalizing the first linear feature vector through the normalization network.

4. The method according to claim 2, characterized in that, The training method for the feature detection model includes: Acquire electromyographic signal sequence samples and an initial feature detection model to be trained, wherein the initial feature detection model includes a convolutional neural network and a recurrent neural network; A convolutional neural network is used to extract multiple second local signal features that have a dependency relationship in the electromyographic signal sequence sample, and a third local signal feature is input into the recurrent neural network, wherein the third local signal feature is obtained by masking the second local signal features; The recurrent neural network outputs initial signal features based on the third local signal features; Obtain the fourth signal feature obtained by quantizing the second local signal feature; A first loss is determined between the initial signal features and the fourth signal features, and the initial feature detection model is corrected using the first loss until a trained feature detection model is obtained.

5. The method according to claim 3, characterized in that, The training method for the action recognition model includes: Obtain an initial action recognition model to be trained, wherein the initial action recognition model includes: a linear network and a normalized network; Obtain electromyographic signal sequence samples and the corresponding sample labels for the electromyographic signal sequence samples; The electromyography (EMG) signal sequence sample is detected to obtain the signal features corresponding to the EMG signal sequence sample. The signal features of the EMG signal sequence sample are linearly calculated through the linear network to obtain a second linear feature vector. The second linear feature vector is then input into the normalization network. The second linear eigenvector is normalized using the normalization network to obtain the initial probability matrix; A second loss is determined between the initial probability matrix and the sample labels, and the initial action recognition model is corrected using the second loss until a trained action recognition model is obtained.

6. The method according to claim 5, characterized in that, The acquisition of electromyographic signal sequence samples and corresponding sample labels includes: Multiple first electromyographic signal samples are obtained from the first sample set, and the first sample label corresponding to the first electromyographic signal sample is obtained; Multiple second electromyographic signal samples are obtained from the second sample set, and the second sample labels corresponding to the second electromyographic signal samples are obtained, wherein the hand gestures in the second electromyographic signal samples are different from the hand gestures in the first electromyographic signal samples. The first electromyography (EMG) signal sample and the second EMG signal sample are spliced ​​together to obtain the EMG signal sequence sample, and the first sample label and the second sample label are spliced ​​together to obtain the sample label.

7. A gesture recognition device, characterized in that, include: The acquisition module is used to acquire the electromyographic signal sequence generated by at least one target object within a preset time. The detection module is used to detect the electromyographic signal sequence and obtain the target signal features corresponding to the electromyographic signal sequence; The calculation module is used to calculate the target signal features to obtain a target probability matrix, wherein the target probability matrix is ​​used to represent the probability distribution of multiple gesture actions at each moment within the preset time period; The decoding module is used to determine the target gesture action corresponding to each moment based on the target probability matrix, and to generate a sequence of gesture actions of the target object within a preset time based on the target gesture action; The decoding module is used to determine the maximum probability corresponding to each moment based on the target probability matrix, and determine the gesture action corresponding to the maximum probability as the target gesture action; detect whether the target gesture actions corresponding to adjacent moments within the preset time are the same; merge the adjacent moments with the same target gesture action, and remove the moments corresponding to the resting state to obtain a gesture action sequence, wherein the moments corresponding to the resting state are the moments when no gesture action is generated.

8. A storage medium, characterized in that, The storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 6 when it is run.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other through the communication bus; wherein: Memory, used to store computer programs; A processor for performing the method of any one of claims 1 to 6 by running a program stored in memory.

Citation Information

Patent Citations

  • Gesture information processing method and device, electronic equipment and storage medium

    CN111209885A

  • Electromyographic gesture recognition method based on multi-stream convolutional neural network

    CN111898526A