A Gesture Recognition Method Based on Weighted Feature Enhancement
By adopting weighted feature enhancement technology in the gesture recognition method, the problem of difficult to effectively utilize the importance of sensor features and the local interactiveness of gesture data is solved, and the accuracy and performance of gesture recognition are improved.
Patent Information
- Application Number
- CN202110696669.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing gesture recognition methods are difficult to effectively consider the importance of different sensor features and the local interactivity of gesture data, resulting in insufficient gesture recognition accuracy and performance.
The gesture recognition method based on weighted feature enhancement is adopted to measure the correlation between different sensor features and gestures through information entropy, a weight calculation model is constructed, the importance of different sensor features is calculated, and the local interaction information of gesture data is extracted through the feature enhancement model.
The accuracy and performance of gesture recognition are improved, the ability to characterize local interactive information of different sensor features and gesture data is enhanced, and the convergence speed of the gesture recognition model is improved.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention discloses a gesture recognition method based on weighted feature enhancement, which relates to the field of information technology, and particularly relates to a gesture recognition method based on sensors. Background Art
[0002] In recent years, the rapid development of perception computing, sensor integration technology and immersive interaction technology has made it possible for all-round perception-based human-computer interaction. Gesture recognition has the characteristics of convenience, naturalness and user-friendliness, bringing new opportunities to human-computer interaction. Interacting with gestures can not only retain the original interaction habits of users, but also enrich the connotation and form of human-computer interaction, and gradually become a new research hotspot. In addition, with the development of mobile communication technology, smart wearable devices can provide exclusive and personalized services for people. Common smart wearable devices (such as data gloves) have the advantages of being light, portable and rich in functions, and are equipped with a large number of sensors, becoming an important tool for human body posture perception.
[0003] Gesture data integrating multiple sensors can comprehensively perceive the real gesture movement trajectory and has local interaction characteristics, that is, the changes in the pose information of adjacent joint points affect each other. For example, the change in the pose information of the upper arm will affect the pose information of the lower arm, and the swing of the lower arm will affect the pose information of the palm. At the same time, different gestures have great differences in the sensitivity of each sensor, which means that the degree of association between the characteristics of each sensor and the gesture class label is different. In addition, gesture data can be divided into: upper arm data, lower arm data, palm data and finger data. There are certain differences in the influence degree of each type of data on gestures. For example, American Sign Language mainly focuses on the bending degree of gestures, while military standard sign language mainly focuses on the arm pose information.
[0004] Therefore, in view of the above characteristics of gesture data, a gesture recognition method that fully considers the importance of features and can represent local interaction characteristics is needed. On the one hand, it can comprehensively consider the importance of different sensors to improve the accuracy of gesture recognition; on the other hand, it can extract the local interaction information of gesture data and enhance the performance of gesture recognition. Summary of the Invention
[0005] In order to solve the above problems, the present invention discloses a gesture recognition method based on weighted feature enhancement. This method fully measures the degree of association between different sensor features and class labels, and combines a feature enhancement model to extract the local interaction information of gesture data, thereby improving the accuracy of gesture recognition and enhancing the performance of gesture recognition.
[0006] The present invention is implemented according to the following technical solutions:
[0007] A gesture recognition method based on weighted feature enhancement, the specific process is as follows:
[0008] Step 1: Measure the correlation between different sensor features and gestures through information entropy, construct a weight calculation model, and calculate the importance of different sensor features through sensor correlation analysis and gesture joint point analysis;
[0009] Step 2: Construct a gesture feature enhancement and fusion model, perform feature enhancement and fusion operations on the weighted gesture data feature values obtained in Step 1, and mine and analyze the local interaction information of the gesture data;
[0010] Step 3: Use long short-term memory units to solve the temporal sequence and long-distance dependence problems of dynamic gestures;
[0011] Step 4: Optimize the model training strategy, accelerate the convergence speed of the recognition model, and improve the gesture recognition accuracy.
[0012] For the specific solution, the specific steps of Step 1 are as follows:
[0013] 1.1. Scan the original gesture data in sequence, and calculate the importance of different sensor features through mutual information. The mutual information I(C j , X i ) between the feature vector X i and the gesture label C j is calculated as follows:
[0014]
[0015] where C i is the i-th gesture in the gesture category set C, X j is the j-th dimension feature of the feature vector X, P(C i ) is the probability that the feature vector belongs to the gesture category C i , P(X j ) represents the probability that the feature item X j appears in the dataset, and P(C i , X j ) represents the ratio of the number of samples in which the feature item X i appears in the gesture category C j to the total number of samples in the entire dataset.
[0016] 1.2. Based on the gesture data of wearable sensors, it can be divided according to the spatial position of the sensors into: upper arm data D u (Upper arm data), forearm data D f (Forearm data), palm of hand data D p (Palm of hand data), and finger data D finger(Finger data). The weights of gesture data of each type (also known as joint point data) are the mean of the mutual information of each eigenvector belonging to this type of data, and its calculation method is as follows:
[0017]
[0018] Among them, k represents the feature dimension of each type of joint point data. The value of i is {1, 2, 3, 4}, representing the weights of upper arm data, forearm data, palm data, and finger data respectively. softmax() is an activation function used to normalize the feature weights, so that
[0019] The weighted eigenvalue F obtained through the weight calculation model weighted , is as follows:
[0020]
[0021] 1.3. Repeat steps 1.1 and 1.2 until the gesture data weights of all sampling points are calculated.
[0022] For the specific scheme, the specific steps of step 2 are as follows:
[0023] 2.1. Scan the obtained weighted gesture data eigenvalues in sequence, and use the element-by-element addition of adjacent joint point data to capture the local interaction information of the gesture data and achieve feature enhancement. Specifically:
[0024]
[0025] Among them ⊙ represents the element-by-element addition operator. The value of i is {1, 2, 3, 4}, representing the upper arm weighted feature, forearm weighted feature, palm weighted feature, and finger weighted feature respectively. In addition, there is also local interaction between each finger. Therefore, the finger data and its adjacent finger data need to perform local interaction information enhancement operations.
[0026] 2.2. Through the method of feature fusion, comprehensively represent the real hand posture and motion information, and merge the enhanced feature vectors into more discriminative features. Specifically:
[0027]
[0028] Among them, represents the enhanced gesture feature vector.
[0029] For the specific scheme, the specific steps of step 3 are as follows:
[0030] 3.1. According to the hidden layer state h at the previous moment t-1and the input gesture feature vector m at the current moment t Calculate the forgetting gate information amount, which determines the gesture information amount that the long short-term memory network (LSTM) cell needs to ignore. Specifically:
[0031] f t = σ(W f · [h t-1 , m t + b f )
[0032] Among them, f t is the information amount forgotten at the current moment, σ() is the sigmoid non-linear activation function, the parameter W f is the forgetting weight, h t-1 is the output of the previous moment, m t is the input gesture feature vector at the current moment, and b f represents the forgetting bias amount.
[0033] 3.2. Calculate the input information i of the gesture data t and the alternative state information for update Specifically:
[0034] i t = σ(W i · [h t-1 , m t + b i )
[0035]
[0036] Among them, σ() is the sigmoid non-linear activation function, tanh() is the activation function, the parameter W i is the input weight, h t-1 is the output of the previous moment, m t is the input gesture feature vector at the current moment, b i represents the input bias amount, W c is the state weight, and b c represents the state bias amount.
[0037] 3.3. Calculate the output information o t and the hidden layer state h t , specifically:
[0038] o t = σ(W o [h t-1 , m t + b o )
[0039] h t= o t * tanh(C t )
[0040] where tanh() is the activation function, σ() is the sigmoid non-linear activation function, the parameter W o is the output weight, h t-1 is the output of the previous moment, m t is the gesture feature vector input at the current moment, b o represents the output bias, C t is the memory at the current moment.
[0041] For the specific solution, the specific steps of step 4 are as follows:
[0042] 4.1. When building the recognition model, perform batch normalization on each layer of the network to reduce the sensitivity of the network to the initial weights and improve the convergence speed of the gesture recognition model.
[0043] 4.2. Combine the softmax loss function and the Fisher linear criterion to redefine the loss function during model training, specifically:
[0044] L = L s + θL f
[0045] where L s is the softmax loss function, θ ∈ [0, 1] is used to control the fusion degree of the softmax loss function and the Fisher criterion, L f is the loss function based on the Fisher linear criterion, specifically:
[0046]
[0047] where δ ∈ [1e - 3, 0.1] is the discriminant parameter, n represents the number of gesture samples with label y i , m is the number of gesture types, and μ represents the sample mean.
[0048] Further solution:
[0049] (1) Calculate the mean of the gesture data, specifically:
[0050]
[0051] where μ B is the mean of each batch of gesture data. For each batch of gesture data m represents the size of each batch, D i is the data item in gesture batch B.
[0052] (2) Calculate the variance of the gesture data, specifically:
[0053]
[0054] Among them, δ is the variance of the calculated gesture data, and μ B is the mean value of gesture data for each batch, and D i is the data item in gesture data batch B.
[0055] (3) Normalize the gesture data, specifically:
[0056]
[0057] Among them, is the gesture data after normalization.
[0058] (4) Perform translation and scaling processing, specifically:
[0059]
[0060] Among them, f i is the translation or scaling amount of the gesture, γ and β represent learnable parameters, D i is the data item in gesture batch B, is the gesture data after normalization.
[0061] Advantages of the present invention:
[0062] (1) The present invention fully considers the sensitivity differences of different sensors for gesture acquisition, comprehensively measures the correlation between different sensor features and gestures, decomposes gesture data into upper arm data, forearm data, palm data, and finger data, solves the differences in the importance of different features, and creates a weight calculation model, which can effectively improve the gesture recognition accuracy;
[0063] (2) The present invention fully considers the local spatial interactivity of gesture data, that is, the changes in the pose information of adjacent joint points affect each other, mines the local interaction information of adjacent joint points, realizes the weighted feature representation and feature enhancement of gestures, and enhances the gesture recognition performance;
[0064] (3) The present invention improves the normalization method of each layer of the network by optimizing the batch normalization method, performs translation and scaling processing on the gesture feature data, reduces the sensitivity of the network to the initial weights, and improves the convergence speed of the gesture recognition model;
[0065] (4) The present invention fully considers the weight influence of different parts in the hand pose representation features, and can accurately represent the subtle motion features of different parts of the upper limb online in the virtual display field;
[0066] (5) The present invention fully considers the weight influence of different parts in the hand gesture representation features, can flexibly represent the relationship between the movements of different parts of the upper limb and the gesture features, and can assist in accurately recognizing the upper limb movements of the upper arm disabled persons in the field of medical rehabilitation, and assist in examining the degree of upper limb function recovery;
[0067] (6) By optimizing and improving the network normalization model, the present invention improves the recognition speed of the gesture model. In the field of industrial control, it can not only achieve flexible control of the robotic arm, but also enhance the flexibility of the manipulator; BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The drawings, as a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute an improper limitation to the present invention. Obviously, the drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.
[0069] In the drawings:
[0070] Att Figure 1 is a schematic diagram of the research framework of a gesture recognition method based on weighted feature enhancement disclosed by the present invention.
[0071] Att Figure 2 is a schematic diagram of the feature enhancement fusion model in a gesture recognition method based on weighted feature enhancement disclosed by the present invention.
[0072] Att Figure 3 is a schematic diagram of the long short-term memory unit structure in a gesture recognition method based on weighted feature enhancement disclosed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The present invention discloses a gesture recognition method based on weighted feature enhancement. The following will clearly describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. The following embodiments are helpful for those skilled in the art to better understand the present invention, but do not limit the present invention in any form. It should be noted that all other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0074] Figure 1 is a research framework diagram of a gesture recognition method based on weighted feature enhancement disclosed by the present invention, mainly including the following steps:
[0075] Step 1: Measure the correlation between different sensor features and gestures through information entropy, construct a weight calculation model, and calculate the importance of different sensor features through correlation analysis and joint point analysis.
[0076] 1.1. Scan the original gesture data sequentially, and calculate the importance of different sensor features through mutual information. The formula for mutual information I(C i , X j ) is as follows:
[0077]
[0078] where C i is the i-th gesture in the gesture category set C, X j is the j-th dimension feature of the feature vector X, P(C i ) is the probability that the feature vector belongs to the gesture category C i , P(X j ) represents the probability that the feature item X j appears in the dataset, and P(C i , X j ) represents the ratio of the number of samples with the feature item X i appearing in the gesture category C j to the number of samples in the entire dataset.
[0079] 1.2. Normalize the weighted gesture data feature values. The formula is as follows:
[0080]
[0081] where k represents the feature dimension of various joint point data, and the value of i is {1, 2, 3, 4}, representing the weights of the upper arm data, forearm data, palm data, and finger data respectively. softmax() is an activation function used to normalize the feature weights, so that
[0082] 1.3. Repeat Step 1 and Step 2 until the weights of the gesture data at all sampling points are calculated.
[0083] Step 2: Enhance and fuse the weighted gesture data features obtained in part (1) through a feature enhancement and fusion model, extract the local interaction information of the gesture data, and improve the accuracy of gesture recognition. Its structure diagram is as Figure 2 shown.
[0084] 2.1. Scan the weighted gesture data feature values obtained in part (1) sequentially, and capture the local interaction information of the gesture data by adding adjacent joint point data element by element to achieve feature enhancement. The formula is as follows:
[0085]
[0086] Among them ⊙ represents the symbol of element-wise addition operation. The value range of i is {1, 2, 3, 4}, which respectively represent the weighted feature of the upper arm, the weighted feature of the forearm, the weighted feature of the palm, and the weighted feature of the fingers. In addition, there is also local interaction between each finger. Therefore, the finger data and its adjacent finger data need to perform local interaction information enhancement operation.
[0087] 2.2. Comprehensively represent the real hand gesture and motion information through feature fusion, and merge the enhanced feature vectors into more discriminative features. The calculation formula is as follows:
[0088]
[0089] Among them, represents the enhanced gesture feature vector.
[0090] Step 3. Use the long short-term memory network LSTM to solve the temporal sequence and long-distance dependence problems of dynamic gestures. Its structure diagram is as Figure 3 shown.
[0091] 3.1. Calculate the forgetting gate information amount according to the hidden layer state h t-1 at the previous moment and the input feature vector m t at the current moment, and determine the gesture information amount that the LSTM unit needs to ignore. The calculation formula is as follows:
[0092] f t = σ(W f ·[h t-1 , m t +b f )
[0093] 3.2. Calculate the input information i t of the gesture data and the alternative state information for update. The calculation formula is as follows:
[0094] i t = σ(W i ·[h t-1 , m t +b i )
[0095]
[0096] 3.3. Based on the LSTM unit state information, calculate the output information o t and the hidden layer state h t . The calculation formula is as follows:
[0097] o t = σ(W o [h t-1 ,m t +b o )
[0098] h t = o t *tanh(C t )
[0099] Step 4. During the model training process, the batch normalization algorithm is adopted to reduce the sensitivity of the network to the initial weights; based on the softmax function and combined with the Fisher linear criterion, a loss function is constructed to optimize the gesture recognition model.
[0100] 4.1. When building the recognition model, batch normalization processing is performed on each layer of the network to reduce the sensitivity of the network to the initial weights and improve the convergence speed of the gesture recognition model.
[0101] 4.2. The softmax loss function and the Fisher linear criterion are combined to re - define the loss function during model training. The formula is as follows:
[0102] L = L s + θL f
[0103] where L s is the softmax loss function, θ ∈ [0, 1] is used to control the fusion degree of the softmax loss function and the Fisher criterion, and L f is the loss function based on the Fisher linear criterion. The formula is as follows:
[0104]
[0105] where δ ∈ [1e - 3, 0.1] is the discrimination parameter, n represents the number of gesture samples with label y i , m is the number of gesture types, and μ represents the sample mean.
[0106] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A gesture recognition method based on weighted feature enhancement and fusion, characterized in that: Step 1: Measure the correlation between different sensor features and gestures through information entropy, construct a weight calculation model, and calculate the importance of different sensor features through sensor correlation analysis and gesture joint point analysis; Step 2: Construct a gesture feature enhancement and fusion model, perform feature enhancement and fusion operations on the weighted gesture data feature values obtained in Step 1, and mine and analyze the local interaction information of gesture data; Step 3: Use long short-term memory units to solve the temporal sequence and long-distance dependence problems of dynamic gestures; Step 4: Optimize the model training strategy, accelerate the convergence speed of the recognition model, and improve the accuracy of gesture recognition; Among them, the specific steps of Step 2 are as follows: 2.
1. Sequentially scan the obtained weighted gesture data feature values, and use the method of adding adjacent joint point data element by element to capture the local interaction information of gesture data and achieve feature enhancement. Specifically: , Among them, , represents the element-wise addition operator, i takes values from {1, 2, 3, 4}, representing the weighted feature of the upper arm, the weighted feature of the forearm, the weighted feature of the palm, and the weighted feature of the fingers respectively; in addition, there is also local interaction between each finger. Therefore, the finger data and its adjacent finger data need to perform local interaction information enhancement operations; 2.
2. Comprehensively represent the real hand posture and motion information through feature fusion, and merge the enhanced feature vectors into more discriminative features. Specifically: , Among them, represents the enhanced gesture feature vector; The specific steps of Step 4 are as follows: 4.
1. When building the recognition model, perform batch normalization on each layer of the network to reduce the sensitivity of the network to the initial weights and accelerate the convergence speed of the gesture recognition model; 4.
2. Combine the softmax loss function and the Fisher linear criterion to redefine the loss function during model training. Specifically: , Among them, is the softmax loss function, which is used to control the fusion degree of the softmax loss function and the Fisher criterion, is the loss function based on the Fisher linear criterion, specifically: , where is a discrimination parameter, n represents the number of gesture samples with label y i , m is the number of gesture types, represents the sample mean.
2. The gesture recognition method based on weighted feature enhancement according to claim 1, characterized in that: The specific steps of Step 1 are as follows: 1.
1. Correlation analysis. Specifically: Scan the original gesture data in sequence, and calculate the importance of different sensor features through mutual information. The mutual information between the feature vector Xj and the gesture label Cj is as follows: The calculation formula is as follows: , Among them, Ci is the set of gesture categories C the i th gesture, Xj is the X th j dimensional feature, is the probability that the feature vector belongs to the gesture category Ci, represents the probability that the feature term Xj appears in the dataset, represents the ratio of the number of samples in which the feature term Xj appears in the gesture category Ci to the number of samples in the entire dataset; 1.
2. Joint point analysis. Specifically: Based on the gesture data of wearable sensors, according to the spatial position of the sensors, it can be divided into: upper arm data, forearm data, palm data, and finger data; then the weights of each type of gesture data, that is, joint point data, are the mean of the mutual information of each feature vector belonging to this type of data, and its calculation method is as follows: , Among them, k represents the data feature dimension of various joint points. The value range of i is {1, 2, 3, 4}, representing the weight of upper arm data, the weight of forearm data, the weight of palm data, and the weight of finger data respectively. softmax() is an activation function used to normalize the feature weights, so that ; The weighted eigenvalue obtained by the weight calculation model , is as follows: , Among them, Du represents the data of the upper arm, D f represents the data of the forearm, D p represents the data of the palm, D finger represents the data of the fingers; 1.
3. Repeat Steps 1.1 and 1.2 until the weights of the gesture data at all sampling points are calculated.
3. The gesture recognition method based on weighted feature enhancement according to claim 1, characterized in that: The specific steps of Step 3 are as follows: 3.
1. Calculate the forget gate information based on the hidden layer state h at the previous moment t-1 and the input gesture feature vector at the current moment m t Determine the amount of gesture information that the long short-term memory network LSTM unit needs to ignore. Specifically: , Among them, f t is the amount of information forgotten at the current moment, is the sigmoid non-linear activation function, and the parameter w f is the forgetting weight, h t-1 is the output of the previous moment, m t is the gesture feature vector input at the current moment, b f represents the forgetting bias; 3.
2. Calculate the input information of gesture data i t and alternative status information for update , specifically:[[]] , , Among them, tanh is the activation function, and the parameter W i is the input weight, b i represents the input bias, W c is the state weight, b c represents the state bias; 3.
3. Calculate the output information based on the LSTM cell state information O t and the hidden layer state h t Specifically, , , Among them, the parameter W 0 is the output weight, and b0 represents the output bias C t is the memory at the current moment.
4. The gesture recognition method based on weighted feature enhancement according to claim 1, wherein, The batch normalization algorithm in Step 4.1 is specifically: (1) Calculate the mean of gesture data. Specifically: , Among them, is the mean of the gesture data for each batch. For each batch of gesture data , m represents the size of each batch, D i is the data item in gesture batch B; (2) Calculate the variance of gesture data. Specifically: , Among them, is the variance of the calculated gesture data, is the mean of the gesture data for each batch, is the data item in gesture data batch B; (3) Normalize the gesture data. Specifically: , Among them, is the gesture data after normalization; (4) Perform translation and scaling processing. Specifically: , Among them, is the gesture translation or scaling amount, , represents the learnable parameter, is the data item in the gesture batch B, is the gesture data after normalization.