Method for improving gesture recognition accuracy
By combining the results of hand detection and classification algorithms for verification, and using the hand tracking mechanism to confirm the consistency of the recognition results, the problem of occlusion and recognition errors in the gesture recognition method based on hand landmark is solved, and the recognition accuracy and credibility are achieved.
Patent Information
- Application Number
- CN202311519831.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
The hand gesture recognition method based on landmark is likely to cause itself to obstruct due to hand posture, resulting in inaccurate prediction of landmark points, which in turn leads to incorrect recognition results and lacks a secondary verification process to correct the results.
The detection and hand-plug part-type method is adopted, and the hand position is positioned and the gesture category is credibility scored is performed. The results of hand detection and classification algorithm are used to verify the accuracy of the recognition results, and the consistency of the recognition results is confirmed in continuous frames through the hand tracking mechanism.
The error recognition rate of gesture recognition is effectively reduced, the recognition accuracy is improved, and the credibility of the recognition results is enhanced through multi-frame judgment.
Smart Images

Figure CN120014692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a method for improving gesture recognition accuracy. Background Art
[0002] In the existing technology, with the increasing popularity of AI applications, smart homes and smart cameras have also appeared in the consumer field. With the popularization of smart devices, more and more human-computer interaction needs have been proposed, such as voice interaction, gesture interaction, etc. Gesture recognition is also increasingly used as a basic application in gesture interaction. Gesture recognition based on hand landmarks is prone to inaccurate landmark point prediction due to various self-occlusions caused by the hand, which in turn leads to errors in subsequent recognition results.
[0003] In other words, hand recognition based on hand landmarks is prone to occlusion due to hand posture. The occluded fingers cannot be seen in reality, so landmark prediction is prone to deviation, leading to incorrect judgment and misrecognition in the subsequent hand posture analysis that relies on the relative position relationship of the landmark points. In addition, there is no secondary verification process, which makes it impossible to correct the results.
[0004] In addition, commonly used technical terms include:
[0005] 1. Gesture recognition: Given a picture, identify and analyze whether there is hand information in the picture, and identify what kind of hand gesture it is? And return the detected hand pixel coordinate information and hand recognition results.
[0006] 2. Hand detection algorithm: detect whether there is a human hand in an image, and return the pixel coordinate information of the human hand.
[0007] 3. Hand classification algorithm: detect the hand category in a picture with a hand and return the result.
[0008] 4. Tracking algorithm: Track the target position according to the coordinate information of the hand detection algorithm in 2, and return the unique tracking ID information of the same target. Summary of the invention
[0009] In order to solve the above problems, the purpose of this application is to: adopt detection plus hand classification, the detection algorithm will locate the hand position, and perform credibility scores for different gesture categories. Based on the detection position information, the hand image data is deducted and then the hand classification algorithm is performed. The hand postures are classified and recognized by two algorithms at the same time. When they are all the same gestures, the final recognition result is artificial. At the same time, hand tracking is performed based on the detection position information, and judgment can be made for multiple consecutive frames. If the results of consecutive frames are consistent, the accuracy of the recognition result can be further determined.
[0010] Specifically, the present invention provides a method for improving gesture recognition accuracy, the method comprising the following steps:
[0011] S1. Start: The entire process starts, the device is powered on, and the running program is executed to start the entire business process;
[0012] S2. Get image: Get a frame of image data from the image sensor. You can get a frame of image data by calling the media SDK get_frame interface.
[0013] S3, hand detection: performing hand detection processing on the obtained one frame of image data, wherein the hand detection processing adopts a CNN deep learning detection model to output hand coordinate frame information and gesture categories, and the gesture categories can be increased or decreased according to actual conditions;
[0014] When the deep learning detection model returns a confidence value for a hand, the confidence value is the reliability of identifying the hand, and the number range is 0 to 1. When it is greater than the set threshold, the hand is detected. The threshold is an empirical value and is assumed to be 0.5; if a hand target is detected, the detected position information and category information are saved in detection_result;
[0015] When the confidence value of the hand returned by the detection model is not greater than the set threshold, the process returns to step S2 to wait for the acquisition of the next frame of image data;
[0016] S4, hand classification: according to the hand detection information obtained in step S3, the hand image information is extracted from the whole image, the hand image information is the pixel coordinates of the upper left corner and the lower right corner of the hand, and is extracted from the whole image as the ROI area, the extraction method adopts the memcpy method to copy the corresponding pixel data, and the extracted image is scaled to 224x224 pixels, the scaled image is sent to the hand classification algorithm to output the classification result, and the classification result is saved in class_result;
[0017] S5, hand tracking: tracking is performed according to the hand frame position information obtained in step S3, and the tracking method is not limited as long as it can ensure that the same hand can be identified as the same tracking ID in each frame detection;
[0018] At the same time, determine whether the categories are consistent based on the category result in step S3 and the classification result in step S4?
[0019] If the results are consistent, the hand gesture of the tracked image is determined to be the result and recorded in the output list, which is an array variable storing gesture results; if they are inconsistent, return to step S2 to continue the next frame result recognition;
[0020] When the same hand information is judged as the same hand posture for multiple times in a row, the hand recognition result is output and fed back to the upper-layer business for display output by other application businesses;
[0021] S6. Determine whether to end the process? When the user thinks it is necessary to end the method, the entire process can be ended by calling the end API; if the user thinks it is not necessary to end the method, continue to acquire images and repeat steps S2, S3, S4, and S5.
[0022] In step S3, the gesture categories can be set to include: victory gesture, palm, fist, OK, and non-specified gesture. In this method, five gestures are used: victory gesture, palm, fist, OK, and non-specified gesture.
[0023] In step S4, the hand classification algorithm adopts a CNN deep learning model, inputs a picture and outputs the confidence of the five gesture categories judged in step S2 in the picture. When the confidence is greater than or equal to a set threshold, the classification result is used as the output result of the hand classification. The threshold is an empirical value, assuming it is 0.5.
[0024] In step S5, the tracking method includes using Kalman filtering plus Hungarian matching algorithm to achieve tracking, and the specific implementation process is as follows:
[0025] Perform Hungarian matching of position information based on the hand detection result in step S3 and the previous detection result to calculate which hands have appeared in the previous frame and are recorded as Ft, which hands appear for the first time and are recorded as Fst, and which hands are not detected and are recorded as Fst;
[0026] If it is Ft, put the coordinate frame of the hand into the list to be output and update the hand frame information of the ID in the historical hand queue;
[0027] If it is Fst, a random unique ID information is assigned to the hand, and the coordinate frame of the hand is put into the list to be output;
[0028] If it is Fst, first determine whether the ID hand has exceeded the maximum life cycle. The maximum life cycle is an empirical value and is set to 5 frames in the present invention. If it exceeds, the ID hand will be deleted from the historical hand queue. If it does not exceed, a coordinate frame is predicted through Kalman filtering based on the historical hand frame information as the current coordinate frame result and put into the list to be output to confirm the tracking ID of the target.
[0029] In step S5, the determination method includes:
[0030] If the category results determined by the hand detection and hand classification algorithms in steps S3 and S4 are consistent, the tracked hand gesture is determined to be the result and recorded in the output list as an array variable for storing gesture results.
[0031] In step S5, the same hand information is continuously presented multiple times as an empirical value, assuming that it is set to 3 frames.
[0032] In the step S5, the display output of the other application services includes drawing the results into an image.
[0033] Therefore, the advantages of this application are:
[0034] 1. Propose a solution and implementation process to improve hand recognition in order to reduce the error rate of gesture recognition.
[0035] 2. It is proposed to simultaneously verify and match the category results in the detection results and the results in the classification calculation. If the matches are consistent, the results are considered credible, which greatly reduces the misrecognition rate.
[0036] 3. A hand tracking process is proposed to count the detection results of each frame. If the same result appears multiple times in a row, it is considered to be the final result, which can further increase the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.
[0038] Figure 1 It is a flow chart of the present application method. DETAILED DESCRIPTION
[0039] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail in conjunction with the accompanying drawings.
[0040] This application proposes a method to improve the accuracy of gesture recognition, such as Figure 1 As shown, the following steps are included:
[0041] S1. Start: The entire process starts, the device is powered on, and the running program is executed to start the entire business process;
[0042] S2. Get image: Get a frame of image data from the image sensor. You can get a frame of image data by calling the media SDK get_frame interface.
[0043] S3, hand detection: performing hand detection processing on the obtained one frame of image data, wherein the hand detection processing adopts CNN deep learning detection model to output hand coordinate frame information and gesture category, wherein the gesture category can be increased or decreased according to actual conditions; the gesture category can be set to include: victory gesture, palm, fist, OK, non-specified gesture, and this method adopts the five gestures of victory gesture, palm, fist, OK, and non-specified gesture;
[0044] When the deep learning detection model returns a confidence value for a hand, the confidence value is the reliability of identifying the hand, and the number range is 0 to 1. When it is greater than the set threshold, the hand is detected. The threshold is an empirical value and is assumed to be 0.5; if a hand target is detected, the detected position information and category information are saved in detection_result;
[0045] When the confidence value of the hand returned by the detection model is not greater than the set threshold, the process returns to step S2 to wait for the acquisition of the next frame of image data;
[0046] S4, hand classification: according to the hand detection information obtained in step S3, the hand image information is extracted from the whole image, the hand image information is the pixel coordinates of the upper left corner and the lower right corner of the hand, and is extracted from the whole image as the ROI area, the extraction method adopts the memcpy method to copy the corresponding pixel data, and the extracted image is scaled to 224x224 pixels, the scaled image is sent to the hand classification algorithm, the classification result is output, and the classification result is saved in class_result;
[0047] The hand classification algorithm adopts a CNN deep learning model, inputs an image and outputs the confidence of the five gesture categories judged in step S2 in the image. When the confidence is greater than or equal to a set threshold, the classification result is used as the output result of the hand classification. The threshold is an empirical value, assuming it is 0.5.
[0048] S5, hand tracking: tracking is performed according to the hand frame position information obtained in step S3, and the tracking method is not limited as long as it can ensure that the same hand can be identified as the same tracking ID in each frame detection;
[0049] At the same time, determine whether the categories are consistent based on the category result in step S3 and the classification result in step S4?
[0050] If the results are consistent, the hand gesture of the tracked image is determined to be the result and recorded in the output list, which is an array variable storing the gesture results; if they are inconsistent, return to step S2 to continue the next frame result recognition;
[0051] When the same hand information is judged as the same hand posture for multiple consecutive times, the hand recognition result is output and fed back to the upper-layer business for other application businesses to display and output; the same hand information is judged as the same hand posture for multiple consecutive times, which is an empirical value, assuming that it is set to 3 frames; the display output of other application businesses includes drawing the result into the image;
[0052] In step S5, the tracking method includes using Kalman filtering plus Hungarian matching algorithm to achieve tracking, and the specific implementation process is as follows:
[0053] Perform Hungarian matching of position information based on the hand detection result in step S3 and the previous detection result to calculate which hands have appeared in the previous frame and are recorded as Ft, which hands appear for the first time and are recorded as Fst, and which hands are not detected and are recorded as Fst;
[0054] If it is Ft, put the coordinate frame of the hand into the list to be output and update the hand frame information of the ID in the historical hand queue;
[0055] If it is Fst, a random unique ID information is assigned to the hand, and the coordinate frame of the hand is put into the list to be output;
[0056] If it is Fst, first determine whether the ID hand has exceeded the maximum life cycle. The maximum life cycle is an empirical value and is set to 5 frames in the present invention. If it exceeds, delete the ID hand from the historical hand queue. If it does not exceed, predict a coordinate frame as the current coordinate frame result through Kalman filtering based on the historical hand frame information and put it into the list to be output to confirm the tracking ID of the target.
[0057] In step S5, the determination method includes:
[0058] If the category results determined by the hand detection and hand classification algorithms in steps S3 and S4 are consistent, the tracked hand gesture is determined to be the result and recorded in the output list as an array variable storing gesture results;
[0059] S6. Determine whether to end the process? When the user thinks it is necessary to end the method, the entire process can be ended by calling the end API; if the user thinks it is not necessary to end the method, continue to acquire images and repeat steps S2, S3, S4, and S5.
[0060] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for improving gesture recognition accuracy, characterized in that: The method comprises the following steps: S1. Start: The entire process starts, the device is powered on, and the running program is executed to start the entire business process; S2. Get image: Get a frame of image data from the image sensor. You can get a frame of image data by calling the media SDK get_frame interface. S3, hand detection: performing hand detection processing on the obtained one frame of image data, wherein the hand detection processing adopts a CNN deep learning detection model to output hand coordinate frame information and gesture categories, and the gesture categories can be increased or decreased according to actual conditions; When the deep learning detection model returns a confidence value for a hand, the confidence value is the reliability of identifying the hand, and the number range is 0 to 1. When it is greater than the set threshold, the hand is detected. The threshold is an empirical value and is assumed to be 0.5; if a hand target is detected, the detected position information and category information are saved in detection_result; When the confidence value of the hand returned by the detection model is not greater than the set threshold, the process returns to step S2 to wait for the acquisition of the next frame of image data; S4, hand classification: according to the hand detection information obtained in step S3, the hand image information is extracted from the whole image, the hand image information is the pixel coordinates of the upper left corner and the lower right corner of the hand, and is extracted from the whole image as the ROI area, the extraction method adopts the memcpy method to copy the corresponding pixel data, and the extracted image is scaled to 224x224 pixels, the scaled image is sent to the hand classification algorithm to output the classification result, and the classification result is saved in class_result; S5, hand tracking: tracking is performed according to the hand frame position information obtained in step S3, and the tracking method is not limited as long as it can ensure that the same hand can be identified as the same tracking ID in each frame detection; At the same time, determine whether the categories are consistent based on the category result in step S3 and the classification result in step S4? If the results are consistent, the hand gesture of the tracked image is determined to be the result and recorded in the output list, which is an array variable storing gesture results; if they are inconsistent, return to step S2 to continue the next frame result recognition; When the same hand information is judged as the same hand posture for multiple times in a row, the hand recognition result is output and fed back to the upper-layer business for display output by other application businesses; S6. Determine whether to end the process? When the user thinks it is necessary to end the method, the entire process can be ended by calling the end API; if the user thinks it is not necessary to end the method, continue to acquire images and repeat steps S2, S3, S4, and S5.
2. A method for improving gesture recognition accuracy according to claim 1, characterized in that: In step S3, the gesture categories can be set to include: victory gesture, palm, fist, OK, and non-specified gesture. In this method, five gestures are used: victory gesture, palm, fist, OK, and non-specified gesture.
3. The method for improving gesture recognition accuracy according to claim 2, characterized in that: In step S4, the hand classification algorithm adopts a CNN deep learning model, inputs a picture and outputs the confidence of the five gesture categories judged in step S2 in the picture. When the confidence is greater than or equal to a set threshold, the classification result is used as the output result of the hand classification. The threshold is an empirical value, assuming it is 0.
5.
4. The method for improving gesture recognition accuracy according to claim 1, characterized in that: In step S5, the tracking method includes using Kalman filtering plus Hungarian matching algorithm to achieve tracking, and the specific implementation process is as follows: Perform Hungarian matching of position information based on the hand detection result in step S3 and the previous detection result to calculate which hands have appeared in the previous frame and are recorded as Ft, which hands appear for the first time and are recorded as Fst, and which hands are not detected and are recorded as Fst; If it is Ft, put the coordinate frame of the hand into the list to be output and update the hand frame information of the ID in the historical hand queue; If it is Fst, a random unique ID information is assigned to the hand, and the coordinate frame of the hand is put into the list to be output; If it is Fst, first determine whether the ID hand has exceeded the maximum life cycle. The maximum life cycle is an empirical value and is set to 5 frames in the present invention. If it exceeds, the ID hand will be deleted from the historical hand queue. If it does not exceed, a coordinate frame is predicted through Kalman filtering based on the historical hand frame information as the current coordinate frame result and put into the list to be output to confirm the tracking ID of the target.
5. The method for improving gesture recognition accuracy according to claim 1, characterized in that: In step S5, the determination method includes: If the category results determined by the hand detection and hand classification algorithms in steps S3 and S4 are consistent, the tracked hand gesture is determined to be the result and recorded in the output list as an array variable for storing gesture results.
6. The method for improving gesture recognition accuracy according to claim 1, characterized in that: In step S5, the same hand information is continuously presented multiple times as an empirical value, assuming that it is set to 3 frames.
7. The method for improving gesture recognition accuracy according to claim 1, characterized in that: In the step S5, the display output of the other application services includes drawing the results into an image.
Citation Information
Patent Citations
Gesture control method, gesture control device, gesture control system and storage medium
CN110297545A
Gesture recognition method and device, computer equipment and storage medium
CN112784810A
Static gesture recognition method and device based on deep learning and automobile
CN114359959A