A gesture recognition method and device, and a storage medium

By performing multiple reduction processing on the image to be processed and hand detection, the problem of insufficient information in gesture recognition is solved, and the accuracy of gesture recognition is improved.

CN114445864BActive Publication Date: 2026-01-16BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210112744.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-29
Publication Date
2026-01-16
Estimated Expiration
2042-01-29

AI Technical Summary

Technical Problem

In existing technologies, static gesture recognition suffers from insufficient information due to the small size of the gesture area within the entire image, thus reducing the accuracy of gesture recognition.

Method used

By performing multiple reduction processing on the image to be processed, multiple reduced images are obtained. Hand detection is then performed to obtain hand detection boxes, thereby determining the hand region image and finally recognizing the target gesture.

Benefits of technology

It improves the accuracy of gesture recognition by increasing the proportion of the gesture in the image, ensuring that enough information is obtained for accurate recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445864B_ABST
    Figure CN114445864B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a gesture recognition method and device, and a storage medium, including: in a case where a to-be-processed image is acquired, processing the to-be-processed image according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images; performing hand detection on the plurality of reduced to-be-processed images to obtain a hand detection frame; acquiring an image in the hand detection frame from the to-be-processed image to obtain a hand region image; and determining a target gesture in the to-be-processed image according to the hand region image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information storage, in particular to a gesture recognition method and device, and a storage medium. BACKGROUND

[0002] Static gesture recognition technology refers to recognizing the posture of a hand at a certain time point through a sensor, such as recognizing a Victory gesture, an OK gesture, and the like of a user.

[0003] In the prior art, when picture information containing a gesture is acquired, the gesture in the picture is directly recognized. When the gesture part occupies a small part of the entire picture area in the acquired picture, the information of the gesture part is very little, thereby reducing the accuracy of gesture recognition. SUMMARY

[0004] To solve the above technical problems, the embodiments of the present application expect to provide a gesture recognition method and device, and a storage medium, which can improve the accuracy of gesture recognition.

[0005] The technical solution of the present application is implemented as follows:

[0006] The embodiments of the present application provide a gesture recognition method, comprising:

[0007] In the case where a to-be-processed image is acquired, the to-be-processed image is processed according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images; and hand detection is performed on the plurality of reduced to-be-processed images to obtain a hand detection frame;

[0008] An image in the hand detection frame is acquired from the to-be-processed image to obtain a hand region image;

[0009] According to the hand region image, a target gesture in the to-be-processed image is determined.

[0010] The embodiments of the present application provide a gesture recognition device, which comprises:

[0011] A processing unit is configured to, in the case where a to-be-processed image is acquired, process the to-be-processed image according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images;

[0012] A detection unit is configured to perform hand detection on the plurality of reduced to-be-processed images to obtain a hand detection frame;

[0013] An acquisition unit is configured to acquire an image in the hand detection frame from the to-be-processed image to obtain a hand region image;

[0014] A determination unit is configured to determine a target gesture in the image to be processed according to the hand region image.

[0015] The embodiment of the present application provides a gesture recognition device, and the device comprises:

[0016] A memory, a processor and a communication bus, the memory communicates with the processor through the communication bus, the memory stores a gesture recognition program executable by the processor, and when the gesture recognition program is executed, the gesture recognition method described above is executed by the processor.

[0017] The embodiment of the present application provides a storage medium, and a computer program is stored in the storage medium and applied to a gesture recognition device, and the computer program is characterized in that the computer program is executed by a processor to implement the gesture recognition method described above.

[0018] The embodiment of the present application provides a gesture recognition method and device and a storage medium, and the gesture recognition method comprises the following steps: in the case that an image to be processed is acquired, the image to be processed is processed according to a plurality of preset image reduction rules to obtain a plurality of reduced images to be processed; hand detection is performed on the plurality of reduced images to be processed to obtain a hand detection frame; an image in the hand detection frame is acquired from the image to be processed to obtain a hand region image; and a target gesture in the image to be processed is determined according to the hand region image. The gesture recognition device adopts the above method implementation scheme, in the case that the image to be processed is acquired, the gesture recognition device reduces the image to be processed according to a plurality of preset image reduction rules to obtain a plurality of reduced images to be processed, obtains a hand detection frame according to the plurality of reduced images to be processed, and determines a hand region image from the image to be processed by using the hand detection frame. In the hand region image, a gesture part region accounts for a large proportion of the entire hand region image, so that the information of the gesture part obtained by using the gesture detection model of the gesture recognition device is increased. According to the hand region image with large hand information, the gesture recognition model can accurately recognize the gesture in the image to be processed, thereby improving the accuracy of gesture recognition. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A gesture recognition method flowchart is provided for the embodiment of the present application.

[0020] Figure 2 A detection frame diagram of a gesture detection model when detecting a gesture is provided for the embodiment of the present application.

[0021] Figure 3 An exemplary sample hand region image schematic diagram is provided for the embodiment of the present application.

[0022] Figure 4An exemplary gesture recognition model provided by an embodiment of the present application is used to identify a gesture, and a recognition block diagram is shown in the following figure;

[0023] Figure 5 An exemplary gesture recognition structure diagram provided by an embodiment of the present application is shown in the following figure;

[0024] Figure 6 An exemplary gesture recognition device provided by an embodiment of the present application is shown in the following figure Figure 1 ;

[0025] Figure 7 An exemplary gesture recognition device provided by an embodiment of the present application is shown in the following figure Figure 2 . DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0027] The embodiments of the present application provide a gesture recognition method, Figure 1 A gesture recognition method flow provided by an embodiment of the present application is shown in the following figure Figure 1 As shown in the figure, Figure 1 The gesture recognition method can include the following steps.

[0028] S101, in the case of obtaining a to-be-processed image, the to-be-processed image is processed according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images; and hand detection is performed on the plurality of reduced to-be-processed images to obtain a hand detection frame.

[0029] The gesture recognition method provided by the embodiments of the present application is applicable to the scenario of identifying a target gesture in a to-be-processed image.

[0030] In the embodiments of the present application, the gesture recognition device can be implemented in various forms. For example, the gesture recognition device described in the present application can include devices such as a mobile phone, a camera, a tablet computer, a notebook computer, a palm computer, a personal digital assistant (PDA), a portable media player (PMP), a navigation device, a wearable device, a smart bracelet, a pedometer, etc., and devices such as a digital TV, a desktop computer, etc.

[0031] In the embodiments of the present application, the to-be-processed image can be an RGB image; the to-be-processed image can also be a depth image; the to-be-processed image can also be other images; the specific type can be determined according to the actual situation, and the embodiments of the present application do not limit this.

[0032] In the embodiments of the present application, the number of images to be processed can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0033] In the embodiments of the present application, the image to be processed can be an image transmitted by other equipment and received by the gesture recognition device; the image to be processed can also be an image photographed by a camera in the gesture recognition device; the image to be processed can also be an image input by a user into the gesture recognition device; and the specific manner in which the gesture recognition device obtains the image to be processed can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0034] In the embodiments of the present application, the plurality of preset image reduction rules can be image reduction rules configured in the gesture recognition device; the plurality of preset image reduction rules can also be image reduction rules transmitted by other devices and received by the gesture recognition model; the plurality of preset image reduction rules can also be image reduction rules obtained by the gesture recognition model in other manners; and the specific manner in which the gesture recognition model obtains the plurality of preset image reduction rules can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0035] It should be noted that the plurality of preset image reduction rules can be rules for multiple times of downsampling; and one preset image reduction rule corresponds to a rule for one time of downsampling. For example, the number of the plurality of preset image reduction rules can be 5; the number of the plurality of preset image reduction rules can also be 3; and the number of the plurality of preset image reduction rules can also be 10; and the specific number of the plurality of preset image reduction rules can be determined according to actual conditions, and the embodiments of the present application do not limit this.

[0036] In the embodiments of the present application, the number of the plurality of preset image reduction rules can be 5, and the process in which the gesture recognition device processes the image to be processed according to the plurality of preset image reduction rules to obtain a plurality of reduced images to be processed can be that the gesture recognition device performs five times of downsampling processing on the image to be processed to obtain a plurality of reduced images to be processed.

[0037] It should be noted that the gesture recognition device can perform depth-separable convolution on the image to be processed to implement multiple times of downsampling processing on the image to be processed, so as to obtain a plurality of reduced images to be processed.

[0038] For example, if the image to be processed is an image with a resolution of 256*256, the gesture recognition device performs five downsampling processes on the image to be processed to obtain a plurality of reduced images to be processed. After the first downsampling process, the gesture recognition device can obtain an image with a resolution of 128*128. Then, after the second downsampling process, the gesture recognition device can obtain an image with a resolution of 64*64. Then, after the third downsampling process, the gesture recognition device can obtain an image with a resolution of 32*32. Then, after the fourth downsampling process, the gesture recognition device can obtain an image with a resolution of 16*16. Finally, after the fifth downsampling process, the gesture recognition device can obtain an image with a resolution of 8*8. Thus, the plurality of reduced images to be processed are obtained, i.e., the images with resolutions of 128*128, 64*64, 32*32, 16*16, and 8*8.

[0039] In the embodiment of the present application, the gesture recognition device includes a gesture detection model. The hand region image can be an image of a hand region obtained by the gesture recognition device from the image to be processed by using the gesture detection model. Specifically, the gesture recognition device processes the image to be processed according to a plurality of preset image reduction rules to obtain a plurality of reduced images to be processed, and detects the hand region in the plurality of reduced images to be processed to obtain a hand detection frame. The manner includes: the gesture recognition device inputs the image to be processed into the gesture detection model, so that the gesture detection model processes the image to be processed according to a plurality of preset image reduction rules to obtain a plurality of reduced images to be processed; and the gesture detection model detects the hand region in the plurality of reduced images to be processed to obtain a hand detection frame.

[0040] In the embodiment of the present application, the gesture detection model can be a model transmitted by other devices and received by the gesture recognition device. The gesture detection model can also be a model obtained by the gesture recognition device through training. The gesture detection model can also be a model obtained by the gesture recognition device through other manners. The manner in which the gesture recognition device obtains the gesture detection model can be determined according to actual conditions, and the embodiment of the present application does not limit the manner.

[0041] It should be noted that the gesture detection model can be HandDetNet. The gesture detection model can also be a model for detecting a hand region image from an image to be processed. The specific manner can be determined according to actual conditions, and the embodiment of the present application does not limit the manner.

[0042] In the embodiment of the present application, if the gesture detection model is a model obtained by the gesture recognition device through training, the gesture recognition device can first acquire a sample to-be-processed image and a first sample hand region image corresponding to the sample to-be-processed image; the gesture recognition device trains an initial gesture detection model by using the sample to-be-processed image and the first sample hand region image, and obtains the gesture detection model.

[0043] In the embodiment of the present application, the sample to-be-processed image can be a CMU hand image set, a YouTube collected hand image set, or an image containing a hand collected.

[0044] In the embodiment of the present application, the to-be-processed image includes a to-be-processed image of a first perspective and / or a to-be-processed image of a third perspective. Specifically, the to-be-processed image of the first perspective is an image when a user views his own hand from his own angle; and the to-be-processed image of the third perspective is an image when the user views the hand of another user from the perspective of a third party.

[0045] It should be noted that the sample to-be-processed image includes a to-be-processed image of a user viewing his own hand from the user's angle (i.e., a to-be-processed image of the first perspective) and a to-be-processed image of the user viewing his own hand from the perspective of another person (i.e., a to-be-processed image of the third perspective).

[0046] In the embodiment of the present application, the process of the gesture recognition device performing hand detection on the plurality of reduced to-be-processed images to obtain a hand detection frame includes: the gesture recognition device screening a target to-be-processed image meeting a preset resolution from the plurality of reduced to-be-processed images; the gesture recognition device performing feature fusion on the target to-be-processed image to obtain a fused image; and the gesture recognition device performing multi-layer hand detection on the fused image to obtain the hand detection frame.

[0047] In the embodiment of the present application, the process of the gesture recognition device performing multi-layer hand detection on the fused image to obtain a hand detection frame can be that the gesture recognition device performs multi-layer hand detection on the fused image and the target to-be-processed image to obtain the hand detection frame.

[0048] In the embodiment of the present application, the preset resolution can be a resolution configured in the gesture recognition device, the preset resolution can also be a resolution transmitted to the gesture recognition device by another device, and the preset resolution can also be a resolution obtained by the gesture recognition device in another manner. The specific manner in which the gesture recognition device obtains the preset resolution can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0049] It should be noted that the preset resolution can be 32*32 resolution; the preset resolution can also be 16*16 resolution; the preset resolution can also be 16*16 resolution; the preset resolution can also be 8*8 resolution; the preset resolution can also be other resolutions; the specific preset resolution can be determined according to actual conditions, and the embodiments of the application do not limit this.

[0050] It should also be noted that the type of preset resolution can be one, and the type of preset resolution can also be two; the type of preset resolution can also be multiple, for example: the preset resolution includes 32*32 resolution, 16*16 resolution and 8*8 resolution; the specific number of types of preset resolution can be determined according to actual conditions, and the embodiments of the application do not limit this.

[0051] In the embodiments of the application, the number of target processing images is multiple, the pixels of the multiple target processing images are different, and the gesture recognition device performs feature fusion on the target processing images to obtain a fused image. The process can be that the gesture recognition device performs pixel fusion from bottom to top (Bottom-up) and / or pixel fusion from top to bottom (Top-down) on the pixel features of the multiple target processing images according to the Feature Pyramid Network (FPN) and PAN technology, thereby obtaining the fused image; or the gesture recognition device can perform other ways of feature fusion on the multiple target processing images to obtain the fused image; the specific way in which the gesture recognition device performs feature fusion on the multiple target processing images to obtain the fused image can be determined according to actual conditions, and the embodiments of the application do not limit this.

[0052] In the embodiments of the application, the process in which the gesture recognition device performs multi-layer hand detection on the fused image to obtain a hand detection frame includes: the gesture recognition device performs multi-layer hand detection on the fused image to obtain multiple detection frames; and the gesture recognition device screens a hand detection frame from the multiple detection frames.

[0053] In the embodiments of the application, the way in which the gesture recognition device performs multi-layer hand detection on the fused image to obtain multiple detection frames can be a multi-layer detection mechanism of the gesture recognition device MultiScale, which uses resolution to be responsible for large targets and uses large resolution to be responsible for small targets, so as to take into account the detection of near and far hands, thereby obtaining multiple detection frames.

[0054] For example, if the fused image includes an image with a resolution of 32*32, an image with a resolution of 16*16, and an image with a resolution of 8*8, the number of anchor points of the image with a resolution of 32*32 can be set to 2; the number of anchor points of the image with a resolution of 16*16 can be set to 2; the number of anchor points of the image with a resolution of 8*8 can be set to 6; and finally, the number of the plurality of detection boxes that can be obtained is 32*32*2+16*16*2+8*8*6=2944.

[0055] In the embodiments of the present application, the manner in which the gesture recognition apparatus screens the hand detection box from the plurality of detection boxes can be that the gesture recognition apparatus screens the hand detection box from the plurality of detection boxes by using a Non-Maximum Suppression (NMS) manner; or the gesture recognition apparatus can randomly select a detection box from the plurality of detection boxes as the hand detection box; or the gesture recognition apparatus can screen the hand detection box from the plurality of detection boxes by using other manners; and the specific manner can be determined according to actual conditions, which is not limited in the embodiments of the present application.

[0056] In the embodiments of the present application, the process in which the gesture recognition apparatus processes the to-be-processed image according to the plurality of preset image reduction rules to obtain the plurality of reduced to-be-processed images includes: the gesture recognition apparatus pre-processes the to-be-processed image to obtain a pre-processed to-be-processed image; and the gesture recognition apparatus processes the pre-processed to-be-processed image by using the plurality of preset image reduction rules to obtain the plurality of reduced to-be-processed images.

[0057] In the embodiments of the present application, the process in which the gesture recognition apparatus pre-processes the to-be-processed image to obtain the pre-processed to-be-processed image can be that the gesture recognition apparatus adjusts the resolution of the to-be-processed image to obtain a to-be-processed image with an adjusted resolution, which is the pre-processed to-be-processed image; or the gesture recognition apparatus can denoise, filter, or perform other processing on the to-be-processed image to obtain the pre-processed to-be-processed image; and the specific manner can be determined according to actual conditions, which is not limited in the embodiments of the present application.

[0058] In the embodiments of the present application, the process in which the gesture recognition apparatus adjusts the resolution of the to-be-processed image to obtain the to-be-processed image with an adjusted resolution can be that the gesture recognition apparatus adjusts the resolution of the to-be-processed image by using the resolution of the input image configured by the gesture detection model, so as to obtain an image that meets the resolution requirement of the input image of the gesture detection model, that is, the to-be-processed image with an adjusted resolution.

[0059] In the embodiment of the present application, the gesture detection model mainly includes three parts, namely, backbone, neck and head. Specifically, the network structure of the gesture detection model is as shown in the following table: Figure 2 As shown in the table, after the gesture recognition device obtains the to-be-processed image, the gesture recognition device pre-processes the to-be-processed image to obtain a pre-processed to-be-processed image. Then, in the backbone part, the gesture recognition device processes the to-be-processed image according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images, which are respectively a 128*128 resolution image, a 64*64 resolution image, a 32*32 resolution image, a 16*16 resolution image and an 8*8 resolution image. The pre-processed to-be-processed image is a 256*256 RGB image. In the neck part, the gesture recognition device uses the gesture detection model to screen target to-be-processed images from the plurality of reduced to-be-processed images, which are respectively a 32*32 resolution image, a 16*16 resolution image and an 8*8 resolution image. In the head part, the gesture recognition device uses the gesture detection model to perform feature fusion on the target to-be-processed images to obtain a fused image, and performs multi-layer hand detection on the fused image to obtain a plurality of detection boxes, and uses an NMS method to screen a hand detection box from the plurality of detection boxes. The image in the hand detection box is obtained from the to-be-processed image to obtain a hand region image.

[0060] S102, the image in the hand detection box is obtained from the to-be-processed image to obtain a hand region image.

[0061] In the embodiment of the present application, after the gesture recognition device performs hand detection on the plurality of reduced to-be-processed images to obtain a hand detection box, the gesture recognition device can obtain the image in the hand detection box from the to-be-processed image to obtain a hand region image.

[0062] In the embodiment of the present application, the hand region image can be an RGB image; the hand region image can also be a depth image; the hand region image can also be other images; the specific hand region image can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0063] S103, determining a target gesture in the to-be-processed image according to the hand region image.

[0064] In the embodiment of the present application, after the gesture recognition device obtains the image in the hand detection box from the to-be-processed image to obtain a hand region image, the gesture recognition device can determine a target gesture in the to-be-processed image according to the hand region image.

[0065] In the embodiment of the present application, the gesture recognition device comprises a gesture recognition model. The gesture recognition device determines the target gesture in the image to be processed according to the hand region image. The gesture recognition device can input the hand region image into the gesture recognition model, and output the target gesture in the image to be processed by using the gesture recognition model.

[0066] In the embodiment of the present application, the gesture recognition model can be a model transmitted by other devices and received by the gesture recognition device. The gesture recognition model can also be a model obtained by training the gesture recognition device. The gesture recognition model can also be a model obtained by other means by the gesture recognition device. The specific manner in which the gesture recognition device obtains the gesture recognition model can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0067] It should be noted that the gesture recognition model can be HandClsNet. The gesture recognition model can also be a model for determining a target gesture from a hand region image. The specific manner in which the gesture recognition model is determined can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0068] In the embodiment of the present application, if the gesture recognition model is a model obtained by training the gesture recognition device, the gesture recognition model can first obtain a second sample hand region image and a gesture label corresponding to the second sample hand region image. Then, the gesture recognition device trains an initial gesture recognition model by using the second sample hand region image and the gesture label, to obtain the gesture recognition model.

[0069] In the embodiment of the present application, the second sample hand region image can be the same as the first sample hand region image. The second sample hand region image can also be different from the first sample hand region image. The second sample hand region image can also be partially the same as the first sample hand region image. The specific manner in which the second sample hand region image is obtained can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0070] It should be noted that the second sample hand region image can be an image obtained by shooting. The second sample hand region image can also be an image collected from the network. The specific manner in which the second sample hand region image is obtained can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0071] In the embodiment of the present application, the number of categories of the second sample hand region image can be 21 categories. The number of categories of the second sample hand region image can also be 50 categories. The number of categories of the second sample hand region image can be other category numbers. The specific number of categories of the second sample hand region image can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0072] It should be noted that the second sample hand region image can also include 1 type of negative sample, so as to output prompt information without gestures in the case of inputting a non-hand region image.

[0073] For example, if the type of the second sample hand region image can be 21 types, the specific 21 types of sample hand region images can be as shown in the following table: Figure 3 As shown, including: number two gesture, love gesture, number six gesture, number eight gesture, OK gesture, number five gesture, like gesture, fist gesture, heart gesture, middle finger gesture, number zero gesture, number one gesture, number three gesture, number four gesture, number seven gesture, number nine gesture, rock gesture, contempt gesture, little finger gesture, palm gesture, cross finger gesture.

[0074] In the embodiment of the present application, the target gesture can be a gesture label; the target gesture can also be a gesture image; the target gesture can also be other information used to explain the type of gesture; the specific target gesture can be determined according to actual conditions, and the embodiment of the present application does not limit it.

[0075] In the embodiment of the present application, the process of determining the target gesture in the image to be processed by the gesture recognition device according to the hand region image includes: the gesture recognition device processes the hand region image according to a plurality of preset image reduction rules to obtain a plurality of reduced hand images; the gesture recognition device performs pixel fusion processing on the plurality of reduced hand images to obtain a fused hand image; the gesture recognition device adjusts the image resolution of the fused hand image to obtain an adjusted image; and the gesture recognition device determines the target gesture according to the adjusted image.

[0076] In the embodiment of the present application, the gesture recognition device can use a gesture recognition model to perform five times of downsampling processing on the hand region image to obtain five reduced hand images, that is, to obtain a plurality of reduced hand images; or the gesture recognition device can use a gesture recognition model to perform ten times of downsampling processing on the hand region image to obtain ten reduced hand images, that is, to obtain a plurality of reduced hand images; the specific manner can be determined according to actual conditions, and the embodiment of the present application does not limit it.

[0077] It should be noted that the gesture recognition device can use a gesture recognition model to perform full convolution network (Fully Convolutional Networks, FCN) on the hand region image to realize multiple times of downsampling processing on the hand region image, so as to obtain a plurality of reduced hand images.

[0078] For example, if the hand region image is a 256*256 resolution image, the gesture recognition device uses the gesture recognition model to perform five times of downsampling processing on the hand region image to obtain a plurality of reduced hand images. After the gesture recognition device uses the gesture recognition model to perform the first time of downsampling processing on the hand region image, a 128*128 resolution image can be obtained. Then, after the gesture recognition device uses the gesture recognition model to perform the second time of downsampling processing on the 128*128 resolution image, a 64*64 resolution image can be obtained. Then, after the gesture recognition device uses the gesture recognition model to perform the third time of downsampling processing on the 64*64 resolution image, a 32*32 resolution image can be obtained. Then, after the gesture recognition device uses the gesture recognition model to perform the fourth time of downsampling processing on the 32*32 resolution image, a 16*16 resolution image can be obtained. Finally, after the gesture recognition device uses the gesture recognition model to perform the fifth time of downsampling processing on the 16*16 resolution image, an 8*8 resolution image can be obtained, thereby obtaining the plurality of reduced hand images, i.e., the 128*128 resolution image, the 64*64 resolution image, the 32*32 resolution image, the 16*16 resolution image, and the 8*8 resolution image.

[0079] In the embodiment of the present application, the process that the gesture recognition device performs pixel fusion processing on the plurality of reduced hand images to obtain the fused hand image can be: the gesture recognition device screens a plurality of screening images from the plurality of reduced hand images according to a preset screening resolution; the gesture recognition device adjusts the resolutions of the plurality of screening images to a first preset adjustment resolution to obtain a plurality of adjustment resolution images; and then, the gesture recognition device performs pixel fusion processing on the plurality of adjustment resolution images to obtain the fused hand image.

[0080] Specifically, the process that the gesture recognition device performs pixel fusion processing on the plurality of adjustment resolution images to obtain the fused hand image can be that the gesture recognition device performs pixel-by-pixel fusion processing on the plurality of adjustment resolution images to obtain the fused hand image, or that the gesture recognition device performs other fusion processing on the plurality of adjustment resolution images to obtain the fused hand image. The specific process can be determined according to actual conditions, and the embodiment of the present application does not limit the process.

[0081] It should be noted that the preset screening resolution can be 8*8 resolution, or 32*32 resolution, or 128*128 resolution. The specific preset screening resolution can be determined according to actual conditions, and the embodiment of the present application does not limit the preset screening resolution.

[0082] It should be noted that the first preset adjustment resolution can be 8*8 resolution.

[0083] In the embodiment of the present application, the gesture recognition device adjusts the image resolution of the fused hand image to obtain an adjusted image. The gesture recognition device can use a pooling operation to adjust the image resolution of the fused hand image to a first preset resolution, thereby obtaining the adjusted image. The first preset resolution can be a resolution of 1*1. The first preset resolution can also be a resolution of other resolution values. The first preset resolution can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0084] In the embodiment of the present application, the gesture recognition device determines the target gesture according to the adjusted image. The process includes: the gesture recognition device adjusts the channel number of the adjusted image according to a preset channel number to obtain an adjusted image of the preset channel number; the gesture recognition device performs activation processing on the adjusted image of the preset channel number to obtain a preset number of recognition confidence; and the gesture recognition device selects a target confidence with the largest confidence value from the preset number of recognition confidence, and determines the target gesture according to the target adjusted image corresponding to the target confidence.

[0085] In the embodiment of the present application, the preset channel number can be the channel number configured in the gesture recognition device. The preset channel number can also be the channel number transmitted by other devices received by the gesture recognition device. The specific manner of obtaining the preset channel number can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0086] For example, the preset channel number can be 22. The preset channel number can also be 40. The preset channel number can also be other numerical values. The specific manner of obtaining the preset channel number can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0087] In the embodiment of the present application, the gesture recognition device performs activation processing on the adjusted image of the preset channel number to obtain a preset number of recognition confidence. The gesture recognition device can use a softmax activation function to perform activation processing on the adjusted image of the preset channel number to obtain a preset number of recognition confidence. The gesture recognition device can also use other manners to perform activation processing on the adjusted image of the preset channel number to obtain a preset number of recognition confidence. The specific manner of obtaining the preset number of recognition confidence can be determined according to actual conditions, and the embodiment of the present application does not limit this.

[0088] It should be noted that the sum of the values of the preset number of recognition confidence is 1.

[0089] In the embodiment of the present application, the process that the gesture recognition device utilizes the gesture recognition model to perform multiple down-sampling processing on the hand region image to obtain multiple reduced hand images includes: the gesture recognition device adjusts the pixels of the hand region image according to a preset pixel adjustment requirement to obtain an adjusted hand region image; and the gesture recognition device utilizes multiple preset image reduction rules to process the adjusted hand region image respectively to obtain multiple reduced hand images.

[0090] In the embodiment of the present application, the preset pixel adjustment requirement can be a pixel requirement of the gesture recognition model input image.

[0091] For example, the process that the gesture recognition device determines the target gesture in the to-be-processed image according to the hand region image includes: Figure 4 As shown in the figure: after obtaining the hand region image, the gesture recognition device adjusts the pixels of the hand region image according to a preset pixel adjustment requirement (pixel adjustment) to obtain an adjusted hand region image; the gesture recognition device utilizes the gesture recognition model to process the adjusted hand region image according to multiple preset image reduction rules to obtain multiple reduced hand images, which are respectively: an image with a resolution of 128*128, an image with a resolution of 64*64, an image with a resolution of 32*32, an image with a resolution of 16*16 and an image with a resolution of 8*8. The adjusted hand region image is an RGB image with a resolution of 256*256. Then, the gesture recognition device utilizes the gesture recognition model to perform pixel fusion processing on the multiple reduced hand images (the image with a resolution of 8*8, the image with a resolution of 32*32 adjusted to an image with a resolution of 8*8, and the image with a resolution of 128*128 adjusted to an image with a resolution of 8*8) to obtain a fused hand image; the gesture recognition device adjusts the image resolution of the fused hand image to obtain an adjusted image; the gesture recognition device adjusts the channel number of the adjusted image according to a preset channel number (22) to obtain an adjusted image with a preset channel number; the gesture recognition device performs activation processing (softmax 22) on the adjusted image with a preset channel number to obtain a preset number of recognition confidence; the gesture recognition device selects a target confidence with the maximum confidence value from the preset number of recognition confidences, and determines a target gesture according to the target adjusted image corresponding to the target confidence.

[0092] For example, the gesture recognition device includes a gesture detection model and a gesture recognition model, and specifically as shown in the figure: Figure 5As shown: in the training stage, the gesture recognition device trains an initial gesture detection model by using a detection data set (a sample to-be-processed image and a first sample hand region image), to obtain a gesture detection model; and trains an initial gesture recognition model by using a classification data set (a second sample hand region image and a gesture label), to obtain a gesture recognition model. When the gesture recognition device obtains a to-be-processed image, the gesture recognition device pre-processes the to-be-processed image, to obtain a pre-processed to-be-processed image; then the gesture recognition device inputs the pre-processed to-be-processed image into the gesture detection model, and detects a hand region image in the to-be-processed image by using the gesture detection model; subsequently, the gesture recognition device adjusts pixels of the hand region image according to a preset pixel adjustment requirement, to obtain an adjusted hand region image; and the gesture recognition device inputs the adjusted hand region image into the gesture recognition model, to obtain a target gesture in the to-be-processed image.

[0093] It can be understood that, when the gesture recognition device obtains a to-be-processed image, the gesture recognition device performs reduction processing on the to-be-processed image according to a plurality of preset image reduction rules, to obtain a plurality of reduced to-be-processed images, obtains a hand detection frame according to the plurality of reduced to-be-processed images, and determines a hand region image from the to-be-processed image by using the hand detection frame. In the hand region image, a gesture part region accounts for a large proportion of the entire hand region image, so that information of the gesture part obtained by the gesture recognition device by using the gesture detection model is increased. Therefore, the gesture recognition device can accurately recognize a gesture in the to-be-processed image by using the gesture recognition model according to the hand region image with a large amount of hand information, thereby improving the accuracy of gesture recognition.

[0094] Based on the same inventive concept of the gesture recognition method, the embodiment of the present application provides a gesture recognition device 1 corresponding to the gesture recognition method; Figure 6 The gesture recognition device provided by the embodiment of the present application has the component structure as shown in Figure 1 The gesture recognition device 1 can include:

[0095] A processing unit 11, configured to, when a to-be-processed image is obtained, perform processing on the to-be-processed image according to a plurality of preset image reduction rules, to obtain a plurality of reduced to-be-processed images;

[0096] A detection unit 12, configured to perform hand detection on the plurality of reduced to-be-processed images, to obtain a hand detection frame;

[0097] An acquisition unit 13, configured to acquire an image in the hand detection frame from the to-be-processed image, to obtain a hand region image;

[0098] A determination unit 14, configured to determine a target gesture in the to-be-processed image according to the hand region image.

[0099] In some embodiments of the present application, the device further comprises a screening unit and a fusion unit;

[0100] The screening unit is configured to screen a target processing image meeting a preset resolution from the plurality of reduced processing images.

[0101] The fusion unit is configured to perform feature fusion on the target processing image to obtain a fused image.

[0102] The detection unit 12 is configured to perform multi-layer hand detection on the fused image to obtain the hand detection frame.

[0103] In some embodiments of the present application, the number of target processing images is a plurality, and the pixels of the plurality of target processing images are different.

[0104] The fusion unit is configured to fuse the pixel features of the plurality of target processing images in a pixel bottom-up fusion manner and / or a pixel top-down fusion manner to obtain the fused image.

[0105] In some embodiments of the present application, the detection unit 12 is configured to perform multi-layer hand detection on the fused image to obtain a plurality of detection frames.

[0106] The screening unit is configured to screen a hand detection frame from the plurality of detection frames.

[0107] In some embodiments of the present application, the processing unit 11 is configured to perform preprocessing on the processing image to obtain a preprocessed processing image, and perform processing on the preprocessed processing image according to a plurality of preset image reduction rules to obtain the plurality of reduced processing images.

[0108] In some embodiments of the present application, the device further comprises an adjusting unit.

[0109] The processing unit 11 is configured to perform processing on the hand region image according to a plurality of preset image reduction rules to obtain a plurality of reduced hand images.

[0110] The fusion unit is configured to perform pixel fusion processing on the plurality of reduced hand images to obtain a fused hand image.

[0111] The adjusting unit is configured to adjust the image resolution of the fused hand image to obtain an adjusted image.

[0112] The determination unit 14 is configured to determine the target gesture according to the adjusted image.

[0113] In some embodiments of this application, the device further includes an activation unit;

[0114] The adjustment unit is used to adjust the number of channels of the adjusted image according to a preset number of channels to obtain an adjusted image with a preset number of channels;

[0115] The activation unit is used to activate the adjusted image with the preset number of channels to obtain a preset number of recognition confidence scores.

[0116] The filtering unit is used to filter out the target confidence score with the highest confidence score from the preset number of recognition confidence scores;

[0117] The determining unit 14 is used to determine the target gesture based on the target adjustment image corresponding to the target confidence level.

[0118] In some embodiments of this application, the adjustment unit is used to adjust the pixels of the hand region image according to preset pixel adjustment requirements to obtain an adjusted hand region image;

[0119] The processing unit 11 is used to process the adjusted hand region image using the multiple preset image reduction rules to obtain the multiple reduced hand images.

[0120] It should be noted that, in practical applications, the aforementioned processing unit 11, detection unit 12, acquisition unit 13, and determination unit 14 can be implemented by the processor 15 on the gesture recognition device 1, specifically by a CPU (Central Processing Unit), MPU (Microprocessor Unit), DSP (Digital Signal Processor), or FPGA (Field Programmable Gate Array), etc.; the aforementioned data storage can be implemented by the memory 16 on the gesture recognition device 1.

[0121] This application also provides a gesture recognition device 1, such as... Figure 7 As shown, the gesture recognition device 1 includes a processor 15, a memory 16, and a communication bus 17. The memory 16 communicates with the processor 15 through the communication bus 17. The memory 16 stores programs executable by the processor 15. When the program is executed, the gesture recognition method described above is executed by the processor 15.

[0122] In practical applications, the memory 16 can be a volatile memory, such as a Random-Access Memory (RAM), or a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD) or a Solid-State Drive (SSD), or a combination of the above kinds of memories, and provides instructions and data to the processor 15.

[0123] The embodiment of the present application provides a computer readable storage medium, which has a computer program, and the program is executed by the processor 15 to realize the gesture recognition method.

[0124] It can be understood that, in the case that the gesture recognition device acquires the to-be-processed image, the gesture recognition device performs the reducing processing on the to-be-processed image according to a plurality of preset image reducing rules to obtain a plurality of reduced to-be-processed images, obtains a hand detection frame according to the plurality of reduced to-be-processed images, and determines a hand region image from the to-be-processed image by using the hand detection frame. In the hand region image, the gesture part region accounts for a large proportion of the entire hand region image, so that the information of the gesture part obtained by the gesture recognition device by using the gesture detection model is increased. Therefore, the gesture recognition model can accurately recognize the gesture in the to-be-processed image according to the hand region image with a large amount of hand information, thereby improving the accuracy of gesture recognition.

[0125] Those skilled in the art should understand that embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0126] The present application is described with reference to flowcharts and / or block diagrams according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the flowcharts and / or block diagrams. Figure 1one or more processes and / or blocks Figure 1 an apparatus with a function specified in one or more blocks.

[0127] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 one or more processes and / or blocks Figure 1 an apparatus with a function specified in one or more blocks.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 one or more processes and / or blocks Figure 1 an apparatus with a function specified in one or more blocks.

[0129] The above descriptions are only preferred embodiments of the present application, and are not intended to limit the protection scope of the present application.

Claims

1. A gesture recognition method, characterized by, The method comprises: In the case of obtaining a to-be-processed image, the to-be-processed image is processed according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images; and hand detection is performed on the plurality of reduced to-be-processed images to obtain a hand detection frame; the to-be-processed image comprises a to-be-processed image of a first perspective; An image in the hand detection frame is obtained from the to-be-processed image to obtain a hand region image; A target gesture in the to-be-processed image is determined according to the hand region image; The target gesture in the to-be-processed image is determined according to the hand region image, comprising: The hand region image is processed according to a plurality of preset image reduction rules to obtain a plurality of reduced hand images; Pixel fusion processing is performed on the plurality of reduced hand images to obtain a fused hand image; The image resolution of the fused hand image is adjusted to obtain an adjusted image; The target gesture is determined according to the adjusted image; The pixel fusion processing on the plurality of reduced hand images to obtain the fused hand image comprises: A plurality of screening images are screened from the plurality of reduced hand images according to a preset screening resolution; The resolution of the plurality of screening images is adjusted to a first preset adjustment resolution to obtain a plurality of adjustment resolution images; Pixel fusion processing is performed on the plurality of adjustment resolution images to obtain the fused hand image.

2. The method of claim 1, wherein, The hand detection on the plurality of reduced to-be-processed images to obtain a hand detection frame comprises: A target processing image meeting a preset resolution is screened from the plurality of reduced to-be-processed images; Feature fusion is performed on the target processing image to obtain a fused image; Multi-layer hand detection is performed on the fused image to obtain the hand detection frame.

3. The method of claim 2, wherein, The number of target processing images is a plurality, and the pixels of the plurality of target processing images are different; the feature fusion on the target processing image to obtain a fused image comprises: Pixel features of the plurality of target processing images are fused according to a pixel bottom-up fusion manner and / or a pixel top-down fusion manner to obtain the fused image.

4. The method of claim 2, wherein, The multi-layer hand detection on the fused image to obtain the hand detection frame comprises: Multi-layer hand detection is performed on the fused image to obtain a plurality of detection frames; A hand detection frame is screened from the plurality of detection frames.

5. The method of claim 1, wherein, The processing of the to-be-processed image according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images comprises: The to-be-processed image is preprocessed to obtain a preprocessed to-be-processed image; The preprocessed to-be-processed image is processed according to the plurality of preset image reduction rules to obtain the plurality of reduced to-be-processed images.

6. The method of claim 1, wherein, The determination of the target gesture according to the adjusted image comprises: The number of channels of the adjusted image is adjusted according to a preset number of channels to obtain an adjusted image of a preset number of channels; Activation processing is performed on the adjusted image of the preset number of channels to obtain a preset number of recognition confidence levels; Filter a target confidence with the largest confidence value from the preset number of recognition confidences, and determine the target gesture according to a target adjustment image corresponding to the target confidence.

7. The method of claim 1, wherein, The processing of the hand region image according to the plurality of preset image reduction rules comprises: Adjusting pixels of the hand region image according to preset pixel adjustment requirements to obtain an adjusted hand region image; The processing of the adjusted hand region image according to the plurality of preset image reduction rules comprises:

8. A gesture recognition apparatus, characterized by The device comprises: A processing unit configured to, in a case where a to-be-processed image is acquired, process the to-be-processed image according to a plurality of preset image reduction rules to obtain a plurality of reduced to-be-processed images; the to-be-processed image comprises a to-be-processed image of a first perspective angle; A detection unit configured to perform hand detection on the plurality of reduced to-be-processed images to obtain a hand detection frame; An acquisition unit configured to acquire an image in the hand detection frame from the to-be-processed image to obtain a hand region image; A determination unit configured to determine a target gesture in the to-be-processed image according to the hand region image; The device further comprises a fusion unit and an adjustment unit; The processing unit is configured to process the hand region image according to a plurality of preset image reduction rules to obtain a plurality of reduced hand images; The fusion unit is configured to perform pixel fusion processing on the plurality of reduced hand images to obtain a fused hand image; The adjustment unit is configured to adjust an image resolution of the fused hand image to obtain an adjustment image; The determination unit is configured to determine the target gesture according to the adjustment image. The fusion unit is configured to filter a plurality of filtered images from the plurality of reduced hand images according to a preset filtering resolution; adjust resolutions of the plurality of filtered images to a first preset adjustment resolution to obtain a plurality of adjustment resolution images; and perform pixel fusion processing on the plurality of adjustment resolution images to obtain the fused hand image.

9. A gesture recognition apparatus, characterized by The device comprises: A memory, a processor and a communication bus; the memory communicates with the processor through the communication bus; the memory stores a hand gesture recognition program executable by the processor; when the hand gesture recognition program is executed, the processor executes the method according to any one of claims 1 to 7.

10. A storage medium storing a computer program thereon for use in a gesture recognition device, characterized in that, The computer program is executed by the processor to implement the method according to any one of claims 1 to 7. The computer program is executed by the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image recognition method and related device

    CN113536876A

  • Image target detection method and device, equipment and storage medium

    CN113936256A