Gesture recognition method, device and intelligent device

By calculating the difference in video frame images in dynamic gesture recognition and combining it with skin color detection, the problem of low accuracy in dynamic gesture recognition is solved, and the recognition accuracy and user experience are improved.

CN114758268BActive Publication Date: 2025-10-03UBTECH ROBOTICS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210262703.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-10-03
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

Existing dynamic gesture recognition methods have low recognition accuracy when performing gesture recognition, resulting in a poor user experience.

Method used

By obtaining adjacent video frame images in a video clip, calculating their difference results, and performing dynamic gesture recognition when the difference is greater than a preset threshold, combined with skin color area detection, it ensures that only video frame images with dynamic scenes are recognized.

Benefits of technology

Improved the accuracy of dynamic gesture recognition and enhanced user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758268B_ABST
    Figure CN114758268B_ABST
Patent Text Reader

Abstract

This application applies to the field of gesture recognition technology and provides a gesture recognition method, apparatus, and intelligent device, including: obtaining a video clip, the video clip including at least two video frame images; determining the difference between a first video frame image and a second video frame image to obtain a difference result, wherein the first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is the video frame image after the second video frame image; and performing dynamic gesture recognition on the first video frame image where the difference result indicates that the difference meets a condition. This method can improve the accuracy of the dynamic gesture recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gesture recognition technology, and in particular to a gesture recognition method, apparatus, smart device, and computer-readable storage medium. Background Art

[0002] Gestures are a natural form of communication between humans, and gesture recognition is also one of the important research directions of human-computer interaction.

[0003] Gesture recognition can be categorized as static or dynamic. Static gesture recognition involves performing gesture recognition on a single input image, identifying the gesture category within that image. Dynamic gesture recognition involves performing gesture recognition on multiple consecutive images over a period of time, identifying the gesture category within those images. Compared to static gesture recognition, dynamic gesture recognition requires learning the temporal relationship between gestures across multiple frames, a continuous process.

[0004] Video-based dynamic gesture recognition technology doesn't know the start and end frames of a gesture within a video. Therefore, a windowed approach is typically used to predict each segment of the video and output the predicted gesture category. However, this approach often results in a certain error rate in the predictions, resulting in a poor user experience. Summary of the Invention

[0005] The embodiments of the present application provide a computer-readable storage medium that can solve the problem of low recognition accuracy in existing dynamic gesture recognition methods when performing gesture recognition.

[0006] In a first aspect, an embodiment of the present application provides a gesture recognition method, comprising:

[0007] Acquire a video clip, where the video clip includes at least two video frame images;

[0008] determining a difference between a first video frame image and a second video frame image to obtain a difference result, wherein the first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is a subsequent video frame image of the second video frame image;

[0009] Dynamic gesture recognition is performed on the first video frame image whose difference result indicates that the difference meets a condition, wherein the condition includes that the difference is greater than a preset difference threshold.

[0010] In a second aspect, an embodiment of the present application provides a gesture recognition device, comprising:

[0011] A video clip acquisition module is used to acquire a video clip, wherein the video clip includes at least two video frame images;

[0012] a difference result determining module, configured to determine a difference between a first video frame image and a second video frame image to obtain a difference result, wherein the first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is a subsequent video frame image of the second video frame image;

[0013] The dynamic gesture recognition module is configured to perform dynamic gesture recognition on the first video frame image whose difference result indicates that the difference satisfies a condition, wherein the condition includes that the difference is greater than a preset difference threshold.

[0014] In a third aspect, an embodiment of the present application provides an intelligent device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a smart device, enables the smart device to execute the method described in the first aspect above.

[0017] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0018] In an embodiment of the present application, before dynamic gesture recognition is performed on a first video frame image, the difference between the first video frame image and a previous video frame image (i.e., the second video frame image) of the first video frame image is first determined, and dynamic gesture recognition is performed on the first video frame image only when it is determined that the difference between the first video frame image is greater than a preset difference threshold. When the difference between the first video frame image is greater than the preset difference threshold, it indicates that the first video frame image contains a dynamic picture compared with the second video frame image. Therefore, dynamic gesture recognition is only performed on the first video frame image for which the difference result indicates that the difference meets the condition. This ensures that there are no static gestures in the video frame images for which dynamic gesture recognition is performed, thereby improving the accuracy of the obtained dynamic gesture recognition results and thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art.

[0020] Figure 1 This is a flowchart of a gesture recognition method provided by an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of a grayscale image corresponding to a first video frame image provided by an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a grayscale image corresponding to a second video frame image provided by an embodiment of the present application;

[0023] Figure 4 This is provided in one embodiment of the present application Figure 2 and Figure 3 Schematic diagram of the corresponding difference map;

[0024] Figure 5 Another embodiment of this application provides the basis Figure 4 Schematic diagram of the generated binary image;

[0025] Figure 6 This is an embodiment of the present application. Figure 5 Schematic diagram of the binary image after morphological processing;

[0026] Figure 7 is a schematic diagram of a mask image provided by another embodiment of the present application;

[0027] Figure 8 is a schematic diagram of the skin color area provided in the embodiment of the present application;

[0028] Figure 9 is a structural diagram of a gesture recognition device provided in an embodiment of the present application;

[0029] Figure 10 It is a schematic diagram of the structure of the smart device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0030] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0031] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0032] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0033] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0034] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.

[0035] Example 1:

[0036] When using a windowing method to predict gesture categories segment by segment, if a video segment only contains static gestures (note that static gestures here are different from the static gesture recognition mentioned above. Static gestures refer to the user's hand not changing or moving in multiple consecutive frames), for example, if the user only places their hand in the picture but does not perform a dynamic gesture action in the preset category, the dynamic gesture recognition model will also return a prediction result for the gesture category of the video segment based on probability, resulting in a poor user experience.

[0037] In order to enable the dynamic gesture recognition model to perform better in actual use, an embodiment of the present application proposes a gesture recognition method, which does not require training of the dynamic gesture recognition model and can also enable the dynamic gesture recognition model to perform better.

[0038] The gesture recognition method provided in the embodiments of the present application is described below with reference to the accompanying drawings.

[0039] Figure 1 A flowchart of a gesture recognition method provided by an embodiment of the present application is shown, and is described in detail as follows:

[0040] Step S11: obtaining a video clip, where the video clip includes at least two video frame images.

[0041] Among them, the video clip here can be a video clip in a pre-acquired video file, or it can be a currently acquired video clip. For example, when the gesture recognition method of an embodiment of the present application is applied to a robot that can recognize gestures, the started robot will acquire multiple video frame images in real time, and these video frame images constitute the above-mentioned video clip.

[0042] Step S12: determine the difference between the first video frame image and the second video frame image to obtain a difference result, wherein the first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is the next video frame image after the second video frame image.

[0043] In the embodiment of the present application, when the video clip is a video clip from a pre-acquired video file, considering that gesture recognition is offline recognition, i.e., real-time performance is not a high requirement, gesture recognition can be performed on the video frame images of the video clip in a forward-to-backward order. That is, the second video frame image of the video clip is first processed as the first video frame image of the embodiment of the present application, and then the third video frame image of the video clip is processed as the first video frame image of the embodiment of the present application, and so on. If the video clip is the currently acquired video clip, the most recently acquired video frame image is generally used as the aforementioned first video frame image.

[0044] Since intelligent devices such as robots usually store the video streams they collect, after determining the first video frame image, the previous video frame image of the first video frame image can be obtained from the stored video stream (i.e., video clip), and the obtained video frame image is used as the second video frame image. Afterwards, the first video frame image is compared with the second video frame image, for example, the pixel values ​​at corresponding positions of the two video frame images are compared to determine the difference between the two video frame images. The difference here refers to the difference between the two video frame images. For example, if the pixel values ​​at the same corresponding position are different, it indicates that there is a difference between the two video frame images at this position.

[0045] Step S13 , performing dynamic gesture recognition on the first video frame image whose difference result indicates that the difference meets a condition, wherein the condition includes that the difference is greater than a preset difference threshold.

[0046] Specifically, when the absolute value difference between the pixel values ​​at corresponding positions in the first video frame image and the second video frame image is greater than 0, it indicates that there is a difference between the first video frame image and the second video frame image. In this case, the difference condition is satisfied including: the absolute value difference between the pixel values ​​at the corresponding positions is greater than a preset first threshold (the preset difference threshold includes the first threshold). Furthermore, the difference condition is considered satisfied only when the number of pixels whose absolute value difference in pixel values ​​is greater than the first threshold is greater than a preset second threshold. In this case, the preset difference threshold includes the first and second thresholds mentioned above.

[0047] In an embodiment of the present application, before dynamic gesture recognition is performed on a first video frame image, the difference between the first video frame image and a previous video frame image (i.e., the second video frame image) of the first video frame image is first determined, and dynamic gesture recognition is performed on the first video frame image only when it is determined that the difference between the first video frame image is greater than a preset difference threshold. When the difference between the first video frame image is greater than the preset difference threshold, it indicates that the first video frame image contains a dynamic picture compared with the second video frame image. Therefore, dynamic gesture recognition is only performed on the first video frame image for which the difference result indicates that the difference meets the condition. This ensures that there are no static gestures in the video frame images for which dynamic gesture recognition is performed, thereby improving the accuracy of the obtained dynamic gesture recognition results and thereby improving the user experience.

[0048] In some embodiments, considering that there is a large difference between the first video frame image and the second video frame image, it may be that the background of the video frame is changing rather than the gesture. Therefore, it is necessary to determine whether the change is a gesture. In this case, the above step S13 includes:

[0049] A1. Detect whether the target area of ​​the target video frame image includes a skin color area, wherein the target video frame image is the first video frame image whose difference result indicates that the difference satisfies the condition, and the target area is the area where there is a difference between the first video frame image and the second video frame image.

[0050] Specifically, it is detected whether there are pixel points in the target area with pixel values ​​that match the skin color of the human body. If so, the area where the pixel points corresponding to these pixel values ​​are located is used as the skin color area.

[0051] A2. Perform dynamic gesture recognition on the target video frame image with skin color area.

[0052] In this embodiment, considering that the user's hands, such as the palm, are typically not obscured by clothing, when the user's gesture changes, the changed pixels typically include pixel values ​​that match skin color. Specifically, in this embodiment, after determining that the difference in the first video frame image satisfies a condition, a further determination is made as to whether the changed region in the first video frame image includes a skin color region. Only after determining that the changed region does the dynamic gesture recognition process on the first video frame image is performed is dynamic gesture recognition performed on the first video frame image, thereby further improving the accuracy of the subsequent dynamic gesture recognition.

[0053] In some embodiments, the above step S12 includes:

[0054] B1. Convert the first video frame image and the second video frame image into grayscale images respectively.

[0055] In this embodiment, in order to facilitate subsequent matching, the first video frame image and the second video frame image are both converted into grayscale images.

[0056] B2. Calculate the absolute difference between the pixel values ​​of the grayscale image of the first video frame image and the pixel values ​​of the grayscale image of the second video frame image to obtain a difference image.

[0057] Specifically, assuming that the first video frame image is represented by cur_image and the second video frame image is represented by pri_gray, the absolute difference between the pixel values ​​of the grayscale image of the first video frame image and the pixel values ​​of the grayscale image of the second video frame image is calculated using the following formula:

[0058] abs_gray=abs(pri_gray-gray).

[0059] In this embodiment, considering that the difference between the pixel values ​​of the grayscale images of two video frame images may be a negative value, and for an 8-bit image, the pixel value is between 0 and 255, therefore, in order to ensure that the image can be displayed accurately and the absolute difference can also reflect the difference between the two pixel values, the absolute difference between the pixel values ​​of the grayscale images is calculated.

[0060] B3. Determine the difference result according to the pixel values ​​of the difference image.

[0061] In this embodiment, since the pixel value of each pixel point in the differential image is the absolute difference between the pixel values ​​of the grayscale images of the first video frame image and the second video frame image, and when the absolute difference is not 0, it indicates that there is a difference between the first video frame image and the second video frame image at this position, the difference result can be accurately determined based on the pixel value of the differential image.

[0062] In some embodiments, the preset difference threshold includes a first threshold and a second threshold, and B3 includes:

[0063] B31. Set the pixel values ​​in the differential image that are greater than the first threshold as first pixel values, and set the pixel values ​​in the differential image that are not greater than the first threshold as second pixel values, to obtain a binary image corresponding to the first video frame image.

[0064] In this embodiment, the first pixel value is usually set to 255, and the second pixel value is usually set to 0. Of course, the first pixel value and the second pixel value can also be set to other values ​​between 0 and 255, which is not limited here.

[0065] In this embodiment, considering that there will be differences between the first video frame image and the second video frame image when the light changes or dust exists, and the changes caused by the light changes or dust are small, that is, the absolute difference between the obtained pixel values ​​is small, therefore, the first threshold can be set to 40. In this way, it is possible to avoid judging smaller disturbances (these disturbances are not caused by gesture changes) as differences between the first video frame image and the second video frame image, and it is also possible to determine the actual differences between the first video frame image and the second video frame image.

[0066] In some embodiments, the first threshold may also be determined by:

[0067] The distribution pattern of the pixel values ​​of each pixel point in the difference image is statistically analyzed to determine the pixel values ​​with the highest concentration of distribution. A pixel value is then selected from these pixel values ​​as the first threshold. For example, if the majority of the pixel values ​​in the difference image are around 50, the first threshold is set to 50. This setting can improve the accuracy of the set first threshold.

[0068] In some embodiments, noise points in the difference image may be filtered first, and then a binary image corresponding to the first video frame image may be generated according to the difference image from which the noise points have been filtered.

[0069] Specifically, the possible types of noise points in the difference image are first determined. A corresponding filtering method is then selected based on the noise point type. Finally, the difference image is denoised using the selected filtering method. For example, considering that salt and pepper noise is common in difference images, and median filtering is particularly suitable for removing salt and pepper noise, median filtering can be used to denoise the difference image. Subsequently, a binary image is generated from the denoised difference image.

[0070] B32. If the number of first pixel values ​​in the binary image corresponding to the first video frame image is greater than the second threshold, a difference result indicating that the difference satisfies a condition is obtained; otherwise, a difference result indicating that the difference does not satisfy the condition is obtained.

[0071] Specifically, considering that when the user's gesture changes, more pixels change, the second threshold value cannot be set too small. In some embodiments, the second threshold value can be set to 200.

[0072] In steps B31 and B32 above, the difference between the first video frame image and the second video frame image is considered to meet the condition only when the number of first pixel values ​​in the binary image corresponding to the first video frame image is greater than the second threshold. Since the first pixel values ​​are obtained by the absolute difference value greater than the first threshold, a large number of first pixel values ​​indicates a large difference between the first video frame image and the second video frame image. In other words, the difference result obtained using this method is more accurate.

[0073] In some embodiments, before step B32, the process includes:

[0074] Perform morphological processing on the binary image. Correspondingly, step B32 specifically includes:

[0075] If the number of first pixel values ​​in the morphologically processed binary image corresponding to the first video frame image is greater than the second threshold, a difference result indicating that the difference meets the condition is obtained; otherwise, a difference result indicating that the difference does not meet the condition is obtained.

[0076] In the embodiment of the present application, since the binary image obtained by threshold segmentation often shows incomplete or fragmented object shapes, morphological processing can make them fuller or remove redundant pixels. Therefore, counting the number of first pixel values ​​in the binary image after morphological processing is more accurate. The morphological processing in the embodiment of the present application includes erosion and dilation operations.

[0077] In some embodiments, step A1 includes:

[0078] A11. Determine a target area of ​​the target video frame image according to the binary image corresponding to the target video frame image.

[0079] In this embodiment, since the target video frame image and its corresponding binary image have the same resolution, each pixel of the target video frame image has a corresponding pixel in the binary image of the target video frame image. In this embodiment, if the pixel value of a pixel in the binary image corresponding to the target video frame image is a first pixel value, the pixel value of the pixel corresponding to the pixel in the target video frame image remains unchanged. If the pixel value of a pixel in the binary image corresponding to the target video frame image is a second pixel value, the pixel value of the pixel corresponding to the pixel in the target video frame image becomes 0. For example, assuming that the first pixel value is 255 and the second pixel value is 0, the pixel value at (0, 0) of the target video frame image is 200, and the pixel value at (0, 0) of the binary image corresponding to the target video frame image is 0, then when a new image is determined based on the target video frame image and its corresponding binary image, the pixel value at (0, 0) of the new image is set to 0. Of course, if the pixel value at (0, 0) of the binary image corresponding to the target video frame image is 255, then when determining a new image based on the target video frame image and its corresponding binary image, the pixel value at (0, 0) of the new image is set to 200. By setting the pixel values ​​of each pixel point of the new image in this way, as long as the pixel values ​​of the pixels of the new image are not 0, the area corresponding to these pixels is the target area of ​​the target video frame image.

[0080] A12. Convert the pixel points corresponding to the target area into YUV space to obtain new pixel points.

[0081] The above-mentioned YUV space is also called YCrCb space.

[0082] In this embodiment, since the YUV space is more suitable for determining the skin color area, the pixels of the target area need to be converted to the YUV space. For example, if the target video frame image is in the red, green, and blue (RGB) space, the target video frame image needs to be converted to the YUV space.

[0083] A13. If the new pixel point exists within the preset elliptical area, it is determined that the target area includes a skin color area.

[0084] The size of the elliptical area may be determined empirically, or may be determined according to the size of a new image (ie, an image determined according to the target video frame image and its corresponding binary image).

[0085] Since pixels with pixel values ​​matching skin color are gathered in an elliptical area after being converted into YUV space, it is only necessary to determine whether there are new pixels in the elliptical area to quickly determine whether the target area contains a skin color area.

[0086] In some embodiments, the target new pixel point is a new pixel point within the preset elliptical area, and within the target area, the area corresponding to the target new pixel point is a skin color area. Step A2 includes:

[0087] It is detected whether the skin color area of ​​the target video frame image is larger than a preset area threshold. If so, dynamic gesture recognition is performed on the target video frame image.

[0088] Specifically, the number of new pixels within the preset elliptical area is accumulated to obtain the size of the skin color area.

[0089] In the embodiment of the present application, since a smaller skin color area indicates that the first video frame image does not have a changed gesture, performing dynamic gesture recognition only on target video frame images with a larger skin color area can improve the accuracy of dynamic gesture recognition.

[0090] In some embodiments, the gesture recognition method provided by the embodiments of the present application further includes:

[0091] Dynamic gesture recognition is not performed on the first video frame image for which the difference result indicates that the difference does not meet the condition.

[0092] Specifically, when the difference between the first video frame image and the second video frame image does not meet the condition, it indicates that the difference between the first video frame image and the second video frame image is small, that is, compared with the second video frame image, there is no change in the first video frame image. In this case, dynamic gesture recognition is not performed on the first video frame image, which can improve the accuracy of the subsequent dynamic gesture recognition results.

[0093] In order to more clearly describe the gesture recognition method provided in the embodiments of the present application, a specific example is given below for description.

[0094] (1) Calculate the grayscale difference between two video frame images

[0095] 1) Assume that the grayscale image of the first video frame is gray (such as Figure 2 As shown), the grayscale image of the second video frame is pri_gary (as shown Figure 3 As shown), calculate the difference between the grayscale image of pri_gray and gray (as shown Figure 4 As shown), the absolute difference between the pixel values ​​of the two grayscale images is obtained:

[0096] abs_gray=abs(pri_gray-gray) (1)

[0097] 2) Perform median filtering on the difference image to filter out the noise points in the difference image and obtain a new difference image blur_gray.

[0098] 3) Binarize blur_gray. Specifically, set the binarization threshold threshold, modify the pixel values ​​(also called pixel grayscale values) greater than the binarization threshold to 255, and modify the pixel values ​​less than the binarization threshold to 0, and you can get the binarized grayscale image, as shown in the following example: Figure 5 As shown in Figure 2. Since the binary image obtained by threshold segmentation often shows incomplete object shape, morphological processing can be used to make it fuller or remove redundant pixels. Figure 5 After morphological processing, we get Figure 6 The schematic diagram shown.

[0099] After obtaining the binary image after the erosion and dilation operation (denoted as mask_gray), you can traverse the image pixels and count the sum of the number of pixels with a grayscale value of 255, sum_p. When sum_p is greater than the second threshold (such as 200), it can be considered that the pixel values ​​of the first video frame image and the second video frame image have changed significantly, that is, there is a moving part in the picture.

[0100] (2) Skin color detection within the changing area

[0101] By using the method in step (1), it is possible to determine whether there is a moving part in the image. However, simply determining whether there is a moving part is not enough. It is necessary to further determine whether the moving part contains skin color, so as to filter out some erroneous judgments caused by changes in other objects.

[0102] First, combine the mask_gray obtained in step (1) with the first video frame image (assuming it is represented by cur_image), that is, in cur_image, the pixel values ​​corresponding to the areas where the pixel grayscale value of mask_gray is 0 are also changed to 0, and we can get Figure 7 The mask_image shown.

[0103] Secondly, the mask_image is calculated to see if there are skin-colored pixels. The principle is to convert the RGB image to the YCRCB space, and the skin-colored pixels will be gathered into an elliptical area. Specifically, the size of the elliptical area is defined first, and then each RGB pixel is converted to the YCrCb space. Then the converted pixel is compared to see if it is within the defined elliptical area. If so, it is determined that these pixels are the pixels corresponding to the skin. The skin-colored area is as follows: Figure 8 shown.

[0104] Finally, set the region threshold skin_thre, traverse the image, and calculate the number of pixel values ​​in the skin color area. When the number of pixel values ​​is greater than skin_thre, it is considered that skin color exists in the moving part.

[0105] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0106] Example 2 :

[0107] Corresponding to the gesture recognition method of the above embodiment, Figure 9 A structural block diagram of a gesture recognition device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0108] Reference Figure 9 The gesture recognition device 9 includes: a video clip acquisition module 91, a difference result determination module 92 and a dynamic gesture recognition module 93.

[0109] The video segment acquisition module 91 is configured to acquire a video segment, where the video segment includes at least two video frame images.

[0110] The difference result determination module 92 is used to determine the difference between the first video frame image and the second video frame image to obtain a difference result, wherein the above-mentioned first video frame image and the above-mentioned second video frame image are adjacent video frame images obtained from the above-mentioned video clip, and the above-mentioned first video frame image is the next video frame image after the above-mentioned second video frame image.

[0111] The dynamic gesture recognition module 93 is configured to perform dynamic gesture recognition on the first video frame image whose difference result indicates that the difference satisfies a condition, wherein the condition includes that the difference is greater than a preset difference threshold.

[0112] In an embodiment of the present application, before dynamic gesture recognition is performed on a first video frame image, the difference between the first video frame image and a previous video frame image (i.e., the second video frame image) of the first video frame image is first determined, and dynamic gesture recognition is performed on the first video frame image only when it is determined that the difference between the first video frame image is greater than a preset difference threshold. When the difference between the first video frame image is greater than the preset difference threshold, it indicates that the first video frame image contains a dynamic picture compared with the second video frame image. Therefore, dynamic gesture recognition is only performed on the first video frame image for which the difference result indicates that the difference meets the condition. This ensures that there are no static gestures in the video frame images for which dynamic gesture recognition is performed, thereby improving the accuracy of the obtained dynamic gesture recognition results and thereby improving the user experience.

[0113] In some embodiments, the dynamic gesture recognition module 93 includes:

[0114] A skin color area detection unit is used to detect whether the target area of ​​the target video frame image contains a skin color area, wherein the above-mentioned target video frame image is the above-mentioned first video frame image whose difference result indicates that the difference meets the condition, and the above-mentioned target area is the area where there is a difference between the above-mentioned first video frame image and the above-mentioned second video frame image.

[0115] The dynamic gesture recognition unit is used to perform dynamic gesture recognition on the target video frame image with skin color area.

[0116] In some embodiments, the difference result determination module 92 includes:

[0117] The grayscale image conversion unit is used to convert the first video frame image and the second video frame image into grayscale images respectively.

[0118] The difference map determining unit is used to calculate the absolute difference between the pixel values ​​of the grayscale image of the first video frame image and the pixel values ​​of the grayscale image of the second video frame image to obtain a difference map.

[0119] The difference result determining unit is configured to determine the difference result according to the pixel values ​​of the difference image.

[0120] In some embodiments, the preset difference threshold includes a first threshold and a second threshold, and the difference result determination unit includes:

[0121] The binary image generation unit corresponding to the first video frame image is used to set the pixel values ​​in the above-mentioned differential image that are greater than the above-mentioned first threshold to the first pixel value, and set the pixel values ​​in the above-mentioned differential image that are not greater than the above-mentioned first threshold to the second pixel value, so as to obtain the binary image corresponding to the above-mentioned first video frame image.

[0122] A different difference result generating unit is used to obtain a difference result indicating that the difference meets the condition if the number of first pixel values ​​in the binary image corresponding to the above-mentioned first video frame image is greater than the above-mentioned second threshold, otherwise, obtain a difference result indicating that the difference does not meet the condition.

[0123] In some embodiments, the binary image generation unit corresponding to the above-mentioned first video frame image is specifically used to: after filtering the noise points in the differential image, set the pixel values ​​in the differential image of the filtered noise points that are greater than the above-mentioned first threshold to the first pixel value, and set the pixel values ​​in the differential image of the filtered noise points that are not greater than the above-mentioned first threshold to the second pixel value, to obtain the binary image corresponding to the above-mentioned first video frame image.

[0124] In some embodiments, the gesture recognition device 9 further includes:

[0125] The morphological processing module is used to perform morphological processing on binary images.

[0126] Correspondingly, the above-mentioned different difference result generating units are specifically used for:

[0127] If the number of first pixel values ​​in the morphologically processed binary image corresponding to the first video frame image is greater than the second threshold, a difference result indicating that the difference meets the condition is obtained; otherwise, a difference result indicating that the difference does not meet the condition is obtained.

[0128] In some embodiments, the skin color area detection unit includes:

[0129] The target area determination unit is used to determine the target area of ​​the target video frame image according to the binary image corresponding to the target video frame image.

[0130] The new pixel point determination unit is used to convert the pixel points corresponding to the target area into the YUV space to obtain new pixel points.

[0131] The skin color area determination unit is configured to determine whether the target area includes a skin color area if the new pixel point exists within a preset elliptical area.

[0132] In some embodiments, the target new pixel point is a new pixel point within the preset elliptical area, and within the target area, the area corresponding to the target new pixel point is a skin color area. The dynamic gesture recognition unit is specifically configured to:

[0133] It is detected whether the skin color area of ​​the target video frame image is larger than a preset area threshold. If so, dynamic gesture recognition is performed on the target video frame image.

[0134] In some embodiments, the gesture recognition device further includes:

[0135] The non-response module is configured to not perform dynamic gesture recognition on the first video frame image whose difference indicated by the difference result does not meet the condition.

[0136] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0137] Example 3:

[0138] Figure 10 This is a schematic diagram of the structure of a smart device provided in one embodiment of the present application. Figure 10 As shown, the smart device 10 of this embodiment includes: at least one processor 100 ( Figure 10Only one processor is shown in the figure), a memory 101, and a computer program 102 stored in the above-mentioned memory 101 and capable of running on the above-mentioned at least one processor 100. When the above-mentioned processor 100 executes the above-mentioned computer program 102, the steps in any of the above-mentioned method embodiments are implemented.

[0139] The smart device 10 can be a computing device such as a robot, a mobile phone, a desktop computer, a notebook, a PDA, or a cloud server. The smart device can include, but is not limited to, a processor 100 and a memory 101. Those skilled in the art will understand that Figure 10 This is merely an example of the smart device 10 and does not constitute a limitation on the smart device 10 . The smart device 10 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0140] The processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0141] In some embodiments, the memory 101 may be an internal storage unit of the smart device 10, such as a hard disk or memory of the smart device 10. In other embodiments, the memory 101 may also be an external storage device of the smart device 10, such as a plug-in hard disk equipped on the smart device 10, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Furthermore, the memory 101 may also include both an internal storage unit of the smart device 10 and an external storage device. The memory 101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 101 may also be used to temporarily store data that has been output or is to be output.

[0142] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0143] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0144] An embodiment of the present application provides a computer program product. When the computer program product is run on a smart device, the smart device can implement the steps in the above-mentioned various method embodiments when executing the computer program product.

[0145] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / smart device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, mobile hard disk, magnetic disk or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0146] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0147] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0148] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0149] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0150] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A gesture recognition method, characterized in that: include: Acquire a video clip, where the video clip includes at least two video frame images; Converting the first video frame image and the second video frame image into grayscale images respectively; Calculating the absolute difference between the pixel values ​​of the grayscale image of the first video frame image and the pixel values ​​of the grayscale image of the second video frame image to obtain a difference image; determining a difference result according to pixel values ​​of the difference image; The first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is a subsequent video frame image of the second video frame image; detecting whether a target area of ​​a target video frame image includes a skin color area, wherein the target video frame image is the first video frame image for which the difference result indicates that a difference satisfies a condition, and the target area is an area where a difference exists between the first video frame image and the second video frame image, and the condition includes that the difference is greater than a preset difference threshold; Performing dynamic gesture recognition on the target video frame image having the skin color area; The preset difference threshold includes a first threshold and a second threshold, and determining the difference result according to the pixel value of the difference image includes: Setting pixel values ​​in the difference image that are greater than the first threshold as first pixel values, and setting pixel values ​​in the difference image that are not greater than the first threshold as second pixel values, to obtain a binary image corresponding to the first video frame image, wherein the first threshold is determined by statistically analyzing the distribution pattern of pixel values ​​of each pixel point in the difference image, determining each pixel value with the most concentrated distribution, and selecting a pixel value from these pixel values ​​as the first threshold; If the number of first pixel values ​​in the binary image corresponding to the first video frame image is greater than the second threshold, a difference result indicating that the difference meets the condition is obtained; otherwise, a difference result indicating that the difference does not meet the condition is obtained.

2. The gesture recognition method according to claim 1, wherein: The detecting whether the target area of ​​the target video frame image includes a skin color area includes: determining a target region of the target video frame image according to a binary image corresponding to the target video frame image; Convert the pixel points corresponding to the target area into YUV space to obtain new pixel points; If the new pixel point exists within the preset elliptical area, it is determined that the target area includes a skin color area.

3. The gesture recognition method according to claim 2, wherein: The target new pixel point is a new pixel point within the preset elliptical area, and within the target area, an area corresponding to the target new pixel point is a skin color area. The performing dynamic gesture recognition on the target video frame image having the skin color area includes: It is detected whether the skin color area of ​​the target video frame image is larger than a preset area threshold. If so, dynamic gesture recognition is performed on the target video frame image.

4. The gesture recognition method according to any one of claims 1 to 3, wherein: The gesture recognition method further includes: Dynamic gesture recognition is not performed on the first video frame image for which the difference result indicates that the difference does not meet the condition.

5. A gesture recognition device, characterized in that: include: A video clip acquisition module is used to acquire a video clip, wherein the video clip includes at least two video frame images; A grayscale image conversion unit, configured to convert the first video frame image and the second video frame image into grayscale images respectively; a difference map determining unit, configured to calculate an absolute difference between pixel values ​​of the grayscale map of the first video frame image and pixel values ​​of the grayscale map of the second video frame image to obtain a difference map; a difference result determining unit, configured to determine a difference result according to pixel values ​​of the difference image; The first video frame image and the second video frame image are adjacent video frame images obtained from the video clip, and the first video frame image is a subsequent video frame image of the second video frame image; a skin color area detection unit, configured to detect whether a target area of ​​a target video frame image includes a skin color area, wherein the target video frame image is the first video frame image for which the difference result indicates that a difference satisfies a condition, and the target area is an area where a difference exists between the first video frame image and the second video frame image, and the condition includes that the difference is greater than a preset difference threshold; a dynamic gesture recognition unit, configured to perform dynamic gesture recognition on the target video frame image having a skin color area; The preset difference threshold includes a first threshold and a second threshold, and the difference result determination unit includes: a binary image generation unit corresponding to the first video frame image, configured to set pixel values ​​in the difference image that are greater than the first threshold as first pixel values, and set pixel values ​​in the difference image that are not greater than the first threshold as second pixel values, to obtain a binary image corresponding to the first video frame image, wherein the first threshold is determined by statistically analyzing the distribution pattern of pixel values ​​of each pixel point in the difference image, determining each pixel value with the most concentrated distribution, and selecting a pixel value from these pixel values ​​as the first threshold; A different difference result generating unit is used to obtain a difference result indicating that the difference meets the condition if the number of first pixel values ​​in the binary image corresponding to the first video frame image is greater than the second threshold, otherwise, obtain a difference result indicating that the difference does not meet the condition.

6. An intelligent device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method and system for extracting gesture image

    CN106503651A

  • Dynamic gesture recognition method and device

    CN109598206A

  • Gesture recognition method and device, terminal equipment and computer readable storage medium

    CN111027395A

  • Video compression method and device and storage medium

    CN112581489A