Gesture detection method and apparatus, electronic device, and storage medium

By combining multiple frames of images and using a neural network model to detect gestures, this method solves the problems of insufficient accuracy and high equipment cost in existing gesture detection technologies, and achieves efficient gesture detection using ordinary cameras.

CN114299612BActive Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111613714.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-11-04
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

Existing gesture detection methods lack accuracy and require specialized image acquisition equipment such as depth cameras or infrared cameras, which are costly.

Method used

By acquiring multiple frames of images and combining information from the previous frame, gestures in the i-th frame can be detected. This method uses a regular camera for data acquisition and combines it with a neural network model for gesture detection, thus avoiding missed detections and reducing costs.

Benefits of technology

It improves the accuracy of gesture detection, reduces reliance on specialized equipment, lowers costs, and is applicable to both general-purpose terminal devices and augmented/virtual reality devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299612B_ABST
    Figure CN114299612B_ABST
Patent Text Reader

Abstract

The present disclosure provides a gesture detection method and device, electronic equipment and storage medium, relates to the technical field of image processing, and particularly relates to the field of artificial intelligence. The specific implementation scheme is: obtaining target information, the target information indicating whether a first target object appears in a target collected image, wherein the first target object represents a hand, the target collected image is at least one image collected at a time earlier than an i-th collected image, i is an integer greater than 1; according to the target information, a first position of the first target object in the i-th collected image is obtained; based on the first position, action feature information of the first target object in the i-th collected image is detected; and according to the action feature information, it is determined whether action information generated by the first target object in the i-th collected image is a predetermined gesture. According to the present disclosure, the detection accuracy can be improved and the cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and in particular, to the technical field of artificial intelligence. BACKGROUND

[0002] In the related art, gesture detection can be performed by two methods. The first method is to detect gesture information by using a single frame image. The second method is to detect gestures by using depth information and thermal information of a hand. The first method has insufficient detection accuracy. In the second method, depth information and thermal information are respectively acquired by using a dedicated image acquisition device, such as a depth camera or an infrared camera. The requirement for the image acquisition device is high. SUMMARY

[0003] The present disclosure provides a gesture detection method and device, an electronic device, and a storage medium.

[0004] According to an aspect of the present disclosure, a gesture detection method is provided, which includes:

[0005] acquiring target information, the target information representing whether a first target object appears in a target acquisition image, wherein the first target object represents a hand, and the target acquisition image is at least one image acquired earlier than an i-th acquisition image, i being an integer greater than 1;

[0006] acquiring, according to the target information, position information of the first target object in the i-th acquisition image;

[0007] detecting, based on the position information, action feature information of the first target object in the i-th acquisition image;

[0008] determining, according to the action feature information, whether action information generated by the first target object in the i-th acquisition image is a predetermined gesture.

[0009] According to another aspect of the present disclosure, a gesture detection device is provided, which includes:

[0010] a first acquisition unit configured to acquire target information, the target information representing whether a first target object appears in a target acquisition image, wherein the first target object represents a hand, and the target acquisition image is at least one image acquired earlier than an i-th acquisition image, i being an integer greater than 1;

[0011] a second acquisition unit configured to acquire, according to the target information, a first position of the first target object in the i-th acquisition image;

[0012] The detection unit is configured to detect action feature information of the first target object in the i-th frame of the captured image based on the first position.

[0013] The first determination unit is configured to determine, according to the action feature information, whether action information generated by the first target object in the i-th frame of the captured image is a predetermined gesture.

[0014] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein,

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of the present disclosure is provided.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the present disclosure.

[0018] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0020] Figure 1 is an application scenario diagram of the embodiment of the present disclosure;

[0021] Figure 2 is a flowchart of the gesture detection method of the embodiment of the present disclosure Figure 1 ;

[0022] Figure 3 is a flowchart of the gesture detection method of the embodiment of the present disclosure Figure 2 ;

[0023] Figure 4 is a flowchart of the gesture detection method of the embodiment of the present disclosure Figure 3 ;

[0024] Figure 4 is a flowchart of the gesture detection method of the embodiment of the present disclosure Figure 6 ;

[0025] Figure 7 FIG. 1 is a schematic diagram of an overall implementation of a gesture detection method according to an embodiment of the present disclosure;

[0026] Figure 8 FIG. 2 is a schematic diagram of a structure of an arm (or hand) and its region detection network model according to an embodiment of the present disclosure;

[0027] Figure 9 FIG. 3 is a schematic diagram of a structure of a gesture detection network according to an embodiment of the present disclosure;

[0028] Figure 10 FIG. 4 is a schematic diagram of a composition of a gesture detection apparatus according to an embodiment of the present disclosure;

[0029] Figure 1 FIG. 5 is a block diagram of an electronic device for implementing a gesture detection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which various details of the present disclosure are set forth to facilitate an understanding. However, it should be appreciated that the embodiments described herein are merely exemplary and are not intended to limit the scope of the present disclosure. It will be apparent to one of ordinary skill in the art that various changes and modifications can be made thereto without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted herein.

[0031] The gesture detection scheme provided by the present disclosure can detect a predetermined gesture generated by a user. The predetermined gesture can be any gesture operation that can be controlled by a gesture operation in an actual application. For example, a scenario in which a terminal device is controlled by a gesture, or a scenario in which a gesture is required for control in a game. The gesture in the present scheme can be considered as a gesture including a palm and fingers, i.e., the gesture is composed of a palm and fingers. It should be noted that, in actual applications, ordinary users can use gestures for operation, and disabled people such as deaf-mutes can also use gestures for operation. The gesture detection scheme of the present disclosure can not only detect gestures generated by ordinary users, but also detect gestures generated by deaf-mutes. For example, a gesture generated by a deaf-mute to express thanks or good wishes.

[0032] The technical solution provided by the present disclosure is that, for the i-th frame of collected image, whether the action information of the first target object in the i-th frame of collected image is a predetermined gesture is detected in combination with the result of whether the first target object appears in the preceding frame image (target collected image) of the i-th frame of collected image. In other words, the detection of the gesture generated by the user is performed in combination with at least two frames of collected images (for example, the i-th frame of collected image and at least one target collected image collected earlier than the i-th frame of collected image). Compared with the gesture detection in the related art by using a single frame of image, the detection accuracy can be greatly improved. In addition, when the gesture appearing in the i-th frame of collected image is detected, whether the target collected image appears the first target object and its influence are considered, the influence factor is considered, the problem of detection error such as missed detection can be effectively avoided, and the detection accuracy is ensured. In addition, the collected image related to the present disclosure can be collected by using only a common image collection device such as a common camera, without the need for a special image collection device such as a depth camera or an infrared camera, so that the investment cost can be greatly reduced, and the practicability of the scheme is stronger.

[0033] It should be noted that the gesture detection method related to the embodiments of the present disclosure can be applied to terminal devices such as mobile phones, notebook computers, tablet computers, desktop computers, and all-in-one computers. In addition, the gesture detection method can also be applied to devices such as augmented reality (AR) devices and virtual reality (VR) devices, and can also be applied to other devices with gesture detection functions, such as servers and game consoles. Figure 1 An application scenario provided by the embodiments of the present disclosure is shown in the following figure. Figure 1 In the figure, it can be regarded as a scene in which a deaf-mute person handles bank business at a bank counter. Figure 1 In the figure, the terminal device 10 of the bank includes at least one camera, and the camera collects images, such as collecting the sign language action 11 of the deaf-mute person. Usually, a complete sign language action can be divided into multiple sub-actions. For example, the three sub-actions are collected by the first frame, the second frame, and the third frame of image. For each collected frame (collection) of image, whether the predetermined gesture is included in each frame of image can be detected by the technical solution of the embodiments of the present disclosure. In the case where a complete sign language action includes three sub-actions, only the first two sub-actions in the complete sign language action can be detected in the first frame and the second frame, and the sign language action is completed in the third frame, and a complete sign language action can be detected.

[0034] It can be understood that, in so many frames of collected images, the image collected by the camera earlier can be regarded as an image collected earlier in time, such as the i-1th frame, and the image collected later can be regarded as an image collected later in time, such as the i-th frame of collected image. In addition, the gesture generated by the user in the i-th frame of collected image can be any reasonable gesture including fingers and palms. For example,Figure 1 In actual application, the gesture can be a sign language action, so as to realize input of deaf-mute on the input content which the deaf-mute wants to input in the terminal device. If the gesture is a sign language action, the predetermined gesture involved in the present solution can be any reasonable sign language action. If no special description is given, the i-th frame image in the present disclosure is the i-th captured image, and the preceding frame image is the preceding captured image.

[0035] Figure 2 The gestures in the above examples are only specific examples, and any other reasonable gesture is within the coverage of the embodiments of the present disclosure.

[0036] The gesture detection method involved in the embodiments of the present disclosure will be further described below.

[0037] The embodiments of the present disclosure provide a gesture detection method, as shown in Figure 4 The method comprises the following steps:

[0038] S201: obtaining target information, the target information indicating whether a first target object appears in a target captured image, wherein the first target object represents a hand, and the target captured image is at least one image captured before the i-th captured image, i being an integer greater than 1;

[0039] It can be understood that the detection of the gesture can be performed frame by frame. For example, for the i-th frame image, the result of whether the first target object appears in the preceding frame image thereof needs to be obtained. The target captured image is one or more images captured by the camera before the i-th captured image. Preferably, the target captured image is an image adjacent to the i-th captured image in time, such as the i-1-th captured image. The hand includes the palm and the fingers. For each frame image captured by the camera, it is necessary to detect frame by frame whether the predetermined part, i.e., the first target object (such as the palm and the fingers), which can generate the predetermined gesture, exists in the image. In this way, compared with the captured image in the later time, the captured image in the earlier time is necessarily analyzed first to determine whether the predetermined part exists, and the predetermined gesture can exist only when the predetermined part appears in the captured image. The result of whether the first target object appears in the target captured image can be saved after the detection of the first target object in the target captured image. The saved information can be read when the i-th captured image is analyzed.

[0040] S202: obtaining a first position of the first target object in the i-th captured image according to the target information;

[0041] In this step, the target information includes two kinds: the first kind is that the first target object appears in the target acquisition image; the second kind is that the first target object does not appear in the target acquisition image. Regardless of which kind of information, for the ith acquisition image, whether the first target object appears in the image before the acquisition time of the ith acquisition image should be considered to affect the detection of the ith acquisition image. This is because: in actual application, if the image before the acquisition time appears at the predetermined part, it indicates that the user appears at the predetermined part possibly because he / she inputs the gesture, and thus the possibility that the predetermined gesture appears in the ith acquisition image is large. In the present scheme, the influence of the detection result of the preceding image on the detection result of the current frame acquisition image (the ith acquisition image) is considered in the detection process of whether the predetermined gesture exists in the ith acquisition image, and thus the problem of detection error such as missed detection can be effectively avoided.

[0042] S203: detecting the motion feature information of the first target object in the ith acquisition image based on the first position;

[0043] In this step, the image region including only the first target object in the ith acquisition image can be cropped or extracted based on the first position of the first target object in the ith acquisition image, and the detection of the motion feature information is performed from the cropped or extracted image. The extraction here is equivalent to distinguishing the target region (the region including the first target object) and the background region (the region other than the first target object) in the overall image (the ith acquisition image). The motion feature information can be any reasonable information, such as the shape, bending degree, motion trend of the joint used by the user when generating the motion, etc.

[0044] S204: determining whether the motion information generated by the first target object in the ith acquisition image is a predetermined gesture according to the motion feature information.

[0045] This step can be regarded as determining whether the motion information generated by the first target object is a predetermined gesture according to the motion feature information of the first target object in the ith acquisition image. The scheme of determining the motion information generated by the first target object according to the motion feature information of the first target object in the ith acquisition image can be regarded as a scheme of obtaining the motion matching the motion feature information. In actual application, the motion possibly generated by the user and the feature information of the motion can be recorded in advance, and the motion corresponding to the motion feature information is searched from the recorded data, which is the expected motion. In addition, the predetermined gesture can be recorded in advance, and whether the motion information exists in the recorded content is searched in the case of determining the motion information generated by the first target object. If it exists, it is determined that the detected motion information is a predetermined gesture; otherwise, it is determined that the detected motion gesture is a non-predetermined gesture, i = i + 1, and the same process as S201-S204 is performed on the ith+1 acquisition image.

[0046] In the detection scheme of whether the predetermined gesture exists in the i-th frame of the captured image in S201-S204, the detection of whether the predetermined gesture exists in the i-th frame of the captured image is performed in combination with two or more frames of captured images, such as the current frame of captured image and the previous frame of captured image. Compared with the scheme of gesture detection by using only a single frame of image in the related art, the detection accuracy can be improved. In addition, in actual application, the probability of the occurrence of the action in the i-th frame of image is high when the action occurs in the previous frame of image. The influence of the information of whether the first target object occurs in the previous frame of captured image on the detection result of the current frame of captured image (i-th frame of captured image) is considered, the problem of detection error such as missed detection can be effectively avoided, and the detection accuracy can be improved. In addition, the captured image in the scheme is only captured by a common image capturing device such as a common camera, and does not need a special image capturing device such as a depth camera or an infrared camera, so that the investment cost can be greatly reduced, and the universality of the scheme is stronger.

[0047] In an optional scheme, after it is determined in S204 that the action information is the predetermined gesture, the method further includes: identifying the gesture type of the predetermined gesture. Different types of gestures have different meanings. Taking the predetermined gesture as a sign language action as an example, sign language action 1 represents the approval of the opinion of the other party or the agreement of the request of the other party, and sign language action 2 represents “thank you”. The sign language action 1 is to make “ok”, and the sign language action 2 is to hold a fist and then stretch out the thumb, towards the other party, and bend the thumb twice. It can be seen that the technical scheme of the disclosure not only can detect whether the predetermined gesture occurs, but also can identify the gesture type, so that the content of the detection is more abundant and reasonable, and the ease of use of the technical scheme of the disclosure is enhanced, and the technical scheme of the disclosure is easily promoted.

[0048] After it is determined in S204 that the action information is the non-predetermined gesture, the action information is saved for subsequent frames. It is mainly considered that the predetermined gesture can be divided into multiple decomposition actions. The action occurring in a frame can be one of the decomposition actions of a complete gesture, and the complete gesture needs to be combined with other decomposition actions. Or, after it is determined that the action information is the non-predetermined gesture, the action feature information is saved, and the action feature of the action occurring in the subsequent frame is used for identification of a complete gesture.

[0049] In an optional scheme, in the case that the target information is that the first target object occurs in the target captured image, the gesture detection scheme of the disclosure is as follows:

[0050] S301: In the case that the target information represents that the first target object occurs in the target captured image, a second position of the first target object in the target captured image is acquired;

[0051] S302: predicting a first position of the first target object in the i-th frame according to the second position;

[0052] The second position and the predicted first position obtained by the foregoing S301 and S302 can be implemented in a gesture detection process of the target image and the predicted first position can be saved. The saved data can be read when the gesture detection of the i-th frame is performed. In an implementation, the second position can be saved and the first position can be predicted according to the second position in the gesture detection process of the i-th frame. It can be understood that the first position and the second position involved in the present solution are positions of the first target object in different images and do not represent the number of positions, which are concepts proposed to distinguish the positions of the first target object in different images.

[0053] The foregoing solution is to estimate or predict the first position of the first target object in the i-th frame according to the second position of the first target object in the target image. Compared with a solution of directly detecting the position of the first target object in the i-th frame, the position prediction or estimation solution can greatly reduce the detection difficulty and reduce the calculation amount. In actual application, the second position of the first target object in the target image can be expanded to obtain the first position of the first target object in the i-th frame. Exemplarily, in actual application, if the position of the first target object in the target image is located in a central region of the target image and has a size of 5*5 pixels, the position of the first target object in the i-th frame is a central region and has a size of 10*10 pixels.

[0054] The foregoing S301-S302 can be further description of obtaining the first position of the first target object in the i-th frame.

[0055] S303: tracking the first target object from the second position in the target image to the first position in the i-th frame.

[0056] The image tracking of the first target object is performed. The tracking is performed from the target image to the i-th frame. Specifically, the tracking of the first target object is performed at the second position of the target image and the tracking is performed to the first position of the i-th frame.

[0057] S304: detecting action change feature information generated when the first target object is tracked from the target image to the i-th frame; and taking the action change feature information as action feature information of the first target object in the i-th frame.

[0058] Detect the action change feature of the first target object in the tracking process, such as the change of the hand action from a posture to another posture. The action feature of the first target object in the target acquisition image can be detected and saved in the process of gesture detection on the target acquisition image. In the process of gesture detection on the current frame (i-th frame) acquisition image, the action feature of the first target object in the previous frame acquisition image is obtained by reading the saved data. The change of the first target object from the action posture and the action bending degree in the previous frame acquisition image to the action posture and the action bending degree in the current frame acquisition image can be regarded as the action change feature information of the first target object from the target acquisition image to the i-th frame acquisition image.

[0059] S305: According to the action change feature information, determine the change action generated by the first target object from the target acquisition image to the i-th frame acquisition image, and determine whether the change action is a predetermined gesture.

[0060] In practical applications, the corresponding relationship between the action feature of each decomposed action in the predetermined gesture generated by the user (the action feature of the decomposed action 1 + the action feature of the decomposed action 2) and the change action (the decomposed action 1 + the decomposed action 2) can be recorded in advance. For example, if the action feature of the decomposed action 1 is detected in the i-1-th frame acquisition image, and the action feature of the decomposed action 2 is detected in the i-th frame acquisition image, according to the record of the corresponding relationship, the change action generated from the target acquisition image, such as the i-1-th frame acquisition image, to the i-th frame acquisition image is the decomposed action 1 + the decomposed action 2.

[0061] In practical applications, the predetermined gesture that the user can generate and the change action generated from the generation to the end of the predetermined gesture can be recorded in advance. When in use, it is checked whether there is a predetermined gesture corresponding to the detected change action in the recorded data. If there is, it means that the detected change action is a predetermined gesture, and if there is not, it means that it is not a predetermined gesture.

[0062] In S301-S305, when the first target object appears in the previous frame acquisition image, the processing flow for the current frame (i-th frame) acquisition image is based on the position tracking of the first target object in the current frame acquisition image and in the previous frame acquisition image. The action change feature of the first target object in at least two frames of acquisition images can be obtained, the detected change action is determined based on the action change feature information, and it is determined whether the change action is a predetermined gesture, which can greatly guarantee the accuracy of gesture detection and avoid missed detection.

[0063] In which, based on the action change feature, the predetermined gesture is detected, the recognition of the predetermined gesture with multiple decomposed actions can be realized, and the recognition accuracy can be guaranteed.

[0064] In practical application, the scheme can detect not only a gesture as a single action, but also a (predetermined) gesture with multiple decomposed actions. The image acquisition in the scheme can be performed by a common image acquisition device such as a common camera, and thus has strong versatility.

[0065] It can be understood that, in practical application, if the i-th acquired image is clear enough, the first position predicted or estimated based on the i-th acquired image can be tracked to the first target object. If the i-th acquired image is not clear enough or the gesture has ended, there is a case that the first target object cannot be tracked to the first position in the i-th acquired image. In this case, the number of acquired images in which the first target object is not tracked to the first target object consecutively from the target acquired image to the first target acquired image is obtained, specifically calculated; it is determined whether the number is greater than or equal to a predetermined threshold; if the number is greater than or equal to the predetermined threshold, the predicted first position is cleared or deleted. If the number is less than the predetermined threshold, the predicted first position is saved or retained. That is, in the case where the number is calculated, it is determined whether the number of acquired images in which the first target object is not tracked consecutively is greater than or equal to the predetermined threshold. If the number is less than the predetermined threshold, it is indicated that the first target object is not tracked in the i-th acquired image because the acquired image is not clear enough, and the first target object can be tracked in the next frame, so the first position needs to be saved for tracking. If the number is greater than or equal to the predetermined threshold, such as two frames, it is indicated that the first target object is not tracked consecutively for multiple frames, and the gesture generated by the user has ended, so the first position is cleared or deleted. For the i+1-th acquired image, the position of the first target object is detected from the i+1-th acquired image. In addition, based on the comparison between the number and the predetermined threshold, it can be determined whether to update the acquisition period used by the image acquisition device such as a camera when acquiring images. The acquisition period is used to determine the acquisition interval between the target acquired image and the i-th acquired image. Further, if the calculated number is greater than or equal to the threshold, the acquisition period is increased, specifically lengthened, such as from originally acquiring one frame of image every 3 minutes to acquiring one frame of image every 5 minutes. It can be understood that, under the increased acquisition period, the acquisition interval between adjacent two frames of acquired images acquired by the image acquisition device is increased. If the calculated number is less than the threshold, the acquisition period does not need to be updated, and the image acquisition continues to be performed according to the original acquisition period.

[0066] The foregoing scheme can avoid the problem of missed detection of a gesture due to occasional unclear acquired images, and improve the detection accuracy.

[0067] Based on the comparison result between the calculated number and the threshold, it can be determined whether to change the acquisition period of the image acquisition device, which can also avoid power consumption caused by image acquisition of the first target object when the gesture has ended, and avoid the problem that the power cannot be saved caused by frequent image acquisition.

[0068] In an alternative, the target information of the first target object not appearing in the target acquisition image is collected, such as Figure 5 As shown in the method comprises:

[0069] S401: determining whether the first target object appears in the i-th acquisition image;

[0070] If the determination is yes, continue to execute S402;

[0071] If the determination is no, the processing flow of the i-th acquisition image ends, i=i+1, and the same processing flow as the i-th acquisition image is executed for the i+1-th acquisition image.

[0072] It can be understood that, if the target acquisition image is an image adjacent to the i-th acquisition image and collected before the i-th acquisition image, such as the i-1-th acquisition image, and the first target object does not appear in the image, it indicates that the last gesture action generated by the user has ended. The predetermined gesture can appear in the i-th acquisition image. Because the gesture is generated by the palm and fingers, it is necessary to determine whether the first target object appears in the i-th acquisition image. Those skilled in the art can understand that, if the first target object appears in the i-th acquisition image, it does not mean that the predetermined gesture must appear in the image. For example, in actual application, taking the predetermined gesture as a predetermined sign language action, the action of the palm and fingers appearing in the image can be an unconscious action, and does not represent the sign language operation. This determination scheme can make the gesture detection scheme more rigorous and meet the actual application.

[0073] S402: detecting a first position of the first target object in the i-th acquisition image;

[0074] In this step, the contour of the first target object appearing in the i-th acquisition image is detected, and the contour region can be used as the first position of the first target object in the i-th acquisition image.

[0075] S403: detecting an action feature information of the first target object in the first position in the i-th acquisition image;

[0076] In this step, the region image of the first position in the i-th acquisition image is cropped or extracted, and the action feature of the first target object is detected from the cropped or extracted image. The action feature includes but is not limited to the posture of the action, the bending degree of the fingers and the palm in the action, etc.

[0077] S404: determining whether the action information generated by the first target object is a predetermined gesture according to the action feature information.

[0078] In this step, the action information generated by the first target object is determined according to the action feature information. The foregoing can be regarded as obtaining a scheme of the action matching the action feature information. In actual application, the action possibly generated by the user can be recorded in advance in correspondence with the feature information of the action, and the action corresponding to the action feature information in the recorded data is searched to obtain the expected action. The predetermined gesture possibly generated by the user can also be recorded in advance in correspondence with the gesture action corresponding to the predetermined gesture, and when used, it is checked whether the predetermined gesture corresponding to the detected action exists in the recorded data. If the predetermined gesture exists, it indicates that the detected action is the predetermined gesture, and if the predetermined gesture does not exist, it indicates that the detected action is not the predetermined gesture.

[0079] In S401-S404, in a case where the first target object does not appear in the target collection image, it is determined whether the first target object appears in the i-th collection image, in a case where the first target object appears, the position of the first target object in the i-th collection image is detected, the action feature of the first target object at the position in the i-th collection image is detected, and whether the action generated by the user in the i-th collection image is the predetermined gesture is determined according to the action feature. This scheme takes into account the case where a single collection image also appears in actual application. The predetermined gesture is stronger in practicability.

[0080] In the foregoing scheme, S402 can detect the first position of the first target object in the i-th collection image by the following manner: detecting whether a second target object appears in the i-th collection image, the second target object at least including an arm; in a case where the second target object appears in the i-th collection image, detecting the position of the second target object in the i-th collection image; and detecting the position (the first position) of the first target object in the i-th collection image based on the position of the second target object in the i-th collection image. It can be understood that if the color of the background part of the collection image is similar to the color of the hand, such as an image of a hand holding a face, the background image is the face, and the color of the hand is similar to the skin color. In this image, the contour boundary of the hand is not clear enough. It is difficult to accurately detect the position of the hand in the image from this image. However, the arm is connected to the hand, and the arm and the face have a large difference in contour. Thus, the position of the arm in the image, that is, the arm region, can be detected first, and then the hand in the image is detected from the arm region and the region near the arm region. Thus, the appearance of the hand in the image can be accurately detected.

[0081] In the case that the second target object does not appear in the i-th captured image, which is the case that the image capturing device only captures the hand and does not capture the arm in the actual life. In order to avoid missing detection of the hand, the first target object can be directly detected from the i-th captured image (the whole image). If the first target object is detected from the i-th captured image, the position information of the first target object in the i-th captured image is detected. If the first target object is not detected from the i-th captured image, i=i+1, and the same processing procedure as the i-th captured image is performed on the (i+1)-th captured image. In the case that the second target object does not appear in the i-th captured image, the scheme of directly detecting the hand from the whole captured image can avoid missing detection of the hand, further avoid missing detection of the gesture, and further ensure the accuracy of the detection.

[0082] The foregoing scheme is a scheme for processing the i-th captured image, i being a positive integer greater than 1. It can be understood that for the 1st captured image, at least one of the following procedures can be used for processing the 1st captured image:

[0083] Scheme one: detecting whether the first target object appears in the 1st captured image; if the first target object appears, detecting the position of the first target object in the 1st captured image; based on the position of the first target object in the 1st captured image, detecting the action feature information of the first target object in the 1st captured image; and determining whether the action information generated by the first target object is a predetermined gesture according to the action feature information of the first target object in the 1st captured image.

[0084] Scheme two: detecting whether the second target object appears in the 1st captured image, the second target object including the arm; if the second target object appears in the 1st captured image, detecting the position of the second target object in the 1st captured image; based on the position of the second target object in the 1st captured image, detecting whether the first target object appears in the 1st captured image. If the first target object appears, detecting the position of the first target object in the 1st captured image; based on the position of the first target object in the 1st captured image, detecting the action feature information of the first target object in the 1st captured image; and determining whether the action information generated by the first target object is a predetermined gesture according to the action feature information of the first target object in the 1st captured image.

[0085] It can be understood that the scheme shown in scheme one is to directly detect whether the first target object appears in the image collected from the first frame. The scheme shown in scheme two considers the problem that if the background part of the collected image has a part with the same color as the hand, such as the face, the hand is not easy to detect. The hand is detected in the arm region and the nearby region if the arm can be detected. In this way, the problem that the hand cannot be directly detected can be avoided, and the detection accuracy is further ensured.

[0086] The gesture detection method of the present disclosure can also be implemented by at least two trained detection models, such as the first detection model and the second detection model. Now the technical scheme of the gesture detection method implemented by at least two detection models will be described in general. The two detection models described above can be neural network models or deep neural network models.

[0087] As shown in Figure 7 , the method comprises:

[0088] S501: For the i-th frame of collected image, obtain target information indicating whether the first target object appears in the target collected image, wherein the first target object represents a hand, and the target collected image is at least one image collected earlier than the i-th frame of collected image;

[0089] S502: According to the target information, input the i-th frame of collected image into the trained first detection model to obtain the position (first position) of the first target object in the i-th frame of collected image;

[0090] The first detection model in this step is the same as the model shown in Figure 8 . The position of the hand in the i-th frame of collected image is detected by the model. Because the first detection model has strong robustness and stability, the detection of the position of the hand by the detection model with strong robustness and stability can greatly improve the detection accuracy.

[0091] Further, in the case that the target information indicates that the target collected image appears the first target object, the image region of the first target object in the i-th frame of collected image predicted for the i-th frame of collected image is input into the first detection model to obtain the position of the first target object in the i-th frame of collected image. In the case that the target information indicates that the target collected image does not appear the first target object, the i-th frame of collected image is input into the first detection model to obtain the position of the first target object in the i-th frame of collected image.

[0092] The first detection model is obtained by taking multiple frames of collected images as training samples and training the position of the first target object in the multiple frames of collected images. The training process is not described in detail in the present disclosure.

[0093] S503: Obtain an image including only the first target object in the i-th frame of the collected image based on the position of the first target object in the i-th frame of the collected image.

[0094] In this step, the region image corresponding to the position in the i-th frame of the collected image can be extracted or cropped to obtain a hand image including only the hand. That is, the hand image in the i-th frame of the collected image is extracted or cropped.

[0095] S504: Input the image including only the first target object in the i-th frame of the collected image into the second detection model to obtain the action feature information of the first target object in the i-th frame of the collected image by the second detection model and identify whether the action generated by the first target object is the result of the predetermined gesture based on the action feature information.

[0096] In this step, the second detection model is the same as the model shown in Figure 8 . The hand image is input into the second detection model to obtain the action feature information in the i-th frame of the image and identify the detection result of whether the action generated by the first target object is the predetermined gesture based on the action feature information. Further, the second detection model at least detects the action feature information of the first target object in the i-th frame of the collected image (implemented by the feature extraction network in Figure 8 ). According to the action feature information of the first target object in the i-th frame of the collected image, the action information generated by the first target object is determined, and whether the action information is the predetermined gesture is determined (implemented by the time sequence feature network and the classification network in Figure 7 ). After the second detection model determines that the action information is the predetermined gesture, the second detection model further detects the gesture type of the predetermined gesture. The detection of the action feature information by the second detection model with robustness and the detection of the predetermined gesture based on the action feature information can improve the detection accuracy.

[0097] In S501-S504, whether the action information generated by the first target object in the i-th frame of the collected image is the predetermined gesture is detected in combination with the result of whether the first target object appears in the previous frame image (target collected image) of the i-th frame of the collected image. Compared with related technologies, the detection of the first position by the first detection model can improve the detection accuracy.

[0098] In the foregoing scheme, a dedicated image collection device is not required, which reduces the investment cost. In addition, two detection models are used for gesture detection. Because the first and second detection models have strong robustness and stability, the use of two detection models with strong robustness and stability for gesture detection can greatly improve the detection accuracy. In the scheme of gesture detection using two detection models, the collected image used can be collected by a normal camera, without the need for a special camera, which is more practical.

[0099] For the two cases that the first target object appears in the target acquisition image and does not appear in the target acquisition image, the following are described respectively.

[0100] Case one: in the case that the target information indicates that the first target object appears in the target acquisition image, the i-th acquisition image is input into the first detection model to obtain the position (first position) of the first target object in the i-th acquisition image. Based on the first position, the first image including only the first target object in the i-th acquisition image is obtained. The first image and the obtained second image are input into the second detection model. The second detection model detects whether the change action feature of the first target object from the target acquisition image to the i-th acquisition image is a predetermined gesture and whether the change action with the change action feature is the predetermined gesture. The second image can be read from the saved data in the processing flow of the i-1-th acquisition image. In the processing flow of the i-1-th acquisition image, the second image is obtained by inputting the target acquisition image into the first detection model to obtain the second position of the first target object in the target acquisition image, and obtaining the second image including only the first target object in the target acquisition image based on the second position.

[0101] In case one, the positions of the first target object in the i-th acquisition image and in the target acquisition image are obtained by the first detection model with good robustness and stability, which can ensure the accuracy of position detection. The detection result of whether the change action of the first target object from the target acquisition image to the i-th acquisition image is a predetermined gesture is obtained by the second detection model with good robustness and stability, which can make the detection result more accurate. In addition, compared with inputting the entire images, i.e., the i-th acquisition image and the target acquisition image, into the second detection model, by inputting part of the images, i.e., the images including only the hand in the i-th acquisition image and the target acquisition image, such as the first image and the second image, into the second detection model, the calculation amount can be greatly reduced and the detection time can be shortened in the detection process.

[0102] It can be understood that in case one, the position of the first target object in the i-th acquisition image is obtained by detection of the first detection model. In actual application, there are alternative schemes for obtaining the position information, such as predicting the position information of the first target object in the i-th acquisition image according to the position information of the first target object in the target acquisition image. For specific prediction schemes, please refer to the relevant description. This alternative scheme does not need to be processed by the first detection model, but can obtain the position information of the first target object in the i-th acquisition image, which is simple and easy to implement, can reduce the calculation amount of position detection, and can speed up the detection.

[0103] In case two, the i-th frame of the collected image is input into the first detection model to obtain the position (first position) of the first target object in the i-th frame of the collected image. Based on the position, a region image including only the first target object in the i-th frame of the collected image is obtained. The region image is input into the second detection model, and the second detection model obtains the action feature of the first target object in the i-th frame of the collected image based on the region image and determines whether the action of the first target object is the predetermined gesture based on the action feature.

[0104] In case two, the position of the first target object in the i-th frame of the collected image is obtained by the first detection model with good robustness and stability, which can ensure the accuracy of the position detection. The action feature of the first target object in the i-th frame of the collected image and whether the action with the action feature is the predetermined gesture are obtained by the second detection model with good robustness and stability, which can make the detection result more accurate. In addition, compared with inputting the entire image, i.e., the i-th frame of the collected image, into the second detection model, inputting the partial image, i.e., the image including only the hand in the i-th frame of the collected image, into the second detection model can greatly reduce the calculation amount and shorten the detection time in the detection process.

[0105] In the foregoing scheme, the scheme of inputting the i-th frame of the collected image into the first detection model to obtain the position of the first target object in the i-th frame of the collected image can be implemented in the following manner:

[0106] First, the i-th frame of the collected image is input into the third detection model. The third detection model is used to detect whether the second target object appears in the i-th frame of the collected image. The second target object at least includes an arm. In the case where the second target object appears in the i-th frame of the collected image, the position of the second target object in the i-th frame of the collected image is obtained by the third detection model. A region image corresponding to the position and the position nearby in the i-th frame of the collected image is obtained. The region image is input into the first detection model to obtain the position information of the first target object in the i-th frame of the collected image. The third detection model is the same as the model shown in FIG. 1. The arm part is detected by the third detection model with good robustness and stability, which can ensure the detection accuracy and further provide a certain guarantee for accurately detecting the position of the first target object in the i-th frame of the collected image. In addition, the hand position in the image is detected based on the detection result of whether the arm position appears, which can avoid the problem of hand detection error caused by the same background color and hand color (for example, the hand is placed on the face). Figure 6

[0107] ​In the foregoing scheme, if the detection result obtained through the detection of the third detection model is that the second target object does not appear in the ith frame of collected images, the entire collected image, such as the ith frame of collected images, is input into the first detection model to obtain a result of whether the first target object can be detected in the ith frame of collected images, so as to avoid the missed detection of the hand. If the first target object can be detected, the ith frame of collected images is input into the first detection model to obtain the position information of the first target object in the ith frame of collected images; if the first target object cannot be detected, i = i + 1, and the i + 1th frame of collected images performs a similar processing procedure as the ith frame of collected images.

[0108] The technical solutions of the embodiments of the present disclosure will be further described below in combination with Figure 7 、 Figure 8 and Figure 6 .

[0109] In combination with the procedure shown in Figure 7 , the camera collects images at a collection period. For a current frame of images (such as the ith frame of collected images) collected by the camera (S601), it is obtained whether a hand appears in a previous frame of images, such as the i-1th frame of collected images. Wherein, whether a hand appears in the i-1th frame of collected images is a result obtained in the procedure of hand detection of the i-1th frame of collected images, and for the convenience of the use of subsequent frames, the result is saved. In the procedure of gesture detection of the i-1th frame of collected images, if a hand action appears in the i-1th frame of collected images but the action is not a predetermined gesture, the action features of the action can be saved, and the position of the hand appearing in the next frame of images, i.e., the ith frame of collected images, is predicted according to the position of the hand appearing in the i-1th frame of collected images, so as to facilitate the gesture detection of subsequent frames. Wherein, the process of prediction can be regarded as: expanding a certain area outward around the position of the hand appearing in the i-1th frame of collected images. Exemplarily, the position of the hand appearing in the i-1th frame of collected images is located in the central region of the image, and the size is 5*5 pixels, and the position of the hand predicted to appear in the ith frame of collected images is the central region, and the size is 10*10 pixels. In the procedure of gesture detection of the ith frame of collected images, each result obtained in the procedure of processing of the i-1th frame of collected images needs to be read.

[0110] Now, in combination with the result of whether a hand appears in the i-1th frame of collected images, the present scheme is described in detail in two processing procedures.

[0111] Processing procedure 1:

[0112] In the processing flow of the i-th frame of the captured image, if no hand appears in the (i-1)-th frame, it indicates that the previous predetermined gesture has ended. A new predetermined gesture may or may not be generated in the i-th frame, depending on the subsequent detection results. Considering that the accuracy of directly detecting the hand from the captured image is poor when there are background parts with similar colors to the hand, such as a face, in the captured image, the present solution first detects whether an arm appears in the i-th frame (S602), and if an arm appears, it needs to detect the arm region in the i-th frame. This process can be implemented by an arm and its region detection network model, which can be considered a third detection model. The composition of this network model is similar to... Figure 7 The model shown is similar, including a feature extraction network, a proposal generation network, a pooling layer, and a classification network. The pooling layer has two inputs: one connected to the feature extraction network and the other to the proposal generation network. A regression network is connected in parallel with the classification network. In its implementation, the i-th frame of the captured image is input to the feature extraction network, which extracts image features such as contours, edges, colors, and textures. These extracted features are then input to the proposal generation network, which uses these features to detect useful parts, such as the arm, in the input image (i.e., the i-th frame) and outlines them with bounding boxes. The outlined image containing only the arm, along with the image features extracted by the feature extraction network, is then input to the pooling layer. The pooling layer performs dimensionality reduction on its input data, reducing the computational load. The pooled image containing only the arm is then input to both the classification and regression networks. The classification network detects whether the input image it receives contains an arm. If it detects that an arm is present, the regression network can detect the position of the arm in the input image. If it detects that the image does not contain an arm, the regression network cannot detect the position of the arm in the image.

[0113] exist Figure 7 In specific implementations, the feature extraction network can be a Convolutional Neural Network (CNN). This network includes multiple convolutional layers, which extract image features through convolutional operations. The proposal generation network can be a Region Proposal Network (RPN). The classification network can be implemented using a classifier such as a softmax function classifier. The regression network can be implemented using location fitting methods or sliding window methods.

[0114] If an arm is detected in the i-th frame of the captured image, it is also necessary to determine whether a hand appears in the image. If a hand appears, a predetermined gesture may occur. If no hand appears, the predetermined gesture will not occur. Further, from the arm and its surrounding area in the i-th frame of the captured image, the presence of a hand is detected (S603), and if a hand is present, its position in the image is detected (S604). This process can be implemented by a hand and its region detection network model (first detection model), the composition of which is similar to... Figure 8 The model shown is similar. In its implementation, the image corresponding to the arm and its surrounding area in the i-th frame is input into the feature extraction network. Compared to directly inputting the entire i-th frame image into the feature extraction network, inputting a partial image, such as the image corresponding to the arm and its surrounding area in the i-th frame, reduces computation. The feature extraction network extracts image features such as contours, edges, colors, and textures from the received images. The extracted image features are input into the proposal box generation network, which uses the extracted image features to detect useful parts in the input image, such as the hand, and circles them with boxes. The circled image containing only the hand and the image features extracted by the feature extraction network are both input into the pooling layer, which performs dimensionality reduction on its own input data to reduce subsequent computation. The image containing only the hand after pooling is input into the classification network and the regression network, respectively. The classification network detects whether the received image contains a hand; if it does, the regression network detects the location of the hand in the received image.

[0115] If no arm is detected in the i-th frame of the image, to avoid missed detections, the detection of whether a hand exists is directly performed from the i-th frame – the entire image (S605). The i-th frame is input into the aforementioned hand and region detection network model. Through the processing of each component in the model, the detection result of whether a hand exists in the i-th frame is obtained. The processing procedure is the same as described above and will not be repeated. If a hand is detected, the detection result also includes the position of the hand in the i-th frame. If no hand exists, then i = i + 1, and the next frame, i.e., the (i+1)-th frame, is processed. This processing procedure is the same as the process for processing the i-th frame and will not be repeated. Compared with the scheme of detecting whether a hand exists from the i-th frame – the entire image, detecting the hand from a portion of the image, i.e., the arm and its surrounding area in the i-th frame, not only reduces the amount of detection computation but also increases the detection accuracy.

[0116] If the hand is detected in the i-th frame, the position of the hand in the i-th frame is extracted as a hand image (S606). Because the hand image and the image involved in the aforementioned arm and hand detection scheme are all color images, the calculation amount of the color image is greater than that of the black and white image. Therefore, the hand image is binarized (S607) to effectively reduce the calculation amount and speed up the entire detection process. The image is converted from color to black and white, and the hand in the black and white image is tracked, and the motion features of the hand such as the motion posture, the bending degree, and the motion direction are extracted. According to the position of the hand in the i-th frame, the possible position of the hand in the next frame, i.e., the i+1-th frame, is predicted. The specific prediction process is described above.

[0117] It can be understood that because the convolutional network has strong robustness and stability, the convolutional network is used to detect whether the hand appears in the image and to detect the position of the hand in the image, which can greatly ensure the detection accuracy.

[0118] Similarly, the arm and the region detection network model are also a convolutional network. The network with strong robustness is used to detect the arm and the position of the arm, which can also ensure the detection accuracy and provide a good foundation for more accurate detection of the hand.

[0119] Process 2:

[0120] In the processing flow of the i-th frame, if the hand appears in the i-1-th frame, the position of the hand in the i-th frame predicted in the processing flow of the i-1-th frame is read (S609), the image corresponding to the position region in the i-th frame is extracted and binarized, the image is converted from a color image to a black and white image, the hand in the black and white image is tracked (S608), and the motion features of the hand are extracted.

[0121] It can be understood that the predicted position is an estimated position, and tracking using the estimated position often has limited accuracy. In the processing flow of the i-th frame of the captured image in the present disclosure, after the image corresponding to the region of the predicted position in the i-th frame of the captured image is extracted, the extracted image can be input into the aforementioned hand and region detection network model to obtain a more accurate position of the hand in the i-th frame of the captured image. The scheme of inputting the extracted image into the hand and region detection network model to obtain the accurate position is described in the related description and will not be repeated here. In the scheme of tracking the hand in the i-th frame of the captured image, the preliminary tracking position in the i-th frame of the captured image can be the position predicted in the processing flow of the i-1-th frame of the captured image. As the actual position of the hand in the i-th frame of the captured image is detected, the tracking of the hand will be performed on the actual position. In this way, accurate tracking of the hand in the image position can be achieved, and preparation for better extraction of action features is made.

[0122] It can be understood that if the hand appears in both the i-1-th frame of the captured image and the i-th frame of the captured image, the tracking of the hand will be from the i-1-th frame of the captured image to the i-th frame of the captured image. The action feature information of the hand appearing in the i-1-th frame of the captured image is read, or the action feature of the hand appearing in the i-1-th frame and the i-2-th frame of the captured image is read, i.e., the action feature information of the hand appearing in at least one previous frame of the image is read. The extraction of the action feature of the hand in the i-th frame of the captured image and all previous frames of the image is achieved by the feature extraction network as described in Figure 8 The action feature of the hand appearing in at least one previous frame of the image and the action feature of the hand in the i-th frame of the captured image are input into the feature extraction network as described in Figure 8The gesture detection network shown can be regarded as a second detection network. The network mainly includes a feature extraction network, a time sequence neural network and a classification network. It can be understood that the data input to the gesture detection network is the image of the position of the hand in the i-th frame of collected images, the feature extraction network extracts the features of the input image itself to obtain the action features of the hand in the i-th frame of collected images. By reading the saved data, the action features of the hand in the i-1-th frame of collected images and other previous frame images are obtained, and the read action feature information and the action features of the hand in the i-th frame of collected images are input to the time sequence neural network. It can be understood that the data input to the time sequence neural network is a sequence of action features, and the time sequence neural network can detect the action change features of the hand from the i-1-th frame of collected images to the i-th frame of collected images based on the sequence of action features, such as changing from the action posture of a certain predetermined gesture decomposition action 1 to the action posture of a decomposition action 2, and detecting the change of the action from the i-1-th frame of collected images to the i-th frame of collected images based on the action change features, such as changing from the decomposition action 1 to the decomposition action 2. The classification network can calculate the probability that the continuous decomposition action (decomposition action 1 + decomposition action 2) generated by the user from the collection of the i-th frame of collected images to the collection of the i-th frame of collected images is a predetermined gesture according to the action change detected by the time sequence neural network. If the calculated probability value is greater than or equal to the preset probability threshold, it is considered that the user has generated a predetermined gesture. If the calculated probability value is less than the preset probability threshold, it is considered that the user has generated a non-predetermined gesture, which can be one of the decomposition actions of the predetermined gesture, and the user has not finished generating the predetermined gesture, so the action features of the non-predetermined gesture action are saved to facilitate the detection of the predetermined gesture by combining the action features of the subsequent frames. Alternatively, the probability of each predetermined gesture of the generated continuous decomposition action is calculated. Such calculation not only obtains the result of whether it is a predetermined gesture, but also obtains the result of which type of gesture the predetermined gesture is. That is, the type of the predetermined gesture generated by the user is recognized (S610). For example, the predetermined gesture is recognized as a "thank you" gesture or a "hello" gesture made by a deaf-mute person. In the foregoing scheme, the non-predetermined gesture action detected in a certain frame can also be saved, and the saved gesture action can be combined with the action detected in the next frame to detect whether a predetermined gesture is generated.

[0123] In a specific implementation, Figure 7 The feature extraction network in the foregoing embodiment can be a CNN; the time sequence neural network can be a long short-term memory network (LSTM); and the classification network can be a classifier such as a softmax classifier.

[0124] It should be noted that the second detection model is a neural network, because the neural network has strong robustness and stability, and the accuracy of using the neural network to detect whether it is a predetermined gesture is greatly improved. At the same time, the action feature information generated in the multiple frames of images is combined to identify the predetermined gesture, which can ensure the identification accuracy.

[0125] To distinguish Figure 8 and Figure 7 two feature extraction networks and two classification networks, the feature extraction network in Figure 8 can be regarded as a first feature extraction network, and the classification network can be regarded as a first classification network. The feature extraction network in Figure 1 can be regarded as a second feature extraction network, and the classification network can be regarded as a second classification network.

[0126] It should be noted that in the processing flow of the ith frame of captured image, if there is a hand in the (i-1)th frame of captured image, the position of the hand predicted in the processing flow of the (i-1)th frame of captured image in the ith frame of captured image can be read, and the image corresponding to the position region in the ith frame of captured image can be extracted. If the hand is not tracked in the extracted image, which is equivalent to that the hand does not appear in the ith frame of captured image, the reason for this situation can be that the ith frame of captured image is not clear and the hand cannot be detected. Another reason can be that the predetermined gesture generated by the user has already ended and no new gesture appears in the ith frame of captured image. The following scheme can be used to distinguish the two situations: in the processing flow of the ith frame of captured image, if the hand is not tracked in the predicted position, it is judged whether the re-detection condition is reached (S611), i.e., whether the number of captured images in which the hand is not tracked consecutively is greater than or equal to a predetermined threshold. If it is less than the threshold, it means that the hand in the (i-1)th frame of captured image can be tracked but the hand in the ith frame of captured image is not tracked, which is caused by the unclear captured image. i = i + 1, the next frame, i.e., the (i+1)th frame of image, can still track the hand, and the position of the first target object predicted in the ith frame of captured image is saved to facilitate subsequent tracking. If it is greater than or equal to the predetermined threshold, e.g., two frames, i.e., the (i-1)th and ith frames of captured image, do not track the hand, which means that the gesture generated by the user has already ended and no new gesture is generated, the position of the hand predicted in the processing flow of the (i-1)th frame of captured image in the ith frame of captured image can be cleared or deleted (S612). Wherein, if two frames consecutively do not track the hand, it means that the last gesture generated by the user has already ended and no new gesture is generated, the image capturing can be stopped. Alternatively, the capturing period used when the camera captures the image can be adjusted, specifically increased, e.g., from 1 frame of image per 1 minute to 1 frame of image per 2 minutes. It can be understood that the increase of the capturing period actually adjusts, specifically increases, the time interval of the camera capturing the image. This scheme of increasing the capturing time interval can avoid the problem that the camera frequently captures but the user actually does not generate a new gesture, resulting in heavy resource processing burden.

[0127] The foregoing scheme of calculating the number of captured images in which the hand is not tracked consecutively can avoid the problem of missed detection of gestures caused by occasional unclear captured images, and improve the detection accuracy. It can also avoid unnecessary problems caused by frequent capturing of the hand when the gesture has already ended.

[0128] In Figure 9In the application scenario shown, in the application scenario of handling bank business, for the speech "Your business has been handled" spoken by the bank business personnel, the bank business handler-deaf-mute A makes a hand gesture action of "thank you". It can be understood that the hand gesture action of "thank you" is to hold a fist and then stretch out the thumb, and the thumb is bent twice towards the other party, that is, the hand gesture action includes the following sub-actions: holding a fist, stretching out the thumb, and bending the thumb twice towards the other party. The camera of the bank terminal collects multiple frames of images, and detects the type of hand gesture made by the deaf-mute based on analysis of the collected multiple frames of images. It is assumed that three frames of images are collected, among which the first frame collects an image of a fist, the second frame collects an image of a stretched-out thumb, and the third frame collects an image of the thumb bent twice towards the other party. Starting from the first frame, the action is detected frame by frame using the foregoing scheme, and after detecting the action appearing in the third frame, the deaf-mute is detected to be making the hand gesture action of "thank you" together with the actions appearing in the first frame and the second frame. After the bank terminal detects the hand gesture action, the meaning represented by the action is made into text data, and the text data is converted into audio data through text conversion to audio technology such as TTS technology. The bank terminal plays the audio data representing "thank you" to realize normal communication between the deaf-mute and ordinary people, greatly facilitating the deaf-mute and improving user experience.

[0129] In another application scenario, such as an AR or VR game scenario, gesture actions are used to realize control of an object to be controlled in an AR or VR device, such as a simulated person displayed on a game interface. It is assumed that friendly forces appear in front of the simulated person, and the game player will make a gesture of indicating that the friendly forces come to his side to discuss countermeasures. It is assumed that the gesture representing this meaning is to open the hand and bend the fingers towards the friendly forces. That is, the gesture includes the following sub-actions: opening the hand and bending the fingers towards the friendly forces. The camera of the AR or VR device collects multiple frames of images, and detects the gesture made by the game player based on analysis of the collected multiple frames of images. It is assumed that two frames of images are collected, among which the first frame collects an image of an open hand, and the second frame collects an image of the fingers bent towards the friendly forces. Starting from the first frame, the action is detected frame by frame using the foregoing scheme, and after detecting the action appearing in the second frame, the game player is detected to be making the gesture of indicating that the friendly forces come to his side together with the action appearing in the first frame. The AR or VR device controls the simulated person to make the same gesture. The consistency of the game player and his simulated person in the game interface is realized, and the game experience is improved.

[0130] The technical scheme of the embodiments of the present disclosure has at least the following advantages:

[0131] Firstly, by detecting the arm first and then detecting the hand, the problem of hand detection error caused by the same background color and hand color (for example, the hand is placed on the face) is avoided; at the same time, the hand detection directly on the collected image is supported, the scheme is reliable and easy; the accuracy of the detection result is improved.

[0132] Secondly, the scheme of detecting the hand and the palm and the position thereof and the scheme of tracking the hand are separated and do not affect each other, which embodies convenience. Among them, the position area of the hand predicted in the previous frame in the next frame image can be taken as the preliminary tracking position of the hand in the next frame image. Compared with directly detecting the actual position of the hand in the next frame image from the entire image and tracking, the preliminary tracking position provides an approximate position of the hand in the image, which provides a certain basis for detecting the actual position. Compared with detecting the position from the entire image, detecting the position from part of the image (the image corresponding to the predicted position area of the hand in the i-th collected image) can greatly reduce the calculation amount.

[0133] Thirdly, when detecting whether the action is a predetermined gesture, the actions of the hand appearing in multiple images are considered, and the actions are combined to detect the predetermined gesture, which can make the result more accurate.

[0134] Fourthly, the technical scheme of the embodiment of the disclosure can not only detect a single predetermined gesture, but also can detect multiple continuous predetermined gestures.

[0135] Fifthly, the setting of the re-detection condition allows some frames to not detect the hand, and the gesture detection has greater fault tolerance.

[0136] The disclosure provides a gesture detection device embodiment, as shown in Figure 10 The device comprises:

[0137] A first acquisition unit 901 is configured to acquire target information, wherein the target information represents whether a first target object appears in a target collected image, the first target object represents a hand, and the target collected image is at least one image collected at a time earlier than an i-th collected image, wherein i is an integer greater than 1.

[0138] A second acquisition unit 902 is configured to acquire position information of the first target object in the i-th collected image according to the target information.

[0139] A detection unit 903 is configured to detect action feature information of the first target object in the i-th collected image based on the position information.

[0140] The first determining unit 904 is used to determine, based on the action feature information, whether the action information generated by the first target object in the i-th frame of the captured image is a predetermined gesture.

[0141] The second acquisition unit 902 is used to acquire a second position of the first target object in the target acquisition image when the target information indicates that the first target object appears in the target acquisition image;

[0142] Predict the first position based on the second position.

[0143] The detection unit 903 is used for

[0144] Based on the second position and the first position, motion change feature information of the first target object from the target acquisition image to the i-th frame acquisition image is determined, and the motion change feature information is used as the motion feature information of the first target object in the i-th frame acquisition image.

[0145] Wherein, the first determining unit 904 is used for

[0146] Based on the action change feature information, determine the change action generated by the first target object from the target acquisition image to the i-th frame acquisition image, and determine whether the change action is the predetermined gesture.

[0147] The second acquisition unit 902 is used for:

[0148] If the target information indicates that the first target object does not appear in the target acquisition image, it is determined whether the first target object appears in the i-th frame acquisition image;

[0149] If the result is confirmed to be yes, then the first position is detected.

[0150] Wherein, the second acquisition unit 902 is used for

[0151] If a second target object appears in the i-th frame of the captured image, the position of the second target object in the i-th frame of the captured image is detected; the second target object includes an arm.

[0152] The first position is detected based on the position of the second target object in the i-th frame of the captured image.

[0153] The device further includes a recognition unit for recognizing the gesture type of the predetermined gesture.

[0154] The device further includes a tracking unit and a third acquisition unit; wherein,

[0155] the tracking unit is configured to track the first target object from the target capture image to the ith capture image;

[0156] the third obtaining unit is configured to, in a case where the tracking unit fails to track the first target object in an image region of the ith capture image corresponding to the first position, obtain a number of capture images in which the first target object is continuously failed to be tracked from the target capture image to the ith capture image.

[0157] based on the number, determine whether to update a capture period used by the image capture device when capturing images, the capture period being used to determine a capture interval between the target capture image and the ith capture image.

[0158] The third obtaining unit is configured to:

[0159] in a case where the number is greater than or equal to a pre-set threshold, increase the capture period; wherein, under the increased capture period, a capture interval at which the image capture device captures adjacent two frames of images is increased.

[0160] The second obtaining unit 902 is configured to

[0161] in a case where the target information indicates that the target capture image appears the first target object, input an image region of the first target object in the ith capture image predicted for the ith capture image into a first detection model to obtain the first position;

[0162] in a case where the target information indicates that the target capture image does not appear the first target object, input the ith capture image into a first detection model to obtain the first position;

[0163] The first detection model is obtained by taking a plurality of capture images as training samples and training a position of the first target object.

[0164] The detection unit 903 and the first determination unit 904 are implemented by a second detection model.

[0165] The second detection model is obtained by taking a plurality of capture images as training samples and training an action feature generated by the first target object, and training whether an action having the action feature generated by the first target object is the predetermined gesture.

[0166] It should be noted that the gesture detection apparatus of the present disclosure, due to its problem-solving principle being similar to the aforementioned gesture detection method, the implementation process and implementation principle of the gesture detection apparatus can be referred to the implementation process and implementation principle of the aforementioned method, and the repeated parts will not be described again.

[0167] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0168] The readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the gesture detection method of the present disclosure. The readable storage medium includes but is not limited to random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), optical fiber, compact disc read only memory (CD-ROM). The computer program product includes a computer program, and the computer program is executed by the processor to implement the gesture detection method of the present disclosure.

[0169] The electronic device includes at least one processor and a memory connected with the at least one processor in communication. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the aforementioned gesture method. The processor includes but is not limited to a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc.

[0170] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present disclosure described and / or claimed in this document.

[0171] As Figures 2 to 8As shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded into a random access memory (RAM) 1003 from the storage unit 1008. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 104.

[0172] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, and the like; an output unit 1007, such as various types of displays, speakers, and the like; a storage unit 1008, such as a magnetic disk, an optical disk, and the like; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0173] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 1001 performs various methods and processes described above, such as Figures 2 to 8 any of the methods illustrated above. For example, in some embodiments, Figures 2 to 8 any of the methods illustrated above can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of Figures 2 to 8 any of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform ​ any of the methods illustrated above by any other appropriate means, such as by means of firmware.

[0174] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0175] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0176] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0177] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0178] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0179] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.

[0180] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technology disclosed in the present disclosure, which are not limited herein.

[0181] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and scope of the disclosure. Any modifications, equivalent substitutions, improvements, and the like, made within the spirit and principles of the disclosure, are intended to be included in the scope of the disclosure.

Claims

1. A gesture detection method, comprising: obtaining target information representing whether a first target object appears in a target capture image, wherein the first target object represents a hand, and the target capture image is at least one image captured earlier than an i-th frame capture image, and i is an integer greater than 1; obtaining a first position of the first target object in the i-th frame capture image according to the target information; detecting motion feature information of the first target object in the i-th frame capture image based on the first position; determining whether motion information generated by the first target object in the i-th frame capture image is a predetermined gesture according to the motion feature information; tracking the first target object from the target capture image to the i-th frame capture image; obtaining a number of capture images in which the first target object is not tracked continuously from the target capture image to the i-th frame capture image, in a case that the first target object is not tracked in an image region of the i-th frame capture image corresponding to the first position; determining whether to update a capture period used by an image capture device to capture images, the capture period being used to determine a capture interval between the target capture image and the i-th frame capture image, based on the number; wherein the obtaining the first position of the first target object in the i-th frame capture image according to the target information comprises: inputting an image region of the first target object in the i-th frame capture image predicted for the i-th frame capture image into a first detection model to obtain the first position, in a case that the target information represents that the first target object appears in the target capture image; and inputting the i-th frame capture image into the first detection model to obtain the first position, in a case that the target information represents that the first target object does not appear in the target capture image; wherein the first detection model is obtained by training a position of the first target object in a plurality of capture images using the plurality of capture images as training samples. The obtaining the first position of the first target object in the i-th frame capture image according to the target information comprises: obtaining a second position of the first target object in the target capture image, in a case that the target information represents that the first target object appears in the target capture image; and predicting the first position according to the second position. The detecting the motion feature information of the first target object in the i-th frame capture image based on the first position comprises: determining motion change feature information of the first target object from the target capture image to the i-th frame capture image based on the second position and the first position, and taking the motion change feature information as the motion feature information of the first target object in the i-th frame capture image. The determining whether the motion information generated by the first target object in the i-th frame capture image is the predetermined gesture according to the motion feature information comprises: ​ ​ ​ ​ ​ ​ ​ ​ 2. The method of claim 1, wherein, ​ ​ ​ 3. The method of claim 2, wherein, ​ ​ 4. The method of claim 3, wherein, ​ According to the action change feature information, determine a change action generated by the first target object from the target capture image to the i-th capture image, and determine whether the change action is the predetermined gesture.

5. The method of claim 1, wherein, The acquiring, according to the target information, of the first position of the first target object in the i-th capture image comprises: In a case where the target information indicates that the first target object does not appear in the target capture image, determine whether the first target object appears in the i-th capture image; In a case where the determination is positive, detect the first position.

6. The method of claim 5, wherein, The detecting of the first position comprises: In a case where a second target object appears in the i-th capture image, detect a position of the second target object in the i-th capture image; the second target object comprises an arm; Based on the position of the second target object in the i-th capture image, detect the first position.

7. The method according to any one of claims 1-6, after determining that the action information generated by the first target object is the predetermined gesture, the method further comprises: Identify a gesture type of the predetermined gesture.

8. The method of claim 1, wherein, The determining, based on the quantity, of whether to update a capture period used by an image capture device when capturing images comprises: In a case where the quantity is greater than or equal to a pre-set threshold, increase the capture period; wherein a capture interval between the target capture image and the i-th capture image captured by the image capture device under the increased capture period is increased.

9. The method of claim 1, wherein, The detecting of the action feature information of the first target object in the i-th capture image, and the determining, according to the action feature information of the first target object in the i-th capture image, of whether the action information generated by the first target object is a predetermined gesture are implemented by a second detection model; The second detection model is obtained by training, with a plurality of capture images as training samples, action features generated by the first target object, and training whether an action with the action features generated by the first target object is the predetermined gesture.

10. A gesture detection device, comprising: a first acquiring unit configured to acquire target information, the target information indicating whether a first target object appears in a target capture image, wherein the first target object represents a hand, and the target capture image is at least one image captured earlier than an i-th capture image, i being an integer greater than 1; a second acquiring unit configured to acquire, according to the target information, a first position of the first target object in the i-th capture image; a detecting unit configured to detect, based on the first position, action feature information of the first target object in the i-th capture image; a first determining unit configured to determine, according to the action feature information, whether action information generated by the first target object in the i-th capture image is a predetermined gesture; a tracking unit configured to track the first target object from the target capture image to the i-th capture image; and a second determining unit configured to determine, based on the action feature information of the first target object in the i-th capture image, whether the action information generated by the first target object is the predetermined gesture. The third acquisition unit is configured to, in a case where the tracking unit fails to track the first target object in an image region corresponding to the first position in the i-th frame of captured images, acquire a number of captured images in which the first target object is continuously failed to be tracked from the target captured image to the i-th frame of captured images. Based on the number, determine whether to update a capture period used by an image capture device when capturing images, the capture period being used to determine a capture interval between the target captured image and the i-th frame of captured images. The second acquisition unit is configured to: in a case where the target information indicates that the target captured image contains the first target object, input an image region of the first target object in the i-th frame of captured images predicted for the i-th frame of captured images into a first detection model to obtain the first position; in a case where the target information indicates that the target captured image does not contain the first target object, input the i-th frame of captured images into the first detection model to obtain the first position. The first detection model is obtained by using a plurality of frames of captured images as training samples and training a position of the first target object.

11. The apparatus of claim 10, wherein the second acquisition unit is configured to, in a case where the target information indicates that the target captured image contains the first target object, acquire a second position of the first target object in the target captured image; predict the first position according to the second position.

12. The apparatus of claim 11, wherein, The detection unit is configured to based on the second position and the first position, determine motion change feature information of the first target object from the target captured image to the i-th frame of captured images, and use the motion change feature information as the motion feature information of the first target object in the i-th frame of captured images.

13. The apparatus of claim 12, wherein, The first determination unit is configured to determine a change motion generated by the first target object from the target captured image to the i-th frame of captured images according to the motion change feature information, and determine whether the change motion is the predetermined gesture.

14. The apparatus of claim 10, wherein, The second acquisition unit is configured to: in a case where the target information indicates that the target captured image does not contain the first target object, determine whether the first target object appears in the i-th frame of captured images; in a case where the determination result is yes, detect the first position.

15. The apparatus of claim 11, wherein, The second acquisition unit is configured to in a case where a second target object appears in the i-th frame of captured images, detect a position of the second target object in the i-th frame of captured images; the second target object includes an arm; based on the position of the second target object in the i-th frame of captured images, detect the first position.

16. The apparatus of any one of claims 10-15, wherein, The apparatus further comprises a recognition unit configured to recognize a gesture type of the predetermined gesture.

17. The apparatus of claim 10, wherein, The third acquisition unit is configured to: in a case where the number is greater than or equal to a pre-set threshold, increase the capture period; wherein, under the increased capture period, a capture interval at which the image capture device captures the target captured image and the i-th frame of captured images is increased.

18. The apparatus of claim 10, wherein, The implementation functions of the detection unit and the first determination unit are implemented by a second detection model; The second detection model is obtained by taking a plurality of collected images as training samples, training action features generated by the first target object, and training whether an action with the action features generated by the first target object is the predetermined gesture. 19.An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-9. 21.A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Processing method and equipment based on gesture recognition control instruction, and readable storage medium

    CN108921101A

  • Information processing method, device and system

    CN111860082A

  • Target tracking method and device, electronic equipment and storage medium

    CN112166435A

  • Posture estimation method and device, computer equipment and storage medium

    CN113449696A