Gesture recognition method and device, electronic equipment, computer readable storage medium and computer program product
By performing gesture detection and detection frame merging and screening of the recognized images, the misidentification problem caused by multiple hand objects in the prior art is solved, and a higher accuracy of gesture recognition is achieved.
Patent Information
- Application Number
- CN202411890162.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
When existing gesture recognition technology exists multiple hand objects in the image, it is easy to cause misidentification, and it is impossible to accurately identify the hand objects that require gesture recognition.
By performing gesture detection on the image to be recognized, multiple hand objects and their corresponding detection boxes are obtained, and detection boxes that meet the merge conditions are merged, the object detection boxes are filtered out, and the hand objects in the object detection box are gesture recognition.
When there are multiple hand objects in the image to be identified, the hand objects that need gesture recognition can be accurately determined, so as to improve the accuracy of gesture recognition and reduce misidentification.
Smart Images

Figure CN119942062A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a gesture recognition method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Gesture recognition technology refers to the technology of using computers to detect and classify hand objects in images. It has a wide range of applications in human-computer interaction, game design, education and teaching, and barrier-free technology. Gesture recognition technology mainly relies on deep learning, which is used to detect and recognize gestures. However, in related technologies, when performing gesture recognition, each hand object in the image is often detected independently, so when there are multiple hand objects in the image, all gestures in the image are recognized, resulting in misrecognition. Summary of the invention
[0003] The embodiments of the present application provide a gesture recognition method, device, electronic device, computer-readable storage medium and computer program product, which can accurately determine the hand object that needs to be gesture recognized when there are multiple hand objects in the image to be recognized, thereby improving the accuracy of gesture recognition.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a gesture recognition method, the method comprising:
[0006] Perform gesture detection on the image to be recognized to obtain multiple hand objects and a first detection frame of each of the hand objects; in response to two of the first detection frames satisfying a merging condition, merge the two first detection frames to obtain a merged detection frame; perform detection frame screening based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is a detection frame that has not been merged in the first detection frame; perform gesture recognition on the hand object in the target detection frame to obtain a target gesture corresponding to the image to be recognized.
[0007] The present application provides a gesture recognition device, including:
[0008] A detection module, used for performing gesture detection on the image to be recognized, and obtaining a plurality of hand objects and a first detection frame of each of the hand objects;
[0009] a merging module, configured to merge the two first detection frames to obtain a merged detection frame in response to the two first detection frames satisfying a merging condition;
[0010] A screening module, configured to screen the detection frames based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is a detection frame that has not been merged in the first detection frame;
[0011] The recognition module is used to perform gesture recognition on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized.
[0012] In the above scheme, the detection module is also used to perform object detection on the image to be identified to obtain multiple objects in the image to be identified; based on the contour of each object, determine the first area corresponding to each object; for each first area, extract the key point features of the object in the first area; classify the first area based on the key point features to obtain the second area corresponding to the hand object in the first area; based on the key point position of the hand object in the second area, generate a first detection frame of the hand object.
[0013] In the above scheme, the merging module is also used to determine, for any two first detection frames, a first distance between the center points of the two first detection frames, an average side length of the target borders in the two first detection frames, and the area of each of the first detection frames; when it is determined based on the first distance and the average side length that the two first detection frames meet the distance condition, and when the areas of the two first detection frames meet the area condition, it is determined that the two first detection frames meet the merging condition.
[0014] In the above scheme, the merging module is also used to determine that the two first detection frames meet the distance condition when the first distance is less than the product of the average side length and the distance coefficient; determine the first area difference of the areas of each of the first detection frames, and the minimum area among the areas of each of the first detection frames; when the first area difference is less than the product of the minimum area and the area coefficient, determine that the areas of the two first detection frames meet the area condition.
[0015] In the above scheme, the merging module is also used to determine the minimum horizontal coordinate and the minimum vertical coordinate in the two first detection frames, and determine the first position based on the minimum horizontal coordinate and the minimum vertical coordinate; determine the maximum horizontal coordinate and the maximum vertical coordinate in the two first detection frames, and determine the second position based on the maximum horizontal coordinate and the maximum vertical coordinate; delete the two first detection frames, and draw a merged detection frame with the first position and the second position as diagonals.
[0016] In the above scheme, the screening module is also used to determine the third detection frame with the largest area among the merged detection frame and the second detection frame, and determine the first area of the third detection frame; for each fourth detection frame, determine the second area difference between the first area and the second area of the fourth detection frame, and determine the third area of the fourth detection frame based on the product of the second area and a preset ratio; wherein the fourth detection frame is the detection frame other than the third detection frame among the merged detection frame and the second detection frame; if the third area of each of the fourth detection frames is smaller than the second area difference, the third detection frame is determined as the target detection frame.
[0017] In the above scheme, the screening module is also used to, if there is a fourth detection frame whose third area is greater than the difference between the second area, use the fourth detection frame whose third area is greater than the difference between the second area as the fifth detection frame; and in the third detection frame and the fifth detection frame, screen out the target detection frame whose center point has the shortest distance from the specified point in the image to be identified.
[0018] In the above scheme, the recognition module is also used to obtain the key point image of the hand object in the target detection frame, and to capture the hand image corresponding to the target detection frame in the image to be recognized; wherein the display styles of the key points corresponding to different fingers of the hand object in the key point image are different; the key point image and the hand image are fused to obtain a fused image; gesture recognition is performed based on the fused image to obtain the target gesture corresponding to the image to be recognized.
[0019] An embodiment of the present application provides an electronic device, the electronic device comprising:
[0020] A memory for storing computer executable instructions or computer programs;
[0021] The processor is used to implement the gesture recognition method provided in the embodiment of the present application when executing the computer executable instructions or computer programs stored in the memory.
[0022] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the gesture recognition method provided in the embodiment of the present application when executed by a processor.
[0023] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the gesture recognition method provided in the embodiment of the present application is implemented.
[0024] The embodiments of the present application have the following beneficial effects:
[0025] Through the above method, when performing gesture recognition on the image to be recognized, gesture detection is first performed on the image to be recognized to obtain multiple hand objects in the image to be recognized and the first detection frame of each hand object. For any two first detection frames, the two first detection frames are merged when it is judged by the merging condition that the two first detection frames meet the merging condition to obtain a merged detection frame. Furthermore, the merged detection frame and the second detection frame that has not been merged are screened to obtain a target detection frame of the hand object that is most likely to be captured by the image to be recognized, and gesture recognition is performed on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized. In this way, when there are multiple hand objects in the image to be recognized, the hand object that makes gesture interaction can be accurately determined, and the influence of interfering gestures can be eliminated, thereby improving the accuracy of gesture recognition for the image to be recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a structural diagram of the gesture recognition system architecture provided by an embodiment of the present application;
[0027] Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0028] Figure 3A is a schematic diagram of a first flow chart of a gesture recognition method provided in an embodiment of the present application;
[0029] Figure 3B is a flowchart of a gesture detection method provided in an embodiment of the present application;
[0030] Figure 3C is a flowchart of a detection frame merging method provided in an embodiment of the present application;
[0031] Figure 3D is a schematic diagram of a first process of the detection frame screening method provided in an embodiment of the present application;
[0032] Figure 3E is a second flow chart of the detection frame screening method provided in an embodiment of the present application;
[0033] Figure 3F is a flowchart of a gesture classification method provided in an embodiment of the present application;
[0034] Figure 4A is a schematic diagram of a hand object in an image to be identified provided by an embodiment of the present application;
[0035] Figure 4B is a first schematic diagram of a first detection frame provided in an embodiment of the present application;
[0036] Figure 5 is a schematic diagram of key points of a hand object provided in an embodiment of the present application;
[0037] Fig. 6A is a second schematic diagram of the first detection frame provided in an embodiment of the present application;
[0038] Figure 6B is a schematic diagram of a target detection frame provided in an embodiment of the present application;
[0039] Figure 7 is a schematic diagram of a hand image and a key point image provided in an embodiment of the present application;
[0040] Figure 8 2 is a schematic diagram of a second flow chart of a gesture recognition method provided in an embodiment of the present application.
[0041] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of superiority or inferiority of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0043] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0044] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0045] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0046] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0047] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0048] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0049] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed may be in real time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0050] Due to the rapid development of deep learning technology, deep learning models have brought a lot of revolutionary impacts to computer vision technology. End-to-end detection using deep learning models has quickly become the mainstream method in many fields such as image classification, object detection, instance segmentation, key point detection, etc. Many deep learning models can simultaneously complete multiple visual tasks, for example, they can complete object detection and key point detection in parallel.
[0051] Among them, gesture recognition technology refers to the technology of using computers to identify and classify the gestures of hand objects in images. It has broad application prospects in the fields of human-computer interaction, game design, education and teaching, and barrier-free technology, and provides a contactless and intuitive way for people to interact with computers.
[0052] Gesture recognition technology mainly relies on deep learning, which can be roughly divided into two methods: end-to-end and two-stage. Among them, the end-to-end recognition method refers to using a deep learning model to directly predict the detection box and corresponding category label corresponding to the gesture from the original image. The advantage of this method is that one model is used to solve the recognition problem, and the deployment cost is low. The disadvantage is that gestures are flexible and there are many similar but semantically different gestures, so the model training is difficult. The two-stage recognition method refers to first using a general gesture detection model to detect and extract the gesture area, and then using the gesture classification model to classify the gesture area. Although the inference speed will be slower, the training difficulty is significantly reduced compared to the end-to-end solution.
[0053] Regardless of which of the above methods is used for gesture recognition, the relevant technology often presupposes during training that there is only one hand object in the image. There is no better post-processing method for situations where there are multiple hand objects in the image or gestures composed of both hands. This may cause all gestures in the image to be recognized, leading to the problem of misrecognition.
[0054] The embodiments of the present application provide a gesture recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can accurately determine the hand object that needs to be gesture recognized when there are multiple hand objects in the image to be recognized, thereby improving the accuracy of gesture recognition.
[0055] It should be noted that, whether an end-to-end recognition method or a two-stage recognition method, gesture recognition can be performed based on the gesture recognition method of the embodiment of the present application.
[0056] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the gesture recognition system 100 provided in an embodiment of the present application. To support a gesture recognition application, the terminal 401 is connected to the server 200 via the network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0057] In some embodiments, the terminal 401 and the server 200 can jointly execute the gesture recognition method as an example for explanation, the terminal 401 is used to respond to the user's gesture recognition instruction, obtain the image to be recognized uploaded by the user, and send a gesture recognition request carrying the image to be recognized to the server 200; the server 200 responds to the gesture recognition request, performs gesture detection on the image to be recognized, and obtains multiple hand objects and a first detection frame of each hand object; in response to the two first detection frames satisfying the merging condition, the two first detection frames are merged to obtain a merged detection frame; based on the merged detection frame and the second detection frame, the detection frame is screened to obtain a target detection frame; wherein the second detection frame is a detection frame that has not been merged in the first detection frame; gesture recognition is performed on the hand object in the target detection frame to obtain a target gesture corresponding to the image to be recognized; the server 200 can send the recognized target gesture to the terminal 401 to display the gesture recognition result to the user.
[0058] In some embodiments, taking the example of terminal 401 independently executing the gesture recognition method, terminal 401 responds to the user's gesture recognition instruction, obtains the image to be recognized uploaded by the user, performs gesture detection on the image to be recognized, and obtains multiple hand objects and the first detection frame of each hand object; in response to the two first detection frames satisfying the merging condition, the two first detection frames are merged to obtain a merged detection frame; detection frame screening is performed based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is the detection frame that has not been merged in the first detection frame; gesture recognition is performed on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized.
[0059] In some embodiments, terminal 401 can be implemented as various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a car terminal, etc., and can also be implemented as a server.
[0060] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0061] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device 400 provided in an embodiment of the present application, Figure 2 The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0062] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0063] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0064] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0065] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0066] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0067] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0068] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;
[0069] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);
[0070] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.
[0071] In some embodiments, the gesture recognition device provided in the embodiments of the present application can be implemented in a software manner. Figure 2 The gesture recognition device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a detection module 4551, a merging module 4552, a screening module 4553, and a recognition module 4554. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0072] In other embodiments, the gesture recognition device provided in the embodiments of the present application may be implemented in hardware. As an example, the gesture recognition device provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the gesture recognition method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application specific integrated circuits (Application Specific Integrated Circuit, ASIC), digital signal processors (Digital Signal Processor, DSP), programmable logic devices (Programmable Logic Device, PLD), complex programmable logic devices (Complex Programmable Logic Device, CPLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other electronic components.
[0073] In some embodiments, the terminal or server can implement the gesture recognition method provided in the embodiments of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a native application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a smart home control APP with gesture recognition function, a game APP, etc.; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0074] The following will describe the gesture recognition method provided by the embodiment of the present application in conjunction with the accompanying drawings. As mentioned above, the electronic device that implements the gesture recognition method of the embodiment of the present application can be the terminal 401, the server 200, or a combination of the two. Therefore, the execution subject of each step will not be repeatedly described below.
[0075] The gesture recognition method of the embodiment of the present application is described by taking the execution subject as the terminal 401 and the terminal 401 adopting the above two-stage recognition method as an example. Here, the terminal 401 may be deployed with a gesture detection model and a gesture classification model.
[0076] See also Figure 3A , Figure 3A is a first flow chart of the gesture recognition method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.
[0077] In step 101, gesture detection is performed on the image to be recognized to obtain a plurality of hand objects and a first detection frame of each hand object.
[0078] The image to be identified includes multiple hand objects and may also include any other objects (such as face objects, etc.). The multiple hand objects in the image to be identified include at least one hand object that needs to be gesture recognized, and there may also be interference objects that do not need to be gesture recognized. Figure 4A is a schematic diagram of a hand object in an image to be identified provided by an embodiment of the present application, see Figure 4A , Figure 4AThe image to be identified shown includes two hand objects (hand object 110 and hand object 120), wherein the gesture made by hand object 110 is the target gesture to be identified for the image to be identified; while hand object 120 does not make a clear gesture and only appears in the image to be identified, so it is an interference object. In the related art, when performing gesture detection on the image to be identified, hand object 120 is often also performed gesture recognition, which will lead to the occurrence of misrecognition problems.
[0079] In the embodiment of the present application, since it is necessary to exclude interfering objects and avoid omission problems, when performing gesture detection on the image to be identified, all hand objects in the image to be identified can be detected and a first detection frame of each hand object can be obtained.
[0080] In some embodiments, Figure 3B is a flowchart of a gesture detection method provided in an embodiment of the present application, see Figure 3B , step 101 can be implemented by steps 1011 to 1015:
[0081] In step 1011, object detection is performed on the image to be identified to obtain multiple objects in the image to be identified.
[0082] In actual implementation, a pre-trained deep learning gesture detection model can be called to perform gesture detection on the object to be identified, or a pre-built image processing method can be directly used to perform gesture detection on the image to be identified, and there is no limitation here.
[0083] In actual implementation, when performing object detection on the image to be identified, the terminal 401 may detect all types of objects in the image to be identified and obtain the position of each object in the image to be identified.
[0084] In step 1012, based on the outline of each object, a first area corresponding to each object is determined.
[0085] In actual implementation, based on the contour of each object and the position of each object in the image to be identified, the first region corresponding to each object in the image to be identified is determined. Here, the first region corresponding to each object only contains one object.
[0086] In step 1013, for each first region, key point features of the object in the first region are extracted.
[0087] In actual implementation, by extracting key points from the objects in the first area, the distribution of the key points can be used to distinguish the hand object from other objects.
[0088] As an example, Figure 5 is a schematic diagram of the key points of the hand object provided in the embodiment of the present application, see Figure 5 When using the gesture detection model to detect hand objects, the parameters of the key point extraction network in the gesture detection model can be adjusted to 21, corresponding to the 21 key points of gesture recognition, which are usually based on joint positions. When the gesture detection model is used to detect gestures on the image to be recognized, 21 key point features of the hand object can be obtained. Similarly, after extracting key points for other objects, the corresponding number of key points of other objects can be obtained.
[0089] In step 1014, the first region is classified based on the key point features to obtain a second region in the first region corresponding to the hand object.
[0090] In actual implementation, the first regions can be classified based on key point features in each first region based on a pre-trained classifier (e.g., a convolutional neural network), that is, the objects in the first region can be classified to obtain the second region corresponding to the hand object in the first region. The other regions in the first region except the second region are deleted and no further processing is performed.
[0091] In step 1015, a first detection frame of the hand object is generated based on the key point positions of the hand object in the second area.
[0092] In actual implementation, for each second area, based on the key point positions extracted for the hand object in the second area and the outline of the hand object, the boundary of the hand object is accurately located, and the position of the boundary is finely adjusted through a regression algorithm, and a first detection frame corresponding to the hand object is generated based on the boundary of the hand object. The shape of the first detection frame can be set based on actual needs, and the first detection frame here can be a quadrilateral. The first detection frame of the hand object can closely surround the hand object to reduce background interference in the image to be identified.
[0093] As an example, Figure 4B is a first schematic diagram of a first detection frame provided in an embodiment of the present application, see Figure 4B ,against Figure 4A After object detection is performed on the image to be identified, a first detection frame 1101 of the hand object 110 and a first detection frame 1201 of the hand object 120 can be obtained.
[0094] Continue to see Figure 3A , and the description will continue with the above step 101.
[0095] In step 102, in response to two first detection frames satisfying a merging condition, the two first detection frames are merged to obtain a merged detection frame.
[0096] In actual implementation, Fig. 6A is a second schematic diagram of the first detection frame provided in an embodiment of the present application, see Fig. 6A Since the target gesture to be recognized in the image to be recognized may be completed by two hand objects, in order to recognize the accurate complete gesture, it is possible to determine whether there is a first detection frame that needs to merge the hand object based on the preset merging conditions, so as to recognize the complete target gesture based on the gestures of the two hand objects.
[0097] In some embodiments, the terminal 401 can determine that the two first detection frames meet the merging condition in the following manner: for any two first detection frames, determine the first distance between the center points of the two first detection frames, the average side length of the target borders in the two first detection frames, and the area of each first detection frame; when it is determined based on the first distance and the average side length that the two first detection frames meet the distance condition, and the areas of the two first detection frames meet the area condition, determine that the two first detection frames meet the merging condition.
[0098] In actual implementation, if the image to be identified includes two or more hand objects, that is, includes two or more first detection frames, it can be determined whether any two first detection frames meet the merging condition for any two first detection frames.
[0099] As an example, the center point position of each first detection frame can be determined according to the actual shape of the first detection frame. Taking the first detection frame as a quadrilateral as an example, the upper left corner coordinates and the lower right corner coordinates of the first detection frame can be obtained, and the center point coordinates of the first detection frame can be calculated based on the upper left corner coordinates and the lower right corner coordinates.
[0100] As an example, the coordinates of the center point of the first detection frame can be calculated by the following formula:
[0101]
[0102] Among them, (x 1 ,y 1 ) represents the coordinates of the upper left corner of the first detection box, (x 2 ,y 2 ) represents the coordinates of the lower right corner of the first detection frame; (x c ,y c ) represents the coordinates of the center point of the first detection frame.
[0103] According to the coordinates of the center points of the two first detection frames, the first distance between the two first detection frames can be determined. For example, the first distance can be calculated by the following formula:
[0104]
[0105] Among them, (x c1 ,y c1 ) and (x c2 ,y c2) represent the center point coordinates of the two first detection frames respectively; d 1 Indicates the first distance.
[0106] In actual implementation, since the positions and visual depths of the two hand objects in the image to be identified are roughly the same when two hands make gestures at the same time, that is, the areas of the first detection frames of the two hand objects are also approximately equal. Therefore, when judging whether any two first detection frames meet the distance condition, the side length and area of the first detection frame can also be obtained.
[0107] As an example, when obtaining the side length of the first detection frame, the side length of the target frame preset for the first detection frame can be obtained, and the average side length of the target frames of the two first detection frames can be calculated based on the side lengths of the target frames in the two first detection frames. Specifically, the average side length of the target frames in the two first detection frames can be calculated using the following formula:
[0108]
[0109] Among them, w 1 and w 2 Respectively represent the side lengths of the target borders in the two first detection frames; d 2 Represents the average side length.
[0110] Then, whether the two first detection frames meet the distance condition is measured according to the average side length and the first distance. If the distance condition is met, it indicates that the distance between the two first detection frames is close enough. For the area condition, if the areas of the two first detection frames meet the area condition, it indicates that the areas of the two first detection frames are approximately equal.
[0111] In actual implementation, when the two first detection frames meet the distance condition and the area condition at the same time, it can be determined that the two first detection frames meet the merging condition, that is, when it is determined based on the first distance and the average side length that the two first detection frames meet the distance condition and the areas of the two first detection frames meet the area condition, it is determined that the two first detection frames meet the merging condition.
[0112] In some embodiments, the terminal 401 may determine whether the two first detection frames satisfy the distance condition and whether the areas of the two first detection frames satisfy the area condition in the following manner: when the first distance is less than the product of the average side length and the distance coefficient, determine that the two first detection frames satisfy the distance condition; determine a first area difference between the areas of each first detection frame and the minimum area among the areas of each first detection frame; when the first area difference is less than the product of the minimum area and the area coefficient, determine that the areas of the two first detection frames satisfy the area condition.
[0113] In actual implementation, when judging the distance condition, a distance coefficient α (the value is usually greater than 1, such as 1.2) can be set for the average side length based on the actual accuracy requirements, and 1 Less than the average side length d 2 When the product of d and the distance coefficient α is obtained, it is determined that the two first detection frames meet the distance condition, that is, when d 1 <α*d 2 , it is determined that the distance condition is met.
[0114] In actual implementation, when judging the area condition, an area coefficient can also be set based on actual needs, and the area difference between the two first detection frames, the minimum area and the area coefficient can be used to judge whether the areas of the two first detection frames meet the area condition. Specifically, the area condition can be judged to be met when the following formula is met:
[0115] |s 1 -s 2 |<0.25*min(s 1 ,s 2 )
[0116] Among them, s 1 and 2 Respectively represent the areas of the two first detection frames; min(s 1 ,s 2 ) means taking the minimum value of the two areas, that is, determining the minimum area.
[0117] Continue to see Fig. 6A , determined based on the above method Fig. 6A After the first detection frame 210 and the first detection frame 220 of the two hand objects meet the distance condition and the areas of the two first detection frames also meet the area condition, the two first detection frames can be merged to regard the two hand objects as a whole for gesture recognition.
[0118] In some embodiments, Figure 3C is a flow chart of the detection frame merging method provided in the embodiment of the present application, see Figure 3C , “merging the two first detection frames to obtain a merged detection frame” in step 102 can be implemented through steps 1021 to 1023.
[0119] In step 1021, the minimum horizontal coordinate and the minimum vertical coordinate in the two first detection frames are determined, and the first position is determined based on the minimum horizontal coordinate and the minimum vertical coordinate.
[0120] In actual implementation, since the merged detection frame obtained after merging the two first detection frames needs to completely cover the two hand objects, the merged detection frame can be drawn based on the highest point and the lowest point of the area occupied by the two first detection frames, so as to cover all areas of the two first detection frames.
[0121] In actual implementation, the shape of the merged detection frame can be a quadrilateral. After determining the minimum horizontal coordinate and the minimum vertical coordinate in the two first detection frames, the first position determined based on the minimum horizontal coordinate and the minimum vertical coordinate can be used as the lower left corner position of the merged detection frame.
[0122] In step 1022, the maximum horizontal coordinate and the maximum vertical coordinate in the two first detection frames are determined, and the second position is determined based on the maximum horizontal coordinate and the maximum vertical coordinate.
[0123] In actual implementation, after determining the maximum horizontal coordinate and the maximum vertical coordinate of the two first detection frames, the second position determined based on the maximum horizontal coordinate and the maximum vertical coordinate can be used as the upper right corner position of the merged detection frame.
[0124] In step 1023, the two first detection frames are deleted, and a merged detection frame is drawn with the first position and the second position as diagonals.
[0125] As an example, Figure 6B is a schematic diagram of the target detection frame provided in the embodiment of the present application, see Fig. 6A and Figure 6B , for Fig. 6A When merging the two first detection frames in , after determining the first position and the second position according to the above method, the two first detection frames can be deleted, and the two hand objects are regarded as a whole. The first position is used as the lower left corner of the target detection frame, and the second position is used as the upper right corner of the target detection frame to draw a merged detection frame for the two hand objects, and the following is obtained: Figure 6B The object detection box shown can completely cover the two hand objects.
[0126] Through the above method, when the target gesture to be identified in the image to be identified is completed by a combination of two or more hand objects, through the above target conditions, a merge judgment is performed on any two first detection frames, which can effectively identify the hand objects involved in the target gesture in independent hand objects, and generate a target detection frame for the corresponding combined hand object, thereby improving the accuracy of target gesture recognition.
[0127] Continue to see Figure 3A , and continue with the above step 102 for explanation.
[0128] In step 103, detection frame screening is performed based on the merged detection frame and the second detection frame to obtain a target detection frame.
[0129] The second detection frame is the detection frame in the first detection frame that has not been merged.
[0130] In actual implementation, considering that there may be interfering gestures in the image to be recognized, after the merging process is completed, the merged detection frame and the second detection frame may be screened to exclude interfering objects.
[0131] In some embodiments, Figure 3D is a schematic diagram of the first process of the detection frame screening method provided in the embodiment of the present application, see Figure 3D , step 103 can be implemented through steps 1031 to 1033.
[0132] In step 1031, a third detection frame having the largest area among the merged detection frame and the second detection frame is determined, and a first area of the third detection frame is determined.
[0133] In actual implementation, when shooting the image to be recognized, the hand object used for gesture interaction usually tends to place the gesture in the center of the camera lens, and is generally closer to the lens than the gesture of the interference object. Therefore, the target detection frame can be screened by the area size of the detection frame.
[0134] As an example, see Figure 4B , Figure 4B The two detection frames in are determined to be unnecessary to be merged after the above merging conditions are determined. Therefore, Figure 4B Both detection frames in belong to the second detection frame, that is, the first detection frame 1101 and the first detection frame 1201 are detection frames that have not been merged. Among them, the hand object in the first detection frame 1101 is making a gesture towards the camera, and the hand object in the first detection frame 1201 is an interference object. Therefore, in Figure 4B The first detection frame 1101 is located closer to the center of the image and has a larger area.
[0135] In actual implementation, if after merging, a merged detection frame and a second detection frame exist in the image to be identified, the detection frame with the largest area may be selected as the third detection frame, and the first area of the third detection frame may be calculated.
[0136] In step 1032, for each fourth detection frame, a second area difference between the first area and the second area of the fourth detection frame is determined, and a third area of the fourth detection frame is determined based on the product of the second area and a preset ratio.
[0137] The fourth detection frame is a detection frame in the merged detection frame and the second detection frame except the third detection frame.
[0138] In actual implementation, in order to improve the accuracy of the determined target detection frame, after determining the third detection frame with the largest area, it is necessary to compare the area of the third detection frame with the fourth detection frame other than the third detection frame in the image to be identified, so that the third detection frame can be used as the target detection frame only when the comparison conditions are met.
[0139] As an example, a preset ratio can be used to determine whether the first area of the third detection frame exceeds the preset ratio of the second area of the fourth detection frame. For example, determine whether the first area of the third detection frame exceeds 30% of the second area of the fourth detection frame. When comparing, the difference between the first area of the third detection frame and the second area of the fourth detection frame (i.e., the second area difference) can be compared with the size of the product of the second area of the fourth detection frame and the preset ratio. Therefore, for each fourth detection frame, the product of the second area of the fourth detection frame and the preset ratio can be determined as the fourth area of the fourth detection frame, and the second area difference between the first area of the third detection frame and the second area of the fourth detection frame can be determined.
[0140] Here, the preset ratio is manually set and can be set according to actual needs without any restrictions.
[0141] In step 1033, if the third area of each fourth detection frame is smaller than the second area difference, the third detection frame is determined as the target detection frame.
[0142] In one case, for each fourth detection frame, the first area of the third detection frame exceeds the preset ratio of the second area of the fourth detection frame, that is, the third area of each fourth detection frame is smaller than the second area difference, then the third detection frame can be directly determined as the target detection frame. Stop screening.
[0143] In some embodiments, Figure 3E is a second flow chart of the detection frame screening method provided in the embodiment of the present application, see Figure 3E After step 1032 is executed, the target detection frame can also be determined through steps 1034 to 1035.
[0144] In step 1034, if there is a fourth detection frame whose third area is greater than the difference between the second area, the fourth detection frame whose third area is greater than the difference between the second area is used as the fifth detection frame.
[0145] In another case, there is a fourth detection frame whose first area cannot exceed a preset proportion of the second area, that is, there is a fourth detection frame whose third area is greater than the difference between the second area, then all fourth detection frames whose third area is greater than the difference between the second area are screened out, and the screened out fourth detection frames are used as the fifth detection frames.
[0146] Here, it can be understood that the area difference between the fifth detection frame and the third detection frame is not much different, and both may be the detection frames of the hand object interacting with the lens. Therefore, it is necessary to continue screening in the fifth detection frame and the third detection frame.
[0147] In step 1035, the target detection frame having the smallest distance from the center point to the designated point in the image to be identified is screened out from the third detection frame and the fifth detection frame.
[0148] In actual implementation, since the hand interacting with the lens is generally closer to the center of the lens, the designated point in the image to be identified can be determined as the center point of the image to be identified, and the target detection frame can be determined by the distance between the center point of each detection frame and the center point of the image to be identified.
[0149] As an example, determine the distance from the center point of the third detection frame to the center point of the image to be identified, as well as the distance from the center point of each fifth detection frame to the center point of the image to be identified, and select the detection frame with the smallest distance as the target detection frame. The target detection frame can be regarded as the detection frame of the hand object that is most biased towards the center of the lens and closer to the camera among all detection frames.
[0150] In the above manner, based on the conditions of area and distance, screening is performed in the merged detection frame and the second detection frame. The hand object in the screened target detection frame is the hand object that is most likely to interact with the camera, that is, the target gesture captured by the camera is most likely to be made by the hand object in the target detection frame. This can improve the recognition accuracy of the target gesture corresponding to the image to be recognized, and enhance the gesture recognition ability of the end-to-end deep learning model or the gesture classification model in the two-stage for the image to be recognized.
[0151] Continue to see Figure 3A , and continue with the above step 103 for explanation.
[0152] In step 104, gesture recognition is performed on the hand object in the target detection frame to obtain a target gesture corresponding to the image to be recognized.
[0153] In actual implementation, after determining the target detection frame in the image to be identified, gesture recognition can be performed on the hand object in the target detection frame, that is, the gestures made by the hand object in the target detection frame are classified to determine what gesture the hand object has made.
[0154] In some embodiments, Figure 3F is a flowchart of the gesture classification method provided in the embodiment of the present application, see Figure 3F , step 104 can be implemented through steps 1041 to 1043.
[0155] In step 1041, a key point image of a hand object in a target detection frame is obtained, and a hand image corresponding to the target detection frame is captured in the image to be identified.
[0156] Among them, the key points corresponding to different fingers of the hand object in the key point image have different display styles.
[0157] In actual implementation, in order to improve the accuracy of gesture recognition, after determining the target detection frame, a hand image that only includes the hand object in the target detection frame can be captured from the image to be recognized, and a key point image extracted from the hand object in the target detection frame when the gesture detection model performs object detection on the image to be recognized can be obtained. The hand image and the key point image are combined to perform gesture recognition on the hand object in the target detection frame.
[0158] As an example, see Figure 7 , Figure 7 is a schematic diagram of a hand image and a key point image provided in an embodiment of the present application, Figure 7 (a) in the figure represents the hand image of the hand object in the corresponding target detection frame captured in the image to be identified. Figure 7 (b) in the figure represents the key point image of the hand object in the target detection frame. The key point image includes 21 key points of the hand object. These key points usually reflect the joint positions of each finger. In order to facilitate the model to distinguish different gestures, the keywords corresponding to different fingers in the key point image can be displayed separately, for example, different colors are used to reflect the key points of different fingers. See Figure 7 In the (b) diagram, the key points of the thumb can be displayed in green, the key points of the index finger can be displayed in red, the key points of the middle finger can be displayed in blue, the key points of the ring finger can be displayed in yellow, the key points of the little finger can be displayed in pink, and the key points of the palm can be displayed in white. In this way, by giving different display styles to the key points of different fingers in the key point image, the importance of the finger position in the hand object can be strengthened when training a deep learning model or a gesture classification model, thereby enhancing the gesture classification ability.
[0159] In step 1042, the key point image and the hand image are fused to obtain a fused image.
[0160] In actual implementation, the key point image and the hand image can be converted into tensors, that is, feature extraction is performed on the two images respectively to obtain the first image feature of the key point image and the second image feature of the hand image, and then the first image feature and the second image feature are feature fused to obtain a fused feature, which can reflect the key information in the key point image and the hand image, and the fused feature reflects the image feature of the fused image. Here, the fusion method can be stacked, and the fusion method can be selected according to the actual situation without specific limitation.
[0161] In step 1043, gesture recognition is performed based on the fused image to obtain a target gesture corresponding to the image to be recognized.
[0162] In actual implementation, if a pre-trained gesture classification model is used for gesture recognition, the fused image can be input into the gesture classification model, and the hand object in the fused image can be classified and recognized through the gesture classification model to obtain the target gesture corresponding to the image to be recognized. The target gesture is the gesture that is most likely to be captured by the image to be recognized.
[0163] Through the above method, when performing gesture recognition on the image to be recognized, gesture detection is first performed on the image to be recognized to obtain multiple hand objects in the image to be recognized and the first detection frame of each hand object. For any two first detection frames, the two first detection frames are merged when it is judged by the merging condition that the two first detection frames meet the merging condition to obtain a merged detection frame. Furthermore, the merged detection frame and the second detection frame that has not been merged are screened to obtain a target detection frame of the hand object that is most likely to be captured by the image to be recognized, and gesture recognition is performed on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized. In this way, when there are multiple hand objects in the image to be recognized, the hand object that makes gesture interaction can be accurately determined, and the influence of interfering gestures can be eliminated, thereby improving the accuracy of gesture recognition for the image to be recognized.
[0164] In a specific embodiment, Figure 8 is a second flow chart of the gesture recognition method provided in the embodiment of the present application, see Figure 8 , the gesture recognition method may also include the following process:
[0165] S201, obtaining an image to be recognized.
[0166] S202, gesture detection.
[0167] Through the pre-trained gesture detection model, gesture detection is performed on the image to be recognized to obtain multiple hand objects and the first detection frame of each hand object. Here, the gesture detection model also extracts the key point features of each hand object, that is, obtains the key point image of each hand object.
[0168] S203, gesture merging.
[0169] For any two first detection frames, when it is determined that the two first detection frames meet the merging condition, the two first detection frames are merged to obtain a merged detection frame. The merging process can refer to the description in the above related embodiments, which will not be repeated here.
[0170] S204, screening target gestures.
[0171] The detection frame is screened based on the merged detection frame and the second detection frame to obtain the target detection frame, that is, the detection frame corresponding to the target gesture; wherein the second detection frame is the detection frame that is not merged in the first detection frame. Here, the screening process can refer to the description in the above-mentioned related embodiments, and will not be repeated here.
[0172] S205, obtaining a hand image.
[0173] S206, obtaining a key point image of the hand.
[0174] S207, gesture recognition.
[0175] In order to improve the accuracy of gesture recognition, after determining the target detection frame, a hand image that only includes the hand object in the target detection frame can be captured from the image to be recognized, and a key point image extracted from the hand object in the target detection frame when the gesture detection model performs object detection on the image to be recognized can be obtained. The hand image and the key point image are combined to perform gesture recognition on the hand object in the target detection frame.
[0176] S208, obtaining recognition results.
[0177] The following is a description of an exemplary structure of a gesture recognition device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the gesture recognition device 455 of the memory 450 may include:
[0178] A detection module 4551 is used to perform gesture detection on the image to be recognized, and obtain multiple hand objects and a first detection frame of each hand object;
[0179] A merging module 4552, configured to merge the two first detection frames to obtain a merged detection frame in response to the two first detection frames satisfying a merging condition;
[0180] A screening module 4553 is used to screen the detection frames based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is a detection frame in the first detection frame that has not been merged;
[0181] The recognition module 4554 is used to perform gesture recognition on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized.
[0182] In some embodiments, the detection module 4551 is also used to perform object detection on the image to be identified to obtain multiple objects in the image to be identified; based on the outline of each object, determine the first area corresponding to each object; for each first area, extract the key point features of the object in the first area; classify the first area based on the key point features to obtain the second area corresponding to the hand object in the first area; based on the key point position of the hand object in the second area, generate a first detection frame of the hand object.
[0183] In some embodiments, the merging module 4552 is also used to determine, for any two first detection frames, a first distance between the center points of the two first detection frames, an average side length of the target borders in the two first detection frames, and an area of each first detection frame; when it is determined based on the first distance and the average side length that the two first detection frames meet the distance condition, and when the areas of the two first detection frames meet the area condition, it is determined that the two first detection frames meet the merging condition.
[0184] In some embodiments, the merging module 4552 is also used to determine that the two first detection frames meet the distance condition when the first distance is less than the product of the average side length and the distance coefficient; determine the first area difference between the areas of each first detection frame, and the minimum area among the areas of each first detection frame; and determine that the areas of the two first detection frames meet the area condition when the first area difference is less than the product of the minimum area and the area coefficient.
[0185] In some embodiments, the merging module 4552 is also used to determine the minimum horizontal coordinate and the minimum vertical coordinate in the two first detection frames, and determine the first position based on the minimum horizontal coordinate and the minimum vertical coordinate; determine the maximum horizontal coordinate and the maximum vertical coordinate in the two first detection frames, and determine the second position based on the maximum horizontal coordinate and the maximum vertical coordinate; delete the two first detection frames, and draw a merged detection frame with the first position and the second position as diagonals.
[0186] In some embodiments, the screening module 4553 is also used to determine the third detection frame with the largest area between the merged detection frame and the second detection frame, and determine the first area of the third detection frame; for each fourth detection frame, determine the second area difference between the first area and the second area of the fourth detection frame, and determine the third area of the fourth detection frame based on the product of the second area and a preset ratio; wherein the fourth detection frame is the detection frame other than the third detection frame in the merged detection frame and the second detection frame; if the third area of each fourth detection frame is smaller than the second area difference, the third detection frame is determined as the target detection frame.
[0187] In some embodiments, the screening module 4553 is also used to, if there is a fourth detection frame whose third area is greater than the difference between the second area, use the fourth detection frame whose third area is greater than the difference between the second area as the fifth detection frame; among the third detection frame and the fifth detection frame, screen out the target detection frame whose center point has the shortest distance from the specified point in the image to be identified.
[0188] In some embodiments, the recognition module 4554 is also used to obtain the key point image of the hand object in the target detection frame, and to capture the hand image corresponding to the target detection frame in the image to be recognized; wherein the display styles of the key points corresponding to different fingers of the hand object in the key point image are different; the key point image and the hand image are fused to obtain a fused image; gesture recognition is performed based on the fused image to obtain the target gesture corresponding to the image to be recognized.
[0189] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device performs the gesture recognition method described in the embodiment of the present application.
[0190] The present application embodiment provides a computer-readable storage medium in which computer executable instructions or computer programs are stored. When the computer executable instructions or computer programs are executed by a processor, the processor will be caused to execute the gesture recognition method provided by the present application embodiment, for example, Figure 3A The gesture recognition method is shown.
[0191] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0192] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0193] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0194] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0195] To summarize, through the embodiments of the present application, when performing gesture recognition on an image to be recognized, gesture detection is first performed on the image to be recognized to obtain multiple hand objects in the image to be recognized and a first detection frame of each hand object. For any two first detection frames, the two first detection frames are merged when it is judged through a merging condition that the two first detection frames meet the merging condition to obtain a merged detection frame. Furthermore, screening is performed in the merged detection frame and the second detection frame that has not been merged to obtain a target detection frame of the hand object that is most likely to be captured by the image to be recognized, and gesture recognition is performed on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized. In this way, when there are multiple hand objects in the image to be recognized, the hand object that performs gesture interaction can be accurately determined, and the influence of interfering gestures can be eliminated, thereby improving the accuracy of gesture recognition for the image to be recognized.
[0196] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A gesture recognition method, characterized in that: The method comprises: Performing gesture detection on the image to be recognized to obtain a plurality of hand objects and a first detection frame of each of the hand objects; In response to the two first detection frames satisfying a merging condition, merging the two first detection frames to obtain a merged detection frame; Performing detection frame screening based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is a detection frame in the first detection frame that has not been merged; Perform gesture recognition on the hand object in the target detection frame to obtain a target gesture corresponding to the image to be recognized.
2. The method according to claim 1, characterized in that: The performing gesture detection on the image to be recognized to obtain a plurality of hand objects and a first detection frame of each of the hand objects includes: Performing object detection on the image to be identified to obtain multiple objects in the image to be identified; Based on the outline of each of the objects, determining a first area corresponding to each of the objects; For each of the first regions, extract key point features of the object in the first region; Classifying the first region based on the key point features to obtain a second region in the first region corresponding to the hand object; Based on the key point positions of the hand object in the second area, a first detection frame of the hand object is generated.
3. The method according to claim 1, characterized in that Before merging the two first detection frames, the method includes: For any two first detection frames, determine a first distance between center points of the two first detection frames, an average side length of target frames in the two first detection frames, and an area of each of the first detection frames; When it is determined based on the first distance and the average side length that the two first detection frames meet a distance condition and the areas of the two first detection frames meet an area condition, it is determined that the two first detection frames meet a merging condition.
4. The method according to claim 3, characterized in that Before determining that the two first detection frames meet a merging condition, the method includes: When the first distance is less than the product of the average side length and the distance coefficient, determining that the two first detection frames meet the distance condition; Determine a first area difference between areas of the first detection frames and a minimum area among areas of the first detection frames; When the first area difference is smaller than the product of the minimum area and the area coefficient, it is determined that the areas of the two first detection frames meet the area condition.
5. The method according to claim 1, characterized in that The step of merging the two first detection frames to obtain a merged detection frame includes: Determine a minimum horizontal coordinate and a minimum vertical coordinate in the two first detection frames, and determine a first position based on the minimum horizontal coordinate and the minimum vertical coordinate; Determine a maximum horizontal coordinate and a maximum vertical coordinate in the two first detection frames, and determine a second position based on the maximum horizontal coordinate and the maximum vertical coordinate; The two first detection frames are deleted, and a merged detection frame is drawn with the first position and the second position as diagonals.
6. The method according to claim 1, characterized in that The performing detection frame screening based on the combined detection frame and the second detection frame to obtain the target detection frame includes: determining a third detection frame having the largest area among the combined detection frame and the second detection frame, and determining a first area of the third detection frame; For each fourth detection frame, determine a second area difference between the first area and the second area of the fourth detection frame, and determine a third area of the fourth detection frame based on a product of the second area and a preset ratio; wherein the fourth detection frame is a detection frame in the merged detection frame and the second detection frame except the third detection frame; If the third area of each of the fourth detection frames is smaller than the second area difference, the third detection frame is determined as the target detection frame.
7. The method according to claim 6, characterized in that The method further comprises: If there is a fourth detection frame whose third area is greater than the difference of the second area, use the fourth detection frame whose third area is greater than the difference of the second area as the fifth detection frame; In the third detection frame and the fifth detection frame, a target detection frame having a minimum distance from a center point to a designated point in the image to be identified is obtained by screening.
8. The method according to claim 1, characterized in that The performing gesture recognition on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized includes: Acquire a key point image of a hand object in the target detection frame, and intercept a hand image corresponding to the target detection frame in the image to be identified; wherein the key points corresponding to different fingers of the hand object in the key point image have different display styles; Fusing the key point image and the hand image to obtain a fused image; Perform gesture recognition based on the fused image to obtain a target gesture corresponding to the image to be recognized.
9. A gesture recognition device, characterized in that: The device comprises: A detection module, used for performing gesture detection on the image to be recognized, and obtaining a plurality of hand objects and a first detection frame of each of the hand objects; a merging module, configured to merge the two first detection frames to obtain a merged detection frame in response to the two first detection frames satisfying a merging condition; A screening module, configured to screen the detection frames based on the merged detection frame and the second detection frame to obtain a target detection frame; wherein the second detection frame is a detection frame that has not been merged in the first detection frame; The recognition module is used to perform gesture recognition on the hand object in the target detection frame to obtain the target gesture corresponding to the image to be recognized.
10. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; The processor is used to implement the gesture recognition method according to any one of claims 1 to 8 when executing the computer executable instructions or computer programs stored in the memory.
11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the gesture recognition method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising computer executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the gesture recognition method according to any one of claims 1 to 8 is implemented.