Hand posture recognition method, hand posture recognition model training method and device
By combining two-dimensional and three-dimensional pose recognition models and constraint loss functions, and adjusting the differences in three-dimensional hand joints between video frames, the problems of inconsistency and flickering in real-time hand pose recognition are solved, and the stability and consistency of the recognition results are improved.
Patent Information
- Application Number
- CN202210887526.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Existing technologies for real-time hand gesture recognition in video suffer from problems such as discontinuous simulated hand gestures and flickering/jumping between multiple video frames, which affect the stability of the recognition results.
By combining two-dimensional and three-dimensional pose recognition models, and adjusting the differences in three-dimensional hand joints between multiple video frames through constraint loss function and chiral feature matching, the continuity and stability of hand pose are ensured.
It improves the stability of hand gesture recognition results, ensures smooth and continuous hand gestures between video frames, and reduces flickering and jitter.
Smart Images

Figure CN115223248B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a hand gesture recognition method, a training method for a hand gesture recognition model, an apparatus, a computer device, and a storage medium. Background Technology
[0002] Hands, with their flexible and adaptable nature, play a vital role in production and daily life. In various application scenarios such as human-computer interaction, virtual reality, and sign language recognition, hand postures can be simulated by recognizing hand gestures in videos or images.
[0003] Currently, related technologies typically employ large-scale feature extraction networks to extract features from hand images, thereby determining the positions of joints in the hand images. Based on the positions of these joints, the positions of the joints in the hand skeleton model are then adjusted to simulate hand posture.
[0004] However, when the above technical solutions are applied to real-time hand gesture recognition of videos, problems such as discontinuity, flickering and jerking of the simulated hand gestures between multiple video frames occur, which greatly affects the stability of the hand gesture recognition results. Summary of the Invention
[0005] This application provides a hand pose recognition method, a training method for a hand pose recognition model, an apparatus, a computer device, and a storage medium, which can effectively improve the stability of hand pose recognition results. The technical solution is as follows:
[0006] On the one hand, a hand pose recognition method is provided, which includes:
[0007] Based on a two-dimensional pose recognition model, the first video frame is processed to determine multiple two-dimensional hand joints in the first video frame. These multiple two-dimensional hand joints are used to describe the two-dimensional pose of the hand in the first video frame.
[0008] Based on multiple two-dimensional hand joints and a three-dimensional pose recognition model of the first video frame, multiple three-dimensional hand joints of the first video frame are determined, and the three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0009] Based on the constraint loss function and multiple 3D hand joints in the second video frame, the multiple 3D hand joints in the first video frame are processed so that the difference between the multiple 3D hand joints in the second video frame and the multiple 3D hand joints in the first video frame satisfies the target condition. The second video frame is a video frame preceding the first video frame.
[0010] On the one hand, a training method for a hand pose recognition model is provided, the method comprising:
[0011] Acquire a sample video of the hand, the two-dimensional hand pose information of the sample video, and the three-dimensional hand pose information of the sample video;
[0012] Based on the two-dimensional pose recognition model included in the initial pose recognition model, the first video frame of the sample video is processed to determine multiple two-dimensional prediction joints of the first video frame. These multiple two-dimensional prediction joints are used to describe the two-dimensional pose of the hand in the first video frame.
[0013] Multiple two-dimensional predicted joints of the first video frame are input into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joints of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0014] Based on the constraint loss function of the initial pose recognition model and multiple 3D prediction joints of the second video frame in the sample video, the multiple 3D prediction joints of the first video frame are processed so that the difference between the multiple 3D prediction joints of the second video frame and the multiple 3D prediction joints of the first video frame satisfies the target condition. The second video frame is a video frame preceding the first video frame.
[0015] Based on the two-dimensional hand posture information, the three-dimensional hand reference information, multiple two-dimensional predicted joints of the first video frame, and multiple three-dimensional predicted joints of the processed first video frame, the model parameters of the initial posture recognition model are adjusted to obtain a hand posture recognition model. The model parameters include the parameters of the two-dimensional posture recognition model and the parameters of the three-dimensional posture recognition model.
[0016] On the one hand, a hand gesture recognition device is provided, the device comprising:
[0017] The two-dimensional recognition module is used to process the first video frame of the video based on the two-dimensional pose recognition model to determine multiple two-dimensional hand joints of the first video frame. These multiple two-dimensional hand joints are used to describe the two-dimensional pose of the hand in the first video frame.
[0018] The three-dimensional recognition module is used to determine multiple three-dimensional hand joints of the first video frame based on multiple two-dimensional hand joints and a three-dimensional pose recognition model, and the three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0019] The constraint module is used to process multiple three-dimensional hand joints of the first video frame based on the constraint loss function and multiple three-dimensional hand joints of the second video frame in the video, so that the difference between the multiple three-dimensional hand joints of the second video frame and the multiple three-dimensional hand joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0020] In one possible implementation, the constraint module includes:
[0021] The determining unit is used to determine, from the multiple three-dimensional hand joints of the first video frame, the first joint that is not occluded by other three-dimensional hand joints;
[0022] The adjustment unit is used to adjust the coordinates of the multiple first joints based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0023] In one possible implementation, the adjustment unit is used for:
[0024] Based on the multiple second joints and the multiple first joints, determine the range of change between the second joints and the first joints;
[0025] If the change is less than the first threshold, the coordinates of the multiple first joints are adjusted based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0026] In one possible implementation, the adjustment unit is used for:
[0027] Based on the distance between the multiple second joints and the multiple first joints and the constraint weights of the constraint loss function, the loss value of the constraint loss function is determined. The constraint weights are used to control the magnitude of the adjustment of the coordinates of the multiple first joints.
[0028] Based on this loss value, the coordinates of the multiple first joints are adjusted to reduce the difference between the coordinates of the first joint and the coordinates of the second joint.
[0029] In one possible implementation, the constraint module is used for:
[0030] Based on multiple three-dimensional hand joints in the first video frame, determine the hand pose type that matches the hand in the first video frame;
[0031] From the multiple three-dimensional hand joints of the first video frame, determine multiple third joints that are related to the hand pose type;
[0032] Based on the multiple third joints corresponding to the third joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, the coordinates of the multiple third joints are adjusted so that the difference between the coordinates of the third joint and the coordinates of the fourth joint is reduced.
[0033] In one possible implementation, the 3D recognition module is used for:
[0034] The pose estimation parameters of the first video frame are obtained by inputting multiple two-dimensional hand joints of the first video frame into the parameter prediction model of the three-dimensional pose recognition model.
[0035] The pose estimation parameters are input into the hand skeleton model of the 3D pose recognition model to obtain multiple 3D hand joints of the first video frame. These multiple 3D hand joints of the first video frame are used to describe the 3D pose of the hand in the first video frame.
[0036] In one possible implementation, the pose estimation parameters include camera parameters and joint rotation parameters. The camera parameters indicate the three-dimensional space in which the hand is located in the first video frame, and the joint rotation parameters indicate how the hand skeleton model determines the deformation of the three-dimensional pose of the hand in the first video frame.
[0037] The device also includes an optimization module, which is used for:
[0038] Based on the multiple three-dimensional hand joints, the camera parameters, and the first video frame, multiple two-dimensional projection joints are determined;
[0039] Based on the multiple two-dimensional projection joints and the multiple two-dimensional hand joints, a projection loss value is determined, which indicates the error between the two-dimensional projection joints and the multiple two-dimensional hand joints;
[0040] Based on the projection loss value, adjust the camera parameters and the joint rotation parameters.
[0041] In one possible implementation, the optimization module is used to:
[0042] Based on the location information of the hand skeleton model, the multiple three-dimensional hand joints of the first video frame are divided into multiple groups of joints with a processing order.
[0043] If the projection loss value determined based on the adjusted camera parameters meets the optimization conditions, then according to the processing order among the multiple sets of joints, the joint rotation parameters corresponding to the multiple sets of joints are adjusted sequentially based on the adjusted camera parameters, the multiple sets of joints, and the first video frame, so as to reduce the projection loss value.
[0044] In one possible implementation, the device further includes a model acquisition module, which is configured to:
[0045] Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined.
[0046] Obtain at least one three-dimensional pose recognition model that matches the chiral features of at least one hand.
[0047] On the one hand, a training device for a hand pose recognition model is provided, the device comprising:
[0048] The acquisition module is used to acquire a sample video of the hand, the two-dimensional hand pose information of the sample video, and the three-dimensional hand pose information of the sample video.
[0049] The two-dimensional prediction module is used to process the first video frame of the sample video based on the two-dimensional pose recognition model included in the initial pose recognition model, so as to determine multiple two-dimensional prediction joints of the first video frame, which are used to describe the two-dimensional pose of the hand in the first video frame.
[0050] The three-dimensional prediction module is used to input multiple two-dimensional predicted joints of the first video frame into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joints of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0051] The constraint module is used to process the multiple three-dimensional prediction joints of the first video frame based on the constraint loss function of the initial pose recognition model and the multiple three-dimensional prediction joints of the second video frame in the sample video, so that the difference between the multiple three-dimensional prediction joints of the second video frame and the multiple three-dimensional prediction joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0052] The adjustment module is used to adjust the model parameters of the initial pose recognition model based on the two-dimensional hand pose information, the three-dimensional hand reference information, multiple two-dimensional predicted joints of the first video frame and multiple three-dimensional predicted joints of the processed first video frame, to obtain a hand pose recognition model. The model parameters include the parameters of the two-dimensional pose recognition model and the parameters of the three-dimensional pose recognition model.
[0053] In one possible implementation, the parameters of the constraint loss function include constraint weights, which are used to control the magnitude of adjustment of the coordinates of the plurality of 3D prediction joints when processing the plurality of 3D prediction joints of the first video frame.
[0054] In one possible implementation, the adjustment module is used to:
[0055] Based on the two-dimensional hand pose information and multiple two-dimensional hand joints of the first video frame, a first loss value is determined. The first loss value indicates the error of the two-dimensional pose recognition model in predicting the two-dimensional pose of the hand in the first video frame.
[0056] Based on the three-dimensional hand pose information and multiple three-dimensional hand joints in the first video frame, a second loss value is determined. The second loss value indicates the error of the three-dimensional pose recognition model in predicting the three-dimensional pose of the hand in the first video frame.
[0057] Based on the first loss value and the second loss value, the model parameters of the initial pose recognition model are adjusted so that the first loss value and the second loss value determined based on the adjusted model parameters are reduced, thereby obtaining the hand pose recognition model.
[0058] In one possible implementation, the device further includes a model acquisition module, which is used for:
[0059] Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined.
[0060] Obtain at least one three-dimensional pose recognition model that matches the chiral features of at least one hand.
[0061] On the one hand, a computer device is provided, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the above-described hand gesture recognition method, or a training method for a hand gesture recognition model.
[0062] On the one hand, a computer-readable storage medium is provided, which stores at least one computer program that is loaded and executed by a processor to implement the above-described hand pose recognition method, or a training method for a hand pose recognition model.
[0063] On the one hand, a computer program product or computer program is provided, which includes program code stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the aforementioned hand gesture recognition method, or the training method for a hand gesture recognition model.
[0064] The technical solution provided in this application, when performing hand pose recognition on video, constrains the differences between the three-dimensional hand joints of multiple video frames to refer to the correlation between the hands in multiple video frames, ensuring that the hand pose determined for the video is coherent and smooth, and greatly improving the stability of the hand pose recognition results. Attached Figure Description
[0065] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0066] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0067] Figure 2 This is a flowchart of a hand gesture recognition method provided in an embodiment of this application;
[0068] Figure 3 This is a schematic diagram of a two-dimensional hand joint provided in an embodiment of this application;
[0069] Figure 4 This is a flowchart of a hand gesture recognition method provided in an embodiment of this application;
[0070] Figure 5 This is a flowchart illustrating the determination of two-dimensional hand joint points according to an embodiment of this application;
[0071] Figure 6 This is a schematic diagram of a hand skeleton model provided in an embodiment of this application;
[0072] Figure 7 This is a flowchart of a joint optimization provided in an embodiment of this application;
[0073] Figure 8 This is a schematic diagram of a hand gesture recognition method provided in an embodiment of this application;
[0074] Figure 9 This is a flowchart of a training method for a hand pose recognition model provided in an embodiment of this application;
[0075] Figure 10 This is a schematic diagram of the structure of a hand posture recognition device provided in an embodiment of this application;
[0076] Figure 11 This is a schematic diagram of the structure of a training device for a hand pose recognition model provided in an embodiment of this application;
[0077] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0078] Figure 13 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0080] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.
[0081] In this application, the term "at least one" means one or more, and "multiple" means two or more, for example, multiple pictures means two or more pictures.
[0082] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the videos and sample videos involved in this application were obtained with full authorization.
[0083] The technical solution provided in this application relates to the field of artificial intelligence and can be applied to various scenarios such as image processing, cloud technology, and big data.
[0084] Artificial Intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning.
[0085] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, character recognition, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, and simultaneous localization and mapping (SLAM).
[0086] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.
[0087] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0088] The implementation environment of the embodiments of this application will be described next.
[0089] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. See also... Figure 1 , Figure 1 The implementation environment shown can be used for training methods or models of hand pose recognition. See [link to relevant documentation]. Figure 1 The implementation environment includes: terminal 101 and server 102.
[0090] The terminal 101 is used to execute the hand gesture recognition method provided in this application embodiment for a video or image, so as to simulate the hand gesture in the video or image.
[0091] In some embodiments, the terminal 101 can perform a hand gesture recognition method based on real-time acquired video. Optionally, the terminal 101 is equipped with an application for implementing the hand gesture recognition method provided in this application, so as to realize the real-time hand gesture recognition process on a mobile device.
[0092] The server 102 is used to execute the training method of the hand posture recognition model provided in the embodiments of this application, thereby providing a hand posture recognition model that can be deployed in the terminal 101 to implement the hand posture recognition method.
[0093] In some embodiments, the hand pose recognition method and the hand pose recognition model training method provided in this application can also be jointly executed by the terminal 101 and the server 102, and this application does not limit this. When the hand pose recognition method is jointly executed by the terminal 101 and the server 102, the terminal 102 can send the real-time acquired video or image to the server 102, and simulate the hand pose in the video or image based on the hand pose recognition result returned by the server 102. When the hand pose recognition model training method is jointly executed by the terminal 101 and the server 102, the terminal 102 can send the real-time acquired video or image to the server 102, thereby providing the server 102 with sample videos for training the hand pose recognition model. In the case where the terminal 101 and the server 102 jointly execute any of the above methods, the server 102 may undertake the main computing work and the terminal 101 may undertake the secondary computing work; or, the server 102 may undertake the secondary computing work and the terminal 101 may undertake the main computing work; or, the server 102 and the terminal 101 may use a distributed computing architecture to perform collaborative computing. This application embodiment does not limit this.
[0094] In some embodiments, the terminal 101 can be connected to the server 102 via a wireless network or a wired network. Optionally, the terminal 101 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Optionally, the server 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0095] In other embodiments, the server 102 can be a node in a blockchain system. After the server 102 trains the hand gesture recognition model, the server 102 can publish the hand gesture recognition model to the blockchain system, that is, store it in the blockchain in the form of blocks, so that other nodes in the blockchain system can apply the hand gesture recognition model.
[0096] After introducing the implementation environment of the embodiments of this application, the hand posture recognition method provided by the embodiments of this application will be described below. Figure 2 This is a flowchart of a hand gesture recognition method provided in an embodiment of this application. The method is executed by a computer device, which can be the aforementioned terminal 101 or server 102. (See also...) Figure 2 The method includes the following steps 201 to 203.
[0097] 201. A computer device processes the first video frame of a video based on a two-dimensional pose recognition model to determine multiple two-dimensional hand joints in the first video frame. These multiple two-dimensional hand joints are used to describe the two-dimensional pose of the hand in the first video frame.
[0098] The video focuses on the hand, meaning that multiple video frames contain the hand. In some embodiments, the hand can belong to any target object with a hand gesture, such as a human, animal, or virtual cartoon character; this application does not limit this. Accordingly, the two-dimensional pose recognition model involved in the embodiments of this application targets the hand skeleton of the target object, such as the human hand skeleton.
[0099] In the embodiments of this application, hand joints are key points used to describe the shape of the hand bones. That is, the shape of the hand bones can be determined based on the positions of multiple hand joints, and the shape of the hand bones is the basis for determining the hand posture.
[0100] In this embodiment, the two-dimensional hand joints are key points used to describe the shape of the hand bones in a two-dimensional image. The computer device processes the first video frame based on the two-dimensional pose recognition model to determine the position of each joint of the hand in the first video frame (two-dimensional image), that is, the coordinates of the multiple two-dimensional hand joints, thereby determining the two-dimensional pose of the hand in the first video frame.
[0101] For ease of understanding, this application provides a schematic diagram of two-dimensional hand joints, see [link / reference]. Figure 3The human hand comprises 21 two-dimensional hand joints, each numbered in sequence corresponding to the hand parts: the joint at the base of the palm is numbered "0"; the four joints of the thumb are numbered "1, 2, 3, 4"; the four joints of the index finger are numbered "5, 6, 7, 8"; the four joints of the middle finger are numbered "9, 10, 11, 12"; the four joints of the ring finger are numbered "13, 14, 15, 16"; and the four joints of the little finger are numbered "17, 18, 19, 20".
[0102] 202. The computer device determines multiple three-dimensional hand joints of the first video frame based on multiple two-dimensional hand joints and a three-dimensional pose recognition model, and the three-dimensional pose recognition model matches the chiral features of the hand in the first video frame.
[0103] Chirality refers to the property that an object cannot be superimposed on its mirror image, describing a structural symmetry of the object. For example, both the human left and right hands are chiral. The chiral feature in this application refers to the type of chirality that a hand conforms to; for example, the left hand conforms to left-handed chirality, and the right hand conforms to right-handed chirality.
[0104] In some embodiments, if the first video frame contains multiple hands with different chiral characteristics, the computer device acquires multiple three-dimensional pose recognition models for the different chiral characteristics to process the two-dimensional hand joints corresponding to the hands with different chiral characteristics respectively. For example, if the first video frame includes an image of the left hand and an image of the right hand, the computer device processes the multiple two-dimensional hand joints corresponding to the left hand based on the three-dimensional pose recognition model of the left hand; and processes the multiple two-dimensional hand joints corresponding to the right hand based on the three-dimensional pose recognition model of the right hand.
[0105] In this embodiment, since the hand has its chiral characteristics, for the hand in the first video frame, this application uses a three-dimensional pose recognition model that matches the chiral characteristics to determine the three-dimensional hand joints, which can improve the accuracy and efficiency of recognizing hand poses.
[0106] 203. The computer device processes the multiple three-dimensional hand joints of the first video frame based on the constraint loss function and the multiple three-dimensional hand joints of the second video frame in the video, so that the difference between the multiple three-dimensional hand joints of the second video frame and the multiple three-dimensional hand joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0107] The loss value of the constraint loss function represents the difference between multiple 3D hand joints in the second video frame and multiple 3D hand joints in the first video frame, i.e., the degree of similarity. Therefore, the constraint loss function can guide the computer device to process the multiple 3D hand joints in the first video frame to ensure that the difference meets the target condition, i.e., that the degree of similarity guarantees the continuity of the hand poses between the two video frames.
[0108] In some embodiments, the objective condition may be: the loss value of the constraint loss function is less than the objective value.
[0109] In this embodiment, considering that the hands in multiple video frames are actually related in terms of movement, by referring to the second video frame whose hand posture has been identified before, and using a constraint loss function to constrain the difference between the three-dimensional hand joints of the current first video frame and the three-dimensional hand joints of the second video frame, the hand posture simulated for the video will not have problems such as flickering or shaking.
[0110] By using the above technical solution, when performing hand posture recognition on video, the correlation between hand movements between video frames is taken into account. By constraining the differences between the three-dimensional hand joints of the video frames, the hand posture simulated from the video will not have problems such as flickering or shaking, ensuring that the hand posture determined from the video is smooth and continuous, and greatly improving the stability of the hand posture recognition results.
[0111] The above steps 201 to 203 provide a brief overview of the hand posture recognition method provided in the embodiments of this application. The hand posture recognition method provided in the embodiments of this application will be described in detail below. Figure 4 This is a flowchart of a hand gesture recognition method provided in an embodiment of this application. The method is executed by a computer device, which can be the aforementioned terminal 101 or server 102. (See also...) Figure 4 The method includes the following steps 401 to 406.
[0112] 401. Computer equipment acquires video.
[0113] This video focuses on the hands; for an explanation of the video, please refer to step 201, which will not be repeated here.
[0114] In some embodiments, the computer device is a terminal, and the video may be a video stored in the terminal. The terminal can read the video from the corresponding storage space in response to a hand gesture recognition command for a specified video. The video may also be a video recorded in real time by the terminal. For example, the terminal may have an application that supports real-time hand gesture recognition, which records the video to be recognized in real time. This application does not limit the method of video acquisition.
[0115] In other embodiments, the computer device is a server, and the video can be a video stored on the server or a video to be used for hand gesture recognition obtained by the server from any terminal. This application does not limit this.
[0116] It should be noted that the computer device in this application acquires video with the user's full authorization. For example, before reading or starting to record video, the terminal displays an authorization prompt, including "Do you allow reading / recording video?", and if the user selects the "agree" option, the terminal acquires the video. As another example, when the user's terminal sends video to the server, an authorization prompt is displayed, including "Do you allow uploading video files to use online services?", and if the user selects the "agree" option, the terminal sends the video to the server so that the server can acquire it.
[0117] 402. The computer device processes the first video frame of the video based on a two-dimensional pose recognition model to determine multiple two-dimensional hand joints of the first video frame, which are used to describe the two-dimensional pose of the hand in the first video frame.
[0118] This step is the same as step 201, and will not be repeated here.
[0119] In some embodiments, the process by which a computer device processes the first video frame based on the two-dimensional pose recognition model includes the following steps 2-1 and 2-2.
[0120] Step 2-1: The computer device performs hand detection on the first video frame of the video based on a two-dimensional pose recognition model to determine at least one hand detection box.
[0121] In some embodiments, hand detection refers to: extracting features from the first video frame, determining the image region containing the hand in the first video frame based on the image features of the first video frame, and then marking the image region with a hand detection bounding box. In some embodiments, the feature extraction process targets visual cues related to the hand in the first video frame, such as the palm outline. In this example, the computer device extracts features from the first video frame based on the two-dimensional pose recognition model, preliminarily determines the position of the palm in the first video frame based on the extracted palm outline features, and then determines the position of the entire hand based on the position of the palm in the first video frame. Optionally, the hand detection bounding box is a rectangle, and the coordinates of the hand detection bounding box in the first video frame (also called rectangle coordinates) can be used to determine the coordinates of multiple two-dimensional hand joints of the hand.
[0122] In some embodiments, the two-dimensional pose recognition model used by the computer device can be trained based on any neural network model that can identify multiple two-dimensional hand joints from a two-dimensional image, such as the OpenPose model or the BlazePalm model. This application embodiment does not limit this.
[0123] Compared to directly detecting multiple joints of the fingers in the hand, the above technical solution can effectively reduce the difficulty of determining the hand position by determining the position of objects with relatively fixed boundaries, such as the palm or fist, in the first video frame, and can also ensure the accuracy of detection to a certain extent.
[0124] Step 2-2: Based on the two-dimensional pose recognition model, the computer device performs key point detection on the hand image within the hand detection box to determine multiple two-dimensional hand joints in the first video frame.
[0125] In some embodiments, the two-dimensional pose recognition model includes a two-dimensional hand joint model for predicting the plurality of two-dimensional hand joints (reference). Figure 3 This two-dimensional hand joint model indicates the position of each two-dimensional hand joint within the hand skeleton. For example, using... Figure 3 Taking the provided 21 two-dimensional hand joint model as an example, the computer device performs feature learning on the hand image within the hand detection box, and based on the learning results of the hand image, predicts the coordinates of the 21 two-dimensional hand joints according to the two-dimensional hand joint model and the coordinates of the hand detection box in the first video frame, thereby obtaining the 21 two-dimensional hand joints of the first video frame.
[0126] The above process is illustrated using a hand in the first video frame as an example. When there are multiple hands in the first video frame, the computer device can detect multiple hand detection boxes, thereby determining the two-dimensional hand joints corresponding to the hand images in the multiple hand detection boxes based on the method provided above. This application embodiment will not elaborate on this.
[0127] In the method described above, the computer device can perform hand detection on each frame of the video using the two-dimensional pose recognition model to determine the hand detection box. In other embodiments, the computer device can perform hand detection every N frames. For multiple intermediate frames between two detections, target tracking is performed based on the hand detection box of the previous frame to obtain the hand detection box of the intermediate frame. Here, N is a positive integer, for example, 15. In this example, step 2-1 above can be replaced by: if a hand detection box exists in the frame preceding the first video frame, the computer device determines the hand detection box of the first video frame based on the hand detection box of the previous frame and the two-dimensional pose recognition model. Optionally, the two-dimensional pose recognition model can be used to implement target tracking. For example, the two-dimensional pose recognition model is trained based on the kernel correlation filter (KCF) algorithm.
[0128] It should be noted that the tracking involved in this application refers to the process of locating a specified object in consecutive video frames in the field of computer vision. For example, the specified object in this application is a hand in a video frame.
[0129] For ease of understanding, this application provides a flowchart for determining two-dimensional hand joint points, see [link / reference]. Figure 5 The process involves several steps: First, inter-frame control is performed on the first video frame 501 to determine if it is an intermediate frame. If it is, hand tracking is performed based on the previous frame to obtain a hand detection box 502 for that frame. If the first video frame 501 is not an intermediate frame, hand detection is performed on it to obtain the hand detection box 502, which is then used to update the parameters involved in the hand tracking process. Keypoint detection is then performed based on the hand detection box 502 to obtain multiple two-dimensional hand joints, and these joints are used to update the parameters involved in the hand tracking process. The keypoint detection result 503 includes multiple two-dimensional hand joints, each based on... Figure 3 The provided numbering indication.
[0130] The above technical solution can effectively improve the speed of determining two-dimensional hand joints in the scenario of real-time hand pose recognition for video, and further enhance the coherence between hand detection results of multiple video frames.
[0131] In other embodiments, during the process of determining the plurality of two-dimensional hand joints, the computer device performs chirality prediction on the first video frame to determine a three-dimensional pose recognition model matching the chiral features of the hand in the first video frame. This process includes steps 2-3 and 2-4 below. Specifically, the computer device can execute steps 2-3 after performing step 2-1 above.
[0132] Steps 2-3: When the computer device detects at least one hand in the first video frame, it determines the chiral features of the at least one hand based on the at least one hand detection frame.
[0133] The definition of chiral characteristics is given in step 202 and will not be repeated here.
[0134] In some embodiments, the computer device can predict the chirality of the hand based on features of the hand image within the hand detection frame. For example, the computer device can classify the hand image based on the overall outline of the hand in the hand image and output the probability that the hand image belongs to the right hand or the left hand, thereby determining the chirality of the hand.
[0135] It should be noted that step 2-3 can be performed after step 2-1 is completed.
[0136] Steps 2-4: The computer device acquires at least one 3D pose recognition model that matches the chiral features of the at least one hand. Refer to step 202, which will not be repeated here.
[0137] Through the above technical solution, this application identifies chiral features during the two-dimensional posture recognition process, thereby pre-obtaining a three-dimensional posture recognition model that matches the chiral features, effectively improving the accuracy and efficiency of hand posture recognition.
[0138] 403. The computer device inputs multiple two-dimensional hand joints of the first video frame into the parameter prediction model of the three-dimensional pose recognition model to obtain the pose estimation parameters of the first video frame.
[0139] In some embodiments, the computer device determines a multidimensional array to indicate the hand in the first video frame based on the plurality of two-dimensional hand joints, and then inputs the multidimensional array into a parameter prediction model. Optionally, the computer device can determine a tensor matrix of the first video frame based on the two-dimensional hand joints, and then input the tensor matrix into the parameter prediction model to calculate parameters describing the hand from a higher dimension (three dimensions), i.e., the pose estimation parameters.
[0140] In some embodiments, changes in hand posture are achieved based on traction and movement between joints. Therefore, the computer device determines the posture estimation parameters based solely on the joints among the plurality of two-dimensional hand joints that can drive the rotation of other joints. For example, Figure 3 The two-dimensional hand joint "0" corresponding to the base of the palm and 15 two-dimensional hand joints "1, 2, 3, 5, 6, 7, 9, 10, 11, 13, 14, 15, 17, 18, 19" which are not fingertips.
[0141] In some embodiments, the pose estimation parameters include camera parameters and joint rotation parameters. The camera parameters indicate the three-dimensional space in which the hand is located in the first video frame, and the joint rotation parameters indicate how the hand skeleton model deforms to determine the three-dimensional pose of the hand in the first video frame. Exemplarily, the pose estimation parameters can be output in the form of a multi-dimensional feature vector.
[0142] In some embodiments, multiple hands in the first video frame have different chiral features. The computer device inputs the two-dimensional hand joints corresponding to the hands with different chiral features into the parameter prediction model of the three-dimensional pose recognition model corresponding to the chiral features, thereby determining the pose estimation parameters corresponding to the hands with different chiral features respectively.
[0143] 404. The computer device inputs the pose estimation parameters into the hand skeleton model of the three-dimensional pose recognition model to obtain multiple three-dimensional hand joints of the first video frame. The multiple three-dimensional hand joints of the first video frame are used to describe the three-dimensional pose of the hand in the first video frame.
[0144] This hand skeleton model provides a method for three-dimensional modeling of hand posture. For example, the hand skeleton model provides the basic three-dimensional structure of the hand skeleton, as shown below. Figure 6 The hand skeleton model is parameterized according to forward kinematics. Based on this, the hand skeleton model can determine the deformation mode of the hand skeleton according to forward kinematics based on the joint acceleration and joint constraint force provided by the input joint rotation parameters, thereby determining the coordinates of the multiple three-dimensional hand joints.
[0145] In this embodiment, the coordinates of the two-dimensional hand joints indicate the positions of each joint in a two-dimensional image. The process of determining the three-dimensional hand joints includes predicting the depth information of each joint in three-dimensional space based on its position in the two-dimensional image, thereby completing the process of abstracting the hand from two-dimensional to three-dimensional. For example, the coordinates of the three-dimensional hand joints can be (x, y, z), where x and y can be determined based on camera parameters and the coordinates of the two-dimensional hand joints, and z can be determined based on depth information. In other embodiments, the two-dimensional hand joints can carry depth-related information, which is used to predict depth information in three-dimensional space during the determination of the three-dimensional hand joints.
[0146] In other embodiments, the three-dimensional hand joints may also be associated with more dimensions of information such as color and reflection intensity, which is not limited in this application.
[0147] In some embodiments, multiple hands in the first video frame have different chiral features. The computer device then inputs the pose estimation parameters corresponding to the hands with different chiral features into the corresponding hand skeleton models. For example, the pose estimation parameters corresponding to the left hand in the first video frame are input into the hand skeleton model corresponding to the left hand. For ease of understanding, this application provides a schematic diagram of a hand skeleton model, see [link to diagram]. Figure 6 , where 601 is the three-dimensional right hand skeleton model, 602 is the three-dimensional left hand skeleton model, and 601 and 602 are in a three-dimensional skeleton coordinate system with 603 as the coordinate axis.
[0148] 405. The computer device processes multiple three-dimensional hand joints of the first video frame based on the constraint loss function and multiple three-dimensional hand joints of the second video frame in the video, so that the difference between the multiple three-dimensional hand joints of the second video frame and the multiple three-dimensional hand joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0149] This step refers to step 203.
[0150] In this embodiment, the first video frame is the M-th frame of the video, where M is an integer greater than 1. In some embodiments, the second video frame is the frame preceding the first video frame. The computer device can obtain the second video frame preceding the first video frame based on the timestamp of the first video frame. In other embodiments, the computer device numbers the multiple video frames based on the timestamps of each video frame in the video, thereby obtaining the second video frame preceding the first video frame according to the order of the numbers. This application does not limit the method of obtaining the second video frame.
[0151] In some other embodiments, the first video frame is the first frame of the video, and the computer device executes step 406 directly after executing step 404.
[0152] In some embodiments, the computer device processes multiple unobstructed three-dimensional hand joints in the first video frame, which can effectively reduce the computational load while improving the stability of hand pose recognition results. In this example, the process of the computer device processing multiple three-dimensional hand joints in the first video frame includes the following steps 5-1 and 5-2, that is, the following steps 5-1 and 5-2 are a possible implementation of step 405.
[0153] Step 5-1: The computer device determines the first joint point that is not occluded by other three-dimensional hand joint points from among the multiple three-dimensional hand joint points in the first video frame.
[0154] The computer device can determine the occlusion relationship between the plurality of three-dimensional hand joints based on their coordinates. In some embodiments, the computer device matches the individual three-dimensional hand joints based on their coordinates, thereby determining the occlusion relationship based on the degree of matching. For example, the degree of matching between two three-dimensional hand joints can be determined based on cosine similarity or Pearson coefficient. In this example, the computer device can identify three-dimensional hand joints whose degree of matching with other three-dimensional hand joints is lower than a target threshold as the first joint.
[0155] Step 5-2: The computer device adjusts the coordinates of the multiple first joints based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0156] In this embodiment, the computer device determines the second joint point based on the number of the first joint point in the hand skeleton model. In some embodiments, multiple three-dimensional hand joint points in any video frame are determined according to... Figure 3 The provided method for numbering the two-dimensional hand joints corresponds to numbers 0-20. If the first joint is numbered 1, then the second joint is the three-dimensional hand joint numbered 1 among the multiple three-dimensional hand joints corresponding to the second video frame. That is, the first joint and the second joint correspond to the same joint in the same hand skeleton model.
[0157] In some embodiments, considering that the range of motion changes between the second video frame and the first video frame may be large, in order to avoid affecting the actual change effect of the motion due to ensuring continuity, step 5-2 can be implemented by the following steps 5-2-1 and 5-2-2.
[0158] Step 5-2-1: The computer device determines the range of change between the second joint point and the first joint point based on the multiple second joint points and the multiple first joint points.
[0159] The magnitude of this change indicates the magnitude of the hand posture change between the second video frame and the first video frame.
[0160] In some embodiments, the computer device can determine the magnitude of the change based on the difference between the coordinates of the second joint and the coordinates of the first joint.
[0161] Step 5-2-2: When the change amplitude is less than the first threshold, the computer device adjusts the coordinates of the multiple first joints based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0162] By judging the change in joint position between consecutive video frames before adjusting the coordinates of the 3D hand joints, it is possible to ensure that the position of the 3D hand joints keeps up with the changes in the gesture when the gesture changes significantly between video frames.
[0163] In some embodiments, the computer device performs the process of adjusting the coordinates of the plurality of first joint points in step 5-2 or step 5-2-2 by the following steps (1) and (2).
[0164] Step (1): The computer device determines the loss value of the constraint loss function based on the distance between the multiple second joints and the multiple first joints and the constraint weight of the constraint loss function. The constraint weight is used to control the magnitude of the adjustment of the coordinates of the multiple first joints.
[0165] In some embodiments, this step (1) can be implemented by the following formula (1).
[0166]
[0167] In formula (1), N is the number of first joints, which is a positive integer; L1 is the constraint loss function; and w1 is the constraint weight. Let be the coordinates of the i-th first joint point in the three-dimensional skeleton coordinate system. Let be the coordinates of the i-th second joint in the three-dimensional skeleton coordinate system.
[0168] In some embodiments, the constraint weight can be set to a fixed value, for example, 0.1.
[0169] Step (2): Based on the loss value, the computer device adjusts the coordinates of the multiple first joints to reduce the difference between the coordinates of the first joint and the coordinates of the second joint.
[0170] In some embodiments, the computer device adjusts the coordinates of the first joint point so that the loss value recalculated based on the adjusted coordinates of the first joint point and the constraint loss function is less than a second threshold. The second threshold is used to measure whether the difference between the first joint point and the second joint point can guarantee that the hand posture determined based on the first joint point and the hand posture determined based on the second joint point are consistent in terms of motion.
[0171] In other embodiments, the computer device processes multiple three-dimensional hand joints in the first video frame that match any preset hand pose type. This can effectively reduce the computational load while improving the stability of the hand pose recognition results, and also improve the effect of the processed three-dimensional hand joints in representing actual hand movements. In this example, the process by which the computer device processes multiple three-dimensional hand joints in the first video frame includes the following steps A to C, that is, steps A to C are another possible implementation of step 405.
[0172] Step A: Based on multiple 3D hand joints in the first video frame, determine the hand pose type that matches the hand in the first video frame.
[0173] In some embodiments, the computer device can acquire multiple preset hand gesture types, such as a clenched fist gesture or an OK gesture. The computer device can match the coordinates of the multiple three-dimensional hand joints with the multiple hand gesture types to initially determine the three-dimensional hand gesture in the first video frame, and then adjust the multiple three-dimensional hand joints with the determined hand gesture type as a reference.
[0174] Step B: From the multiple three-dimensional hand joints in the first video frame, determine multiple third joints that are related to the hand pose type.
[0175] In some embodiments, the hand gesture type corresponds to preset gesture key points, which are associated with the gesture corresponding to that hand gesture type. For example, the gesture key points of the OK gesture include the joints of the thumb and index finger. Based on this, the computer device can determine the plurality of third joints based on the number of the gesture key points in the hand skeleton model.
[0176] Step C: Based on the multiple fourth joints corresponding to the third joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, adjust the coordinates of the multiple third joints to reduce the difference between the coordinates of the third joint and the coordinates of the fourth joint.
[0177] This step refers to steps (1) and (2) above, and will not be repeated here.
[0178] By taking into account the correlation of hand movements between video frames and constraining the differences between the three-dimensional hand joints in the video frames, the hand posture determined from the video is made coherent and smooth, which greatly improves the stability of the hand posture recognition results.
[0179] 406. The computer device optimizes the processed multiple three-dimensional hand joints based on the multiple two-dimensional hand joints.
[0180] In this embodiment of the application, the computer device will also refer to the recognition results of the two-dimensional hand posture, that is, the multiple two-dimensional hand joints, to further optimize the multiple three-dimensional hand joints.
[0181] In some embodiments, the computer device determines a plurality of two-dimensional projection joints based on the plurality of three-dimensional hand joints, the camera parameters, and the first video frame. For example, the computer device projects the plurality of three-dimensional hand joints onto the first video frame based on the camera parameters to obtain the two-dimensional projection joints corresponding to the three-dimensional hand joints. These two-dimensional projection joints can represent the position of the three-dimensional pose predicted by the three-dimensional pose recognition model for the hand in the first video frame within a two-dimensional image.
[0182] In some embodiments, the computer device determines a projection loss value based on the plurality of two-dimensional projection joints and the plurality of two-dimensional hand joints, the projection loss value indicating the error between the two-dimensional projection joints and the plurality of two-dimensional hand joints. In some embodiments, the projection loss value can be determined based on the following formula (2).
[0183]
[0184] In formula (2), N is the number of projection joints in the second dimension, which is a positive integer, and L2 is the projection loss value. These are the coordinates of the two-dimensional projection joints in the two-dimensional image. These are the coordinates of the second-dimensional hand joint points in the two-dimensional image.
[0185] In some embodiments, the computer device adjusts the camera parameters and the joint rotation parameters based on the projection loss value.
[0186] In some embodiments, the process of adjusting the camera parameters and the joint rotation parameters based on the projection loss value can be achieved through the following steps 6-1 and 6-2.
[0187] Step 6-1: The computer device divides the multiple three-dimensional hand joints of the first video frame into multiple groups of joints with a processing order according to the part information of the hand skeleton model.
[0188] In some embodiments, the location information indicates different locations in the hand skeleton model; for example, the location information indicates the joints corresponding to the thumb, index finger, middle finger, ring finger, and little finger, respectively. In this example, the computer device can divide the multiple three-dimensional hand joints into multiple groups of joints in the order of thumb, index finger, middle finger, ring finger, and little finger based on the location information.
[0189] In some embodiments, the joint numbering order corresponding to the hand skeleton model (refer to...) Figure 3 If the information is determined based on the location, then the computer device can directly group the multiple three-dimensional hand joints according to the joint number sequence corresponding to the hand skeleton model. For example, number 1-3 is the first group, number 5-7 is the second group, number 9-11 is the third group, number 13-15 is the fourth group, and number 17-19 is the fifth group.
[0190] Step 6-2: When the projection loss value determined based on the adjusted camera parameters meets the optimization conditions, the computer device, according to the processing order among the multiple sets of joints, adjusts the joint rotation parameters corresponding to the multiple sets of joints in sequence based on the adjusted camera parameters, the multiple sets of joints, and the first video frame, so as to reduce the projection loss value.
[0191] In this embodiment, the computer device employs a step-by-step optimization strategy, first adjusting the camera parameters, and then adjusting the joint rotation parameters based on the adjusted camera parameters. First, the computer device adjusts the camera parameters based on the aforementioned projection loss value to obtain target camera parameters that satisfy the optimization condition, where the optimization condition may be that the projection loss value is less than a third threshold. Second, following this processing order, the computer device projects each group of joints back to the first video frame based on the target camera parameters to obtain the two-dimensional projected joints corresponding to that group of joints. Then, following steps similar to those in formula (2), it determines the projection loss value corresponding to this optimization of that group of joints. By adjusting the joint rotation parameters corresponding to that group of joints, the projection loss value is reduced.
[0192] The above process is illustrated using only one adjustment as an example. Step 6-2 includes multiple adjustments for the multiple sets of joints. In some embodiments, when the projection loss value corresponding to the multiple sets of joints is less than the fourth threshold, the computer device completes the optimization process for the multiple three-dimensional hand joints.
[0193] In related technologies, parallel optimization of multiple 3D hand joints corresponding to the entire hand requires calculating the projection loss value for each 3D hand joint and updating the joint rotation parameters of each 3D hand joint in each optimization iteration. This makes obtaining the optimization result very difficult, resulting in a large computational load and reducing the optimization speed. The technical solution provided in this application groups multiple 3D hand joints, allowing for separate processing of the joints corresponding to each finger, thus greatly reducing the difficulty of determining the optimization result and effectively improving the speed of optimizing 3D hand joints. For example, when the technical solution provided in this application is executed by a computer device configured with a 6-core processor at 2.6 GHz, the optimization time for 3D hand joints can be reduced from 4 milliseconds to 2.5 milliseconds. Therefore, the hand posture recognition method provided in this application can be deployed in mobile devices to efficiently achieve real-time hand posture recognition.
[0194] In some embodiments, steps 405 and 406 above are equivalent to jointly optimizing the output of the 3D pose recognition model. To facilitate understanding of the above process, this application provides a flowchart of the joint optimization process, see below. Figure 7 The principle of this joint optimization process is as described above and will not be repeated here.
[0195] By using the above technical solution, the error caused by the 3D pose recognition model is corrected based on the projection loss value, which further improves the accuracy and stability of hand pose recognition.
[0196] In some embodiments, the computer device obtains the hand pose recognition result of the first video frame based on the optimized plurality of three-dimensional hand joints and the first video frame.
[0197] In some embodiments, the computer device obtains the deformation result of the hand skeleton based on the optimized three-dimensional hand joints, transforms the mesh vertices into the three-dimensional skeletal coordinate system based on the skin matrix, completes the process of the skeleton driving the skin change, and thus presents the hand posture in the three-dimensional skeletal coordinate system in the form of a hand posture model.
[0198] This application also provides a schematic diagram of a hand gesture recognition method to assist in illustrating the technical solutions described in steps 401 to 406 above. See [link / reference] Figure 8First, based on a two-dimensional pose recognition model, hand detection, key point detection, and chirality prediction are performed on the input video frames to obtain two-dimensional hand joints. Then, based on the two-dimensional hand joints and a three-dimensional pose recognition model (including a three-dimensional skeleton model of the left hand and a three-dimensional skeleton model of the right hand), pose estimation parameters (including camera parameters and joint rotation parameters) for determining the three-dimensional hand coordinates are obtained. Based on these pose estimation parameters and the two-dimensional hand joints, joint optimization is performed to finally obtain the three-dimensional hand joints, thereby establishing a hand pose model based on the three-dimensional hand joints.
[0199] In other embodiments, the computer device processes multiple video frames of the video based on the hand posture recognition method provided in steps 401 to 406 above, thereby completing the process of reconstructing the three-dimensional posture of the hand in the video. Then, based on the hand posture model in the three-dimensional skeletal coordinate system, it can realize various applications such as real-time hand special effects animation and real-time motion understanding.
[0200] By using the above technical solution, when performing hand posture recognition on video, the correlation between hand movements between video frames is taken into account. By constraining the differences between the three-dimensional hand joints of the video frames, the hand posture simulated from the video will not have problems such as flickering or shaking, ensuring that the hand posture determined from the video is smooth and continuous, and greatly improving the stability of the hand posture recognition results.
[0201] Furthermore, this application provides a variety of constraint strategies. By processing multiple unoccluded 3D hand joints in the first video frame, the stability of hand pose recognition results can be improved while the computational load can be effectively reduced. By processing multiple 3D hand joints in the first video frame that match any preset hand pose type, the stability of hand pose recognition results can be improved while the computational load can be effectively reduced. In addition, the processed 3D hand joints can be improved to present the effect of actual hand movements.
[0202] Furthermore, this application provides a joint optimization process for the output results of the three-dimensional pose recognition model. By referring to the two-dimensional pose recognition results, the errors caused by the three-dimensional pose recognition model are corrected, thereby further improving the accuracy and stability of hand pose recognition.
[0203] The training method of the hand pose recognition model provided in the embodiments of this application will be introduced next. Figure 9 This is a flowchart illustrating a training method for a hand pose recognition model provided in an embodiment of this application. The method is executed by a computer device, which can be the aforementioned terminal 101 or server 102. (See also...) Figure 9 The method includes the following steps 901 to 905.
[0204] 901. The computer device acquires a sample video of a hand, two-dimensional hand pose information of the sample video, and three-dimensional hand pose information of the sample video.
[0205] The definition of the sample video is given in step 401.
[0206] In some embodiments, the sample video may be a video containing rich annotation information from a dataset related to hand pose recognition. The two-dimensional hand pose information may include the two-dimensional hand joints of each video frame in the sample video; the three-dimensional hand pose information may include the three-dimensional hand joints of each video frame in the sample video.
[0207] In some embodiments, the computer device is a server capable of retrieving a sample video, two-dimensional hand pose information of the sample video, and three-dimensional hand pose information of the sample video from a sample database storing multiple sample videos.
[0208] In some embodiments, the computer device is a terminal that can obtain the sample video, the two-dimensional hand pose information of the sample video, and the three-dimensional hand pose information of the sample video from a sample database via a server.
[0209] 902. The computer device processes the first video frame of the sample video based on the two-dimensional pose recognition model included in the initial pose recognition model to determine multiple two-dimensional prediction joints of the first video frame, which are used to describe the two-dimensional pose of the hand in the first video frame.
[0210] This step refers to step 402 above, and will not be repeated here.
[0211] 903. The computer device inputs multiple two-dimensional predicted joints of the first video frame into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joints of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0212] This step refers to steps 403 and 404 above, and will not be repeated here.
[0213] In some embodiments, before performing step 903, the computer device performs hand detection on the first video frame based on the two-dimensional pose recognition model. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined; at least one three-dimensional pose recognition model matching the chiral features of the at least one hand is obtained. This process is similar to steps 2-3 and 2-4 described above, and will not be repeated here.
[0214] 904. The computer device processes the multiple three-dimensional prediction joints of the first video frame based on the constraint loss function of the initial pose recognition model and the multiple three-dimensional prediction joints of the second video frame in the sample video, so that the difference between the multiple three-dimensional prediction joints of the second video frame and the multiple three-dimensional prediction joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0215] This step refers to step 405 above, and will not be repeated here.
[0216] In some embodiments, the parameters of the constraint loss function include constraint weights, which are used to control the magnitude of adjustment of the coordinates of the plurality of 3D prediction joints when processing the plurality of 3D prediction joints of the first video frame.
[0217] In some embodiments, after the computer device completes step 904, it can also optimize the multiple processed 3D predicted joints in a manner similar to step 406 described above, which will not be elaborated here.
[0218] 905. The computer device adjusts the model parameters of the initial pose recognition model based on the two-dimensional hand pose information, the three-dimensional hand reference information, multiple two-dimensional predicted joints of the first video frame and multiple three-dimensional predicted joints of the processed first video frame to obtain a hand pose recognition model. The model parameters include the parameters of the two-dimensional pose recognition model and the parameters of the three-dimensional pose recognition model.
[0219] In some embodiments, the parameters of the two-dimensional pose recognition model include neural network parameters for feature extraction of the first video frame and parameters involved in determining the hand detection box in the first video frame.
[0220] In some embodiments, the parameters of the three-dimensional pose recognition model include the parameters of the parameter prediction model in step 403 and the parameters of the hand skeleton model in step 404.
[0221] In some embodiments, step 905 can be implemented through the following steps one to three. Step one can be executed after step 902 is completed, and step two can be executed after step 904 is completed.
[0222] Step 1: The computer device determines a first loss value based on the two-dimensional hand pose information and multiple two-dimensional hand joints in the first video frame. The first loss value indicates the error of the two-dimensional pose recognition model in predicting the two-dimensional pose of the hand in the first video frame.
[0223] Step 2: Based on the three-dimensional hand pose information and multiple three-dimensional hand joints in the first video frame, the computer device determines a second loss value. The second loss value indicates the error of the three-dimensional pose recognition model in predicting the three-dimensional pose of the hand in the first video frame.
[0224] Step 3: The computer device adjusts the model parameters of the initial posture recognition model based on the first loss value and the second loss value, so that the first loss value and the second loss value determined based on the adjusted model parameters are reduced, thereby obtaining the hand posture recognition model.
[0225] In some embodiments, the computer device first adjusts the parameters of the two-dimensional pose recognition model based on the first loss value. If the parameters of the two-dimensional pose recognition model cause the first loss value to satisfy an iteration stopping condition, then, based on the adjusted parameters of the two-dimensional pose recognition model and the second loss value, the device adjusts the parameters of the three-dimensional pose recognition model to ensure that the second loss value satisfies an iteration stopping condition. The iteration stopping condition may be that the loss value is less than a specified threshold, or that the number of iterations reaches a target number.
[0226] In some embodiments, the computer device trains three-dimensional pose recognition models with different chiral feature matching, for example, training three-dimensional pose recognition models corresponding to the left hand and the right hand respectively, so as to improve the relevance and accuracy of the three-dimensional pose recognition models.
[0227] The hand pose recognition model obtained through the above technical solution takes into account the correlation of hand movements between video frames when performing hand pose recognition on video. By constraining the differences between the three-dimensional hand joints of the video frames, the hand pose simulated on the video will not have problems such as flickering or shaking, thus ensuring that the hand pose determined on the video is smooth and continuous, and greatly improving the stability of the hand pose recognition results.
[0228] Furthermore, this application provides a variety of constraint strategies. By processing multiple unoccluded 3D hand joints in the first video frame, the stability of hand pose recognition results can be improved while the computational load can be effectively reduced. By processing multiple 3D hand joints in the first video frame that match any preset hand pose type, the stability of hand pose recognition results can be improved while the computational load can be effectively reduced. In addition, the processed 3D hand joints can be improved to present the effect of actual hand movements.
[0229] Furthermore, this application provides a joint optimization process for the output results of the three-dimensional pose recognition model. By referring to the two-dimensional pose recognition results, the errors caused by the three-dimensional pose recognition model are corrected, thereby further improving the accuracy and stability of hand pose recognition.
[0230] Figure 10 This is a schematic diagram of the structure of a hand gesture recognition device provided in an embodiment of this application. See also... Figure 10 The device includes:
[0231] The two-dimensional recognition module 1001 is used to process the first video frame of the video based on the two-dimensional pose recognition model to determine multiple two-dimensional hand joints of the first video frame, which are used to describe the two-dimensional pose of the hand in the first video frame.
[0232] The 3D recognition module 1002 is used to determine multiple 3D hand joints of the first video frame based on multiple 2D hand joints and a 3D pose recognition model of the first video frame, and the 3D pose recognition model is matched with the chiral features of the hand in the first video frame.
[0233] The constraint module 1003 is used to process multiple three-dimensional hand joints of the first video frame based on the constraint loss function and multiple three-dimensional hand joints of the second video frame in the video, so that the difference between the multiple three-dimensional hand joints of the second video frame and the multiple three-dimensional hand joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0234] In one possible implementation, the constraint module 1003 includes:
[0235] The determining unit is used to determine, from the multiple three-dimensional hand joints of the first video frame, the first joint that is not occluded by other three-dimensional hand joints;
[0236] The adjustment unit is used to adjust the coordinates of the multiple first joints based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0237] In one possible implementation, the adjustment unit is used for:
[0238] Based on the multiple second joints and the multiple first joints, determine the range of change between the second joints and the first joints;
[0239] If the change is less than the first threshold, the coordinates of the multiple first joints are adjusted based on the multiple second joints corresponding to the first joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
[0240] In one possible implementation, the adjustment unit is used for:
[0241] Based on the distance between the multiple second joints and the multiple first joints and the constraint weights of the constraint loss function, the loss value of the constraint loss function is determined. The constraint weights are used to control the magnitude of the adjustment of the coordinates of the multiple first joints.
[0242] Based on this loss value, the coordinates of the multiple first joints are adjusted to reduce the difference between the coordinates of the first joint and the coordinates of the second joint.
[0243] In one possible implementation, the constraint module 1003 is used for:
[0244] Based on multiple three-dimensional hand joints in the first video frame, determine the hand pose type that matches the hand in the first video frame;
[0245] From the multiple three-dimensional hand joints of the first video frame, determine multiple third joints that are related to the hand pose type;
[0246] Based on the multiple third joints corresponding to the third joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, the coordinates of the multiple third joints are adjusted so that the difference between the coordinates of the third joint and the coordinates of the fourth joint is reduced.
[0247] In one possible implementation, the three-dimensional recognition module 1002 is used for:
[0248] The pose estimation parameters of the first video frame are obtained by inputting multiple two-dimensional hand joints of the first video frame into the parameter prediction model of the three-dimensional pose recognition model.
[0249] The pose estimation parameters are input into the hand skeleton model of the 3D pose recognition model to obtain multiple 3D hand joints of the first video frame. These multiple 3D hand joints of the first video frame are used to describe the 3D pose of the hand in the first video frame.
[0250] In one possible implementation, the pose estimation parameters include camera parameters and joint rotation parameters. The camera parameters indicate the three-dimensional space in which the hand is located in the first video frame, and the joint rotation parameters indicate how the hand skeleton model determines the deformation of the three-dimensional pose of the hand in the first video frame.
[0251] The device also includes an optimization module, which is used for:
[0252] Based on the multiple three-dimensional hand joints, the camera parameters, and the first video frame, multiple two-dimensional projection joints are determined;
[0253] Based on the multiple two-dimensional projection joints and the multiple two-dimensional hand joints, a projection loss value is determined, which indicates the error between the two-dimensional projection joints and the multiple two-dimensional hand joints;
[0254] Based on the projection loss value, adjust the camera parameters and the joint rotation parameters.
[0255] In one possible implementation, the optimization module is used to:
[0256] Based on the location information of the hand skeleton model, the multiple three-dimensional hand joints of the first video frame are divided into multiple groups of joints with a processing order.
[0257] If the projection loss value determined based on the adjusted camera parameters meets the optimization conditions, then according to the processing order among the multiple sets of joints, the joint rotation parameters corresponding to the multiple sets of joints are adjusted sequentially based on the adjusted camera parameters, the multiple sets of joints, and the first video frame, so as to reduce the projection loss value.
[0258] In one possible implementation, the device further includes a model acquisition module, which is configured to:
[0259] Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined.
[0260] Obtain at least one three-dimensional pose recognition model that matches the chiral features of at least one hand.
[0261] By using the above technical solution, when performing hand posture recognition on video, the correlation between hand movements between video frames is taken into account. By constraining the differences between the three-dimensional hand joints of the video frames, the hand posture simulated from the video will not have problems such as flickering or shaking, ensuring that the hand posture determined from the video is smooth and continuous, and greatly improving the stability of the hand posture recognition results.
[0262] It should be noted that the hand gesture recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the corresponding steps. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the hand gesture recognition device and the hand gesture recognition method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0263] Figure 11This is a schematic diagram of the structure of a training device for a hand pose recognition model provided in an embodiment of this application. See also... Figure 11 The device includes:
[0264] The acquisition module 1101 is used to acquire a sample video of the hand, two-dimensional hand pose information of the sample video, and three-dimensional hand pose information of the sample video.
[0265] The two-dimensional prediction module 1102 is used to process the first video frame of the sample video based on the two-dimensional pose recognition model included in the initial pose recognition model, so as to determine multiple two-dimensional prediction joints of the first video frame, which are used to describe the two-dimensional pose of the hand in the first video frame.
[0266] The three-dimensional prediction module 1103 is used to input multiple two-dimensional predicted joints of the first video frame into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joints of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame.
[0267] The constraint module 1104 is used to process the multiple three-dimensional prediction joints of the first video frame based on the constraint loss function of the initial pose recognition model and the multiple three-dimensional prediction joints of the second video frame in the sample video, so that the difference between the multiple three-dimensional prediction joints of the second video frame and the multiple three-dimensional prediction joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
[0268] The adjustment module 1105 is used to adjust the model parameters of the initial pose recognition model based on the two-dimensional hand pose information, the three-dimensional hand reference information, multiple two-dimensional prediction joints of the first video frame and multiple three-dimensional prediction joints of the processed first video frame, to obtain a hand pose recognition model. The model parameters include the parameters of the two-dimensional pose recognition model and the parameters of the three-dimensional pose recognition model.
[0269] In one possible implementation, the parameters of the constraint loss function include constraint weights, which are used to control the magnitude of adjustment of the coordinates of the plurality of 3D prediction joints when processing the plurality of 3D prediction joints of the first video frame.
[0270] In one possible implementation, the adjustment module 1105 is used for:
[0271] Based on the two-dimensional hand pose information and multiple two-dimensional hand joints of the first video frame, a first loss value is determined. The first loss value indicates the error of the two-dimensional pose recognition model in predicting the two-dimensional pose of the hand in the first video frame.
[0272] Based on the three-dimensional hand pose information and multiple three-dimensional hand joints in the first video frame, a second loss value is determined. The second loss value indicates the error of the three-dimensional pose recognition model in predicting the three-dimensional pose of the hand in the first video frame.
[0273] Based on the first loss value and the second loss value, the model parameters of the initial pose recognition model are adjusted so that the first loss value and the second loss value determined based on the adjusted model parameters are reduced, thereby obtaining the hand pose recognition model.
[0274] In one possible implementation, the device further includes a model acquisition module, which is used for:
[0275] Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined.
[0276] Obtain at least one three-dimensional pose recognition model that matches the chiral features of at least one hand.
[0277] The hand pose recognition model obtained through the above technical solution takes into account the correlation of hand movements between video frames when performing hand pose recognition on video. By constraining the differences between the three-dimensional hand joints of the video frames, the hand pose simulated on the video will not have problems such as flickering or shaking, thus ensuring that the hand pose determined on the video is smooth and continuous, and greatly improving the stability of the hand pose recognition results.
[0278] It should be noted that the hand pose recognition model training device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the corresponding steps. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the hand pose recognition model training device and the hand pose recognition model training method embodiment provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiment, which will not be repeated here.
[0279] This application provides a computer device including a processor and a memory. The memory is used to store at least one computer program, which is loaded and executed by the processor to implement the above-described hand gesture recognition method or hand gesture recognition model training method.
[0280] Taking computer devices as terminals as an example, Figure 12This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal can be: a personal computer (PC), mobile phone, smartphone, personal digital assistant (PDA), wearable device, pocket PC (PPC), tablet computer, smart car system, smart TV, smart speaker, smart voice interaction device, smart home appliance, in-vehicle terminal, etc. The terminal may also be referred to as user equipment, user terminal, portable terminal, laptop terminal, desktop terminal, or other names.
[0281] Typically, a terminal includes a processor 1201 and a memory 1202.
[0282] Processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1201 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as a central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1201 may also include an AI processor, which is used to handle computational operations related to machine learning.
[0283] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 are used to store at least one instruction, which is executed by the processor 1201 to enable the terminal to implement the hand gesture recognition method or the hand gesture recognition model training method provided in the method embodiments of this application.
[0284] In some embodiments, the terminal may also optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, memory 1202, and peripheral device interface 1203 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1203 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, and a power supply 1208.
[0285] Peripheral interface 1203 can be used to connect at least one input / output (I / O) related peripheral device to processor 1201 and memory 1202. In some embodiments, processor 1201, memory 1202 and peripheral interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1201, memory 1202 and peripheral interface 1203 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0286] The radio frequency (RF) circuit 1204 is used to receive and transmit radio frequency (RF) signals, also known as electromagnetic signals. The RF circuit 1204 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1204 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1204 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or Wireless Fidelity (WiFi) networks. In some embodiments, the RF circuit 1204 may also include circuitry related to Near Field Communication (NFC), which is not limited in this application.
[0287] Display screen 1205 is used to display a user interface (UI). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1205 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1201 for processing. In this case, display screen 1205 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1205, disposed on the front panel of the terminal; in other embodiments, there may be at least two display screens, disposed on different surfaces of the terminal or in a folded design; in still other embodiments, display screen 1205 may be a flexible display screen, disposed on a curved or folded surface of the terminal. Furthermore, display screen 1205 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 1205 may be made of materials such as Liquid Crystal Display (LCD) or Organic Light-Emitting Diode (OLED).
[0288] The camera assembly 1206 is used to acquire images or videos. Optionally, the camera assembly 1206 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, virtual reality (VR) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1206 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0289] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to achieve voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1207 may also include a headphone jack.
[0290] The power supply 1208 is used to power the various components in the terminal. The power supply 1208 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1208 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0291] In some embodiments, the terminal further includes one or more sensors 1209. The one or more sensors 1209 include, but are not limited to: an acceleration sensor 1210, a gyroscope sensor 1211, a pressure sensor 1212, an optical sensor 1213, and a proximity sensor 1214.
[0292] Accelerometer 1210 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established by the terminal. For example, accelerometer 1210 can be used to detect the components of gravitational acceleration on the three coordinate axes. Processor 1201 can control display screen 1205 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1210.
[0293] The gyroscope sensor 1211 can detect the terminal's orientation and rotation angle. The gyroscope sensor 1211 can work in conjunction with the accelerometer sensor 1210 to collect the user's 3D movements on the terminal. Based on the data collected by the gyroscope sensor 1211, the processor 1201 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0294] The pressure sensor 1212 can be disposed on the side bezel of the terminal and / or the lower layer of the display screen 1205. When the pressure sensor 1212 is disposed on the side bezel of the terminal, it can detect the user's grip signal on the terminal, and the processor 1201 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1212. When the pressure sensor 1212 is disposed on the lower layer of the display screen 1205, the processor 1201 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1205. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0295] Optical sensor 1213 is used to collect ambient light intensity. In one embodiment, processor 1201 can control the display brightness of display screen 1205 based on the ambient light intensity collected by optical sensor 1213. Specifically, when the ambient light intensity is high, the display brightness of display screen 1205 is increased; when the ambient light intensity is low, the display brightness of display screen 1205 is decreased. In another embodiment, processor 1201 can also dynamically adjust the shooting parameters of camera assembly 1206 based on the ambient light intensity collected by optical sensor 1213.
[0296] The proximity sensor 1214, also known as a distance sensor, is typically installed on the front panel of the terminal. The proximity sensor 1214 is used to detect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1214 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1201 controls the display screen 1205 to switch from a screen-on state to a screen-off state; when the proximity sensor 1214 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1201 controls the display screen 1205 to switch from a screen-off state to a screen-on state.
[0297] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0298] Taking computer equipment as a server as an example, Figure 13This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1300 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1301 and one or more memories 1302. The one or more memories 1302 store at least one computer program, which is loaded and executed by the one or more processors 1301 to implement the aforementioned hand gesture recognition method or hand gesture recognition model training method. Of course, the server 1300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1300 may also include other components for implementing device functions, which will not be elaborated here.
[0299] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program that can be executed by a processor to perform the hand gesture recognition method or the hand gesture recognition model training method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.
[0300] In an exemplary embodiment, a computer program product or computer program is also provided, which includes program code stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the above-described hand gesture recognition method or hand gesture recognition model training method.
[0301] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0302] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A hand gesture recognition method, characterized in that, The method includes: Based on a two-dimensional pose recognition model, the first video frame is processed to determine multiple two-dimensional hand joints in the first video frame. These multiple two-dimensional hand joints are used to describe the two-dimensional pose of the hand in the first video frame. Based on multiple two-dimensional hand joints and a three-dimensional pose recognition model of the first video frame, multiple three-dimensional hand joints of the first video frame are determined, and the three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame. Based on the constraint loss function and multiple three-dimensional hand joints in the second video frame of the video, multiple three-dimensional hand joints in the first video frame are processed so that the difference between the multiple three-dimensional hand joints in the second video frame and the multiple three-dimensional hand joints in the first video frame satisfies the target condition. The second video frame is a video frame preceding the first video frame.
2. The method according to claim 1, characterized in that, The step of processing multiple three-dimensional hand joints in the first video frame based on a constraint loss function and multiple three-dimensional hand joints in the second video frame, so that the difference between the multiple three-dimensional hand joints in the second video frame and the multiple three-dimensional hand joints in the first video frame satisfies the target condition, includes: From the multiple three-dimensional hand joints in the first video frame, determine a number of first joints that are not occluded by other three-dimensional hand joints; Based on the multiple second joints corresponding to the multiple first joints in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, the coordinates of the multiple first joints are adjusted.
3. The method according to claim 2, characterized in that, The adjustment of the coordinates of the plurality of first joints corresponding to the plurality of second joints in the plurality of three-dimensional hand joints based on the second video frame and the constraint loss function includes: Based on the plurality of second joint points and the plurality of first joint points, determine the range of change between the second joint points and the first joint points; If the change is less than a first threshold, the coordinates of the multiple first joints are adjusted based on the multiple second joints corresponding to the multiple first joints in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
4. The method according to claim 2 or 3, characterized in that, The adjustment of the coordinates of the plurality of first joints corresponding to the plurality of second joints in the plurality of three-dimensional hand joints based on the second video frame and the constraint loss function includes: Based on the distance between the plurality of second joints and the plurality of first joints and the constraint weights of the constraint loss function, the loss value of the constraint loss function is determined, and the constraint weights are used to control the magnitude of adjustment of the coordinates of the plurality of first joints. Based on the loss value, the coordinates of the plurality of first joints are adjusted to reduce the difference between the coordinates of the first joints and the coordinates of the second joints.
5. The method according to claim 1, characterized in that, The step of processing multiple three-dimensional hand joints in the first video frame based on a constraint loss function and multiple three-dimensional hand joints in the second video frame, so that the difference between the multiple three-dimensional hand joints in the second video frame and the multiple three-dimensional hand joints in the first video frame satisfies the target condition, includes: Based on multiple three-dimensional hand joints in the first video frame, determine the hand pose type that matches the hand in the first video frame. From the multiple three-dimensional hand joints of the first video frame, determine multiple third joints related to the hand posture type; Based on the multiple fourth joints corresponding to the third joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, the coordinates of the multiple third joints are adjusted so that the difference between the coordinates of the third joint and the coordinates of the fourth joint is reduced.
6. The method according to claim 1, characterized in that, The determination of multiple three-dimensional hand joints in the first video frame based on multiple two-dimensional hand joints and a three-dimensional pose recognition model includes: Multiple two-dimensional hand joints of the first video frame are input into the parameter prediction model of the three-dimensional pose recognition model to obtain the pose estimation parameters of the first video frame. The pose estimation parameters are input into the hand skeleton model of the three-dimensional pose recognition model to obtain multiple three-dimensional hand joints of the first video frame. The multiple three-dimensional hand joints of the first video frame are used to describe the three-dimensional pose of the hand in the first video frame.
7. The method according to claim 6, characterized in that, The pose estimation parameters include camera parameters and joint rotation parameters. The camera parameters indicate the three-dimensional space in which the hand is located in the first video frame, and the joint rotation parameters indicate the deformation method by which the hand skeleton model determines the three-dimensional pose of the hand in the first video frame. After determining the multiple three-dimensional hand joints of the first video frame based on the multiple two-dimensional hand joints and the three-dimensional pose recognition model of the first video frame, the method further includes: Based on the multiple three-dimensional hand joints, the camera parameters, and the first video frame, multiple two-dimensional projection joints are determined; Based on the plurality of two-dimensional projection joints and the plurality of two-dimensional hand joints, a projection loss value is determined, wherein the projection loss value indicates the error between the two-dimensional projection joints and the plurality of two-dimensional hand joints; Based on the projection loss value, adjust the camera parameters and the joint rotation parameters.
8. The method according to claim 7, characterized in that, The step of adjusting the camera parameters and the joint rotation parameters based on the projection loss value includes: Based on the location information of the hand skeleton model, the multiple three-dimensional hand joints of the first video frame are divided into multiple groups of joints with a processing order. If the projection loss value determined based on the adjusted camera parameters meets the optimization conditions, the joint rotation parameters corresponding to the multiple sets of joints are adjusted sequentially according to the processing order among the multiple sets of joints, based on the adjusted camera parameters, the multiple sets of joints and the first video frame, so as to reduce the projection loss value.
9. The method according to claim 1, characterized in that, Before determining the multiple three-dimensional hand joints of the first video frame based on the multiple two-dimensional hand joints and the three-dimensional pose recognition model of the first video frame, the method further includes: Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined. Obtain at least one three-dimensional pose recognition model that matches the chiral features of the at least one hand.
10. A training method for a hand pose recognition model, characterized in that, The method includes: Acquire sample videos of the hand, two-dimensional hand pose information of the sample videos, and three-dimensional hand pose information of the sample videos; Based on the two-dimensional pose recognition model included in the initial pose recognition model, the first video frame of the sample video is processed to determine multiple two-dimensional prediction joints of the first video frame. The multiple two-dimensional prediction joints are used to describe the two-dimensional pose of the hand in the first video frame. Multiple two-dimensional predicted joints of the first video frame are input into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joints of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame. Based on the constraint loss function of the initial pose recognition model and multiple 3D prediction joints of the second video frame in the sample video, the multiple 3D prediction joints of the first video frame are processed so that the difference between the multiple 3D prediction joints of the second video frame and the multiple 3D prediction joints of the first video frame satisfies the target condition. The second video frame is a video frame before the first video frame. Based on the two-dimensional hand posture information, the three-dimensional hand posture information, multiple two-dimensional predicted joints of the first video frame, and multiple three-dimensional predicted joints of the processed first video frame, the model parameters of the initial posture recognition model are adjusted to obtain a hand posture recognition model. The model parameters include the parameters of the two-dimensional posture recognition model and the parameters of the three-dimensional posture recognition model.
11. The method according to claim 10, characterized in that, The parameters of the constraint loss function include constraint weights, which are used to control the magnitude of adjustment of the coordinates of the multiple 3D prediction joints when processing the multiple 3D prediction joints of the first video frame.
12. The method according to claim 10, characterized in that, The step of adjusting the model parameters of the initial pose recognition model based on the two-dimensional hand pose information, the three-dimensional hand pose information, multiple two-dimensional predicted joints of the first video frame, and multiple three-dimensional predicted joints of the processed first video frame to obtain the hand pose recognition model includes: Based on the two-dimensional hand pose information and multiple two-dimensional hand joints of the first video frame, a first loss value is determined. The first loss value indicates the error of the two-dimensional pose recognition model in predicting the two-dimensional pose of the hand in the first video frame. Based on the three-dimensional hand pose information and multiple three-dimensional hand joints in the first video frame, a second loss value is determined. The second loss value indicates the error of the three-dimensional pose recognition model in predicting the three-dimensional pose of the hand in the first video frame. Based on the first loss value and the second loss value, the model parameters of the initial pose recognition model are adjusted so that the first loss value and the second loss value determined based on the adjusted model parameters are reduced, thereby obtaining the hand pose recognition model.
13. The method according to claim 10, characterized in that, Before inputting the multiple two-dimensional predicted joints of the first video frame into the three-dimensional pose recognition model included in the initial pose recognition model, the method further includes: Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined. Obtain at least one three-dimensional pose recognition model that matches the chiral features of the at least one hand.
14. A hand gesture recognition device, characterized in that, The device includes: A two-dimensional recognition module is used to process the first video frame of a video based on a two-dimensional pose recognition model to determine multiple two-dimensional hand joints in the first video frame. The multiple two-dimensional hand joints are used to describe the two-dimensional pose of the hand in the first video frame. A 3D recognition module is used to determine multiple 3D hand joints in the first video frame based on multiple 2D hand joints and a 3D pose recognition model, wherein the 3D pose recognition model is matched with the chiral features of the hand in the first video frame. The constraint module is used to process multiple three-dimensional hand joints of the first video frame based on a constraint loss function and multiple three-dimensional hand joints of the second video frame in the video, so that the difference between the multiple three-dimensional hand joints of the second video frame and the multiple three-dimensional hand joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame.
15. The apparatus according to claim 14, characterized in that, The constraint module includes: The determining unit is used to determine, from the multiple three-dimensional hand joints in the first video frame, a plurality of first joints that are not occluded by other three-dimensional hand joints; The adjustment unit is used to adjust the coordinates of the plurality of first joints based on the plurality of second joints corresponding to the plurality of first joints in the plurality of three-dimensional hand joints of the second video frame and the constraint loss function.
16. The apparatus according to claim 15, characterized in that, The adjustment unit is used for: Based on the plurality of second joint points and the plurality of first joint points, determine the range of change between the second joint points and the first joint points; If the change is less than a first threshold, the coordinates of the multiple first joints are adjusted based on the multiple second joints corresponding to the multiple first joints in the multiple three-dimensional hand joints of the second video frame and the constraint loss function.
17. The apparatus according to claim 15 or 16, characterized in that, The adjustment unit is used for: Based on the distance between the plurality of second joints and the plurality of first joints and the constraint weights of the constraint loss function, the loss value of the constraint loss function is determined, and the constraint weights are used to control the magnitude of adjustment of the coordinates of the plurality of first joints. Based on the loss value, the coordinates of the plurality of first joints are adjusted to reduce the difference between the coordinates of the first joints and the coordinates of the second joints.
18. The apparatus according to claim 14, characterized in that, The constraint module is used for: Based on multiple three-dimensional hand joints in the first video frame, determine the hand pose type that matches the hand in the first video frame. From the multiple three-dimensional hand joints of the first video frame, determine multiple third joints related to the hand posture type; Based on the multiple fourth joints corresponding to the third joint in the multiple three-dimensional hand joints of the second video frame and the constraint loss function, the coordinates of the multiple third joints are adjusted so that the difference between the coordinates of the third joint and the coordinates of the fourth joint is reduced.
19. The apparatus according to claim 14, characterized in that, The three-dimensional recognition module is used for: Multiple two-dimensional hand joints of the first video frame are input into the parameter prediction model of the three-dimensional pose recognition model to obtain the pose estimation parameters of the first video frame. The pose estimation parameters are input into the hand skeleton model of the three-dimensional pose recognition model to obtain multiple three-dimensional hand joints of the first video frame. The multiple three-dimensional hand joints of the first video frame are used to describe the three-dimensional pose of the hand in the first video frame.
20. The apparatus according to claim 19, characterized in that, The pose estimation parameters include camera parameters and joint rotation parameters. The camera parameters indicate the three-dimensional space in which the hand is located in the first video frame, and the joint rotation parameters indicate the deformation method by which the hand skeleton model determines the three-dimensional pose of the hand in the first video frame. The device also includes an optimization module, which is used to: Based on the multiple three-dimensional hand joints, the camera parameters, and the first video frame, multiple two-dimensional projection joints are determined; Based on the plurality of two-dimensional projection joints and the plurality of two-dimensional hand joints, a projection loss value is determined, wherein the projection loss value indicates the error between the two-dimensional projection joints and the plurality of two-dimensional hand joints; Based on the projection loss value, adjust the camera parameters and the joint rotation parameters.
21. The apparatus according to claim 20, characterized in that, The optimization module is used for: Based on the location information of the hand skeleton model, the multiple three-dimensional hand joints of the first video frame are divided into multiple groups of joints with a processing order. If the projection loss value determined based on the adjusted camera parameters meets the optimization conditions, the joint rotation parameters corresponding to the multiple sets of joints are adjusted sequentially according to the processing order among the multiple sets of joints, based on the adjusted camera parameters, the multiple sets of joints and the first video frame, so as to reduce the projection loss value.
22. The apparatus according to claim 14, characterized in that, The device further includes a model acquisition module, the model acquisition module being used for: Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined. Obtain at least one three-dimensional pose recognition model that matches the chiral features of the at least one hand.
23. A training device for a hand pose recognition model, characterized in that, The device includes: The acquisition module is used to acquire a sample video of the hand, two-dimensional hand pose information of the sample video, and three-dimensional hand pose information of the sample video. The two-dimensional prediction module is used to process the first video frame of the sample video based on the two-dimensional pose recognition model included in the initial pose recognition model, so as to determine multiple two-dimensional prediction joints of the first video frame, the multiple two-dimensional prediction joints being used to describe the two-dimensional pose of the hand in the first video frame. The three-dimensional prediction module is used to input multiple two-dimensional predicted joint points of the first video frame into the three-dimensional pose recognition model included in the initial pose recognition model to obtain multiple three-dimensional predicted joint points of the first video frame. The three-dimensional pose recognition model is matched with the chiral features of the hand in the first video frame. The constraint module is used to process multiple three-dimensional prediction joints of the first video frame based on the constraint loss function of the initial pose recognition model and multiple three-dimensional prediction joints of the second video frame in the sample video, so that the difference between the multiple three-dimensional prediction joints of the second video frame and the multiple three-dimensional prediction joints of the first video frame satisfies the target condition, wherein the second video frame is a video frame preceding the first video frame. The adjustment module is used to adjust the model parameters of the initial pose recognition model based on the two-dimensional hand pose information, the three-dimensional hand pose information, multiple two-dimensional prediction joints of the first video frame, and multiple three-dimensional prediction joints of the processed first video frame, to obtain a hand pose recognition model. The model parameters include the parameters of the two-dimensional pose recognition model and the parameters of the three-dimensional pose recognition model.
24. The apparatus according to claim 23, characterized in that, The parameters of the constraint loss function include constraint weights, which are used to control the magnitude of adjustment of the coordinates of the multiple 3D prediction joints when processing the multiple 3D prediction joints of the first video frame.
25. The apparatus according to claim 23, characterized in that, The adjustment module is used for: Based on the two-dimensional hand pose information and multiple two-dimensional hand joints of the first video frame, a first loss value is determined. The first loss value indicates the error of the two-dimensional pose recognition model in predicting the two-dimensional pose of the hand in the first video frame. Based on the three-dimensional hand pose information and multiple three-dimensional hand joints in the first video frame, a second loss value is determined. The second loss value indicates the error of the three-dimensional pose recognition model in predicting the three-dimensional pose of the hand in the first video frame. Based on the first loss value and the second loss value, the model parameters of the initial pose recognition model are adjusted so that the first loss value and the second loss value determined based on the adjusted model parameters are reduced, thereby obtaining the hand pose recognition model.
26. The apparatus according to claim 23, characterized in that, The device further includes a model acquisition module, which is used for: Based on the two-dimensional pose recognition model, hand detection is performed on the first video frame. If at least one hand is detected in the first video frame, the chiral features of the at least one hand are determined. Obtain at least one three-dimensional pose recognition model that matches the chiral features of the at least one hand.
27. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the hand gesture recognition method as described in any one of claims 1 to 9, or the training method for the hand gesture recognition model as described in any one of claims 10 to 13.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the hand gesture recognition method as described in any one of claims 1 to 9, or the training method for the hand gesture recognition model as described in any one of claims 10 to 13.
29. A computer program product, characterized in that, The computer program product includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the hand gesture recognition method as described in any one of claims 1 to 9, or the training method for a hand gesture recognition model as described in any one of claims 10 to 13.
Citation Information
Patent Citations
Cross-view character recognition method based on shapes and postures under wearable equipment
CN111582036A
Motion capture and redirection method
CN113989928A