A monocular RGB two - hand pose real - time detection method and system

By adopting detection network and manual linear skin model in monocular RGB cameras, combined with projection priors, high-precision real-time detection of dual-hand posture under monocular RGB is achieved, solving the problems of high cost and low accuracy in the existing technology, and achieving efficient and economical dual-hand posture estimation.

CN117953579BActive Publication Date: 2025-05-27BEIJING WEILAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410105762.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-05-27
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect the posture of both hands in real time through a monocular RGB camera, especially when the hand structure is complex and the joint relationship is changing, and traditional methods require multi-view or RGBD cameras, which increases costs.

Method used

The human body image was collected by a monocular RGB camera, and the standard format image was obtained through preprocessing. The pre-configured detection network was used to extract the coordinates of the human body detection frame and the whole body two-dimensional key point coordinates, calculate the two-hand bounding box and input the two-hand posture regression network, and obtain the set of human hand parameters, including confidence, shape posture parameters and projection parameters. Combined with the human hand linear skin model and projection prior, a tracking sequence was constructed to reduce the jitter of the human hand skeleton.

Benefits of technology

High-precision two-hand posture estimation under monocular RGB conditions is achieved, which improves the estimation accuracy of two-hand interaction, hand self-blocking and hand object interaction, avoids the dependence of additional auxiliary information, and reduces the cost and complexity of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117953579B_ABST
    Figure CN117953579B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of computer vision, and discloses a monocular RGB two-handed pose real-time detection method and system. The detection method includes: collecting a human body image through a monocular RGB camera, and obtaining a standard format image through preprocessing of the human body image; sequentially extracting a human body detection frame and full-body two-dimensional key point coordinates in the standard format image based on a pre-configured detection network; calculating a bounding box of both hands based on the full-body two-dimensional key point coordinates, and inputting the bounding box into a two-handed pose regression network to obtain a set of human hand parameters; obtaining a three-dimensional skeleton, a mesh model and two-dimensional positions of the human hands based on the set of human hand parameters, and constructing a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton. The present invention combines deep learning technology to complete the estimation of two-handed poses, introduces a linear skinning model of human hands and a projection prior, and improves the estimation accuracy of two-handed interaction, self-occlusion of hands, and hand-object interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for real-time monocular RGB two-hand pose detection. Background Art

[0002] Real-time hand pose estimation is an important research direction in the field of computer vision. It aims to capture and understand the movement and pose of human hands in real time and accurately by analyzing the hand position and joint angles in images or videos. This technology has a wide range of applications in many fields, including virtual reality, augmented reality, gesture recognition, human-computer interaction, medical imaging, etc.

[0003] To solve this problem, researchers usually use deep learning technologies, especially deep learning models such as convolutional neural networks (CNNs) and vision transformers (VITs). One of the main challenges behind real-time hand pose estimation is to achieve highly accurate pose speculation because the structure of the hand is complex and the relationship between joints is variable. At the same time, RGB images lack three-dimensional information of the scene, resulting in most of the current monocular RGB viewpoint methods on the market being unable to well regress the hand pose. Therefore, for regression accuracy, many methods either adopt multi-viewpoint inputs or use RGBD cameras with depth, but this will significantly increase the usage cost.

[0004] Therefore, how to provide a method and system for real-time monocular RGB two-hand pose detection is an urgent problem to be solved at present. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for real-time monocular RGB two-hand pose detection to solve the above technical problems existing in the prior art.

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.

[0007] According to a first aspect of an embodiment of the present invention, a method for real-time monocular RGB two-hand pose detection is provided.

[0008] In one embodiment, the method for real-time monocular RGB two-hand pose detection includes:

[0009] Collect a human body image through a monocular RGB camera, and obtain a standard format image through preprocessing the human body image;

[0010] Based on a pre-configured detection network, sequentially extract the human detection box and the full-body two-dimensional key point coordinates in the standard format image;

[0011] Based on the full-body two-dimensional key point coordinates, calculate the bounding boxes of both hands, and input the bounding boxes into the two-hand pose regression network to obtain a set of hand parameters, where the set of hand parameters includes confidence, shape and pose parameters, and projection parameters;

[0012] Based on the set of hand parameters, obtain the three-dimensional skeleton, mesh model and two-dimensional position of the hand, and based on the three-dimensional skeleton and two-dimensional skeleton, construct a tracking sequence to reduce the jitter of the hand skeleton.

[0013] In one embodiment, the monocular RGB two-hand pose real-time detection method further includes:

[0014] Before obtaining the set of hand parameters, based on a visual feature extraction backbone network, extract the implicit feature vector of the hand in the bounding box, and connect a decoder to the implicit feature vector to construct a two-hand pose regression network;

[0015] Obtain hand synthetic data and multi-view data, merge them as training data, and train and optimize the two-hand pose regression network based on the training data.

[0016] In one embodiment, the sequentially extracting the human detection box and the full-body two-dimensional key point coordinates in the standard format image based on the pre-configured detection network includes:

[0017] Input the standard format image into the human position detection network, output multiple candidate target detection boxes, and obtain the human detection box by calculating the non-maximum suppression of the candidate target detection boxes;

[0018] Input the human detection box into the human key point detection network, output the heat maps of each joint of the human body, and obtain the full-body two-dimensional key point coordinates by calculating the non-maximum suppression of the heat maps.

[0019] In one embodiment, the calculating the bounding boxes of both hands based on the full-body two-dimensional key point coordinates includes:

[0020] Based on the full-body two-dimensional key point coordinates, calculate the bounding boxes of the left hand and the right hand respectively. If the two hands overlap, merge the bounding boxes of the two hands into one bounding box;

[0021] Based on the remaining full-body two-dimensional key point coordinates after calculating the bounding boxes of both hands, calculate the face bounding box, the human foot bounding box and the human body bounding box to be used as a reference and assistance for human pose evaluation.

[0022] In one embodiment, obtaining the three-dimensional skeleton, mesh model, and two-dimensional position of the human hand based on the set of human hand parameters, and constructing a tracking sequence based on the three-dimensional skeleton and two-dimensional skeleton to reduce the jitter of the human hand skeleton includes:

[0023] Extract the human hands with confidence greater than the preset confidence threshold and their parameters from the set of human hand parameters;

[0024] Input the shape and pose parameters into the human hand linear skinning model to obtain the three-dimensional skeleton and mesh model of the hand;

[0025] Calculate the two-dimensional position of the human hand in the human body image based on the projection parameters;

[0026] Construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton of the hand joints projected onto the human body image. Determine the initial human hand region and the skeleton of the joints through the first frame, and perform quaternion filtering on each skeleton of the human hand joints in subsequent frames to ensure the stability of the skeleton in the time dimension.

[0027] According to the second aspect of the embodiments of the present invention, a monocular RGB two-handed pose real-time detection system is provided.

[0028] In one embodiment, the monocular RGB two-handed pose real-time detection system includes:

[0029] An image acquisition and processing module for acquiring a human body image through a monocular RGB camera and obtaining a standard format image through preprocessing of the human body image;

[0030] A human body detection and positioning module for sequentially extracting the human body detection frame and the full-body two-dimensional key point coordinates in the standard format image based on a pre-configured detection network;

[0031] A human hand parameter calculation module for calculating the bounding boxes of both hands based on the full-body two-dimensional key point coordinates and inputting the bounding boxes into a two-handed pose regression network to obtain a set of human hand parameters, where the set of human hand parameters includes confidence, shape and pose parameters, and projection parameters;

[0032] A human hand pose detection module for obtaining the three-dimensional skeleton, mesh model, and two-dimensional position of the human hand based on the set of human hand parameters, and constructing a tracking sequence based on the three-dimensional skeleton and two-dimensional skeleton to reduce the jitter of the human hand skeleton.

[0033] In one embodiment, the monocular RGB two-handed pose real-time detection system further includes:

[0034] A model construction module, configured to extract an implicit feature vector of a hand in a bounding box based on a vision feature extraction backbone network before obtaining a hand parameter set, and connect a decoder to the implicit feature vector to construct a two-hand pose regression network;

[0035] A model training module, configured to obtain synthetic hand data and multi-view data, combine them as training data, and train and optimize the two-hand pose regression network based on the training data.

[0036] In one embodiment, the human body detection and positioning module includes: a position detection module and a key point detection module, where

[0037] The position detection module is configured to input the standard format image into a human body position detection network, output a plurality of candidate target detection boxes, and obtain a human body detection box by calculating non-maximum suppression of the candidate target detection boxes;

[0038] The key point detection module is configured to input the human body detection box into a human body key point detection network, output heat maps of each joint of the human body, and obtain the full-body two-dimensional key point coordinates by calculating non-maximum suppression of the heat maps.

[0039] In one embodiment, the hand parameter calculation module includes: a bounding box detection module and a reference assistance module, where

[0040] Based on the full-body two-dimensional key point coordinates, calculate the bounding boxes of the left hand and the right hand respectively. If the two hands overlap, merge the bounding boxes of the two hands into one bounding box;

[0041] Based on the remaining full-body two-dimensional key point coordinates after calculating the two-hand bounding boxes, calculate the face bounding box, the human foot bounding box, and the human body bounding box to be used as a reference and assistance for human body pose evaluation.

[0042] In one embodiment, the hand pose detection module includes: a threshold judgment module, a skeleton mesh module, a two-dimensional positioning module, and a tracking sequence module, where

[0043] The threshold judgment module is configured to extract the hands and their parameters in the hand parameter set whose confidence level is greater than a preset confidence level threshold;

[0044] The skeleton mesh module is configured to input the shape and pose parameters into a hand linear skinning model to obtain a three-dimensional skeleton and mesh model of the hand;

[0045] The two-dimensional positioning module is configured to calculate the two-dimensional position of the hand in the human body image based on projection parameters;

[0046] The tracking sequence module is used to construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton projected from the hand joints onto the human image, determine the initial hand region and the skeleton of the joints through the first frame, and perform quaternion filtering on each skeleton of the hand joints in subsequent frames to ensure the stability of the skeleton in the time dimension.

[0047] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0048] The present invention combines deep learning technology to complete the estimation of the postures of both hands. By introducing the linear skinning model of the human hand and the projection prior, the estimation accuracy of hand interaction, hand self-occlusion, and hand-object interaction is improved; most traditional methods only use the information of hand key points to estimate the hand posture, which performs poorly in the case of hand occlusion such as hand interaction, while the present invention introduces the linear skinning model of the human hand as a hand prior, enabling the neural network to have guidance during the training process, thereby improving the performance under severe occlusion interaction; on the other hand, through sufficient training on high-quality data, the neural network of the present invention obtains a stable projection prior, which enables the neural network of the present invention to have good three-dimensional perception in RGB images lacking three-dimensional information without relying on additional assistance such as depth maps and camera parameters, avoiding many three-dimensional ambiguity problems, and thus can improve the convergence speed of algorithm training and the accuracy of the final result. In a normal indoor environment, the system of the present invention can capture the gestures of both hands of the person in the picture in real time and reconstruct a three-dimensional mesh model.

[0049] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0051] Figure 1 is a flowchart of a monocular RGB real-time two-handed pose detection method shown according to an exemplary embodiment;

[0052] Figure 2 is a system principle block diagram of a monocular RGB real-time two-handed pose detection system shown according to an exemplary embodiment;

[0053] Figure 3 is a flowchart of human body image processing shown according to an exemplary embodiment;

[0054] Figure 4 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following description and the accompanying drawings fully disclose specific embodiments herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. In this document, terms such as "first", "second", etc. are only used to distinguish one element from another, without requiring or implying any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a structure, device or equipment comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the structure, device or equipment comprising the said element. The various embodiments herein are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0056] In this document, terms such as "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing this document and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention. In the description of this document, unless otherwise specified and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0057] In this document, unless otherwise stated, the term "plurality" means two or more.

[0058] In this document, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.

[0059] In this document, the term "and / or" is an associative relationship describing an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.

[0060] It should be understood that although the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0061] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0062] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0063] Figure 1 An embodiment of a monocular RGB two-handed gesture real-time detection method of the present invention is shown.

[0064] In this alternative embodiment, the monocular RGB two-handed gesture real-time detection method includes:

[0065] Step S101: Collect a human body image through a monocular RGB camera, and obtain a standard format image through preprocessing of the human body image;

[0066] Step S103: Based on a pre-configured detection network, sequentially extract the human body detection box and the full-body two-dimensional key point coordinates in the standard format image;

[0067] Step S105: Based on the full-body two-dimensional key point coordinates, calculate the bounding boxes of the two hands, and input the bounding boxes into a two-handed gesture regression network to obtain a set of hand parameters, where the set of hand parameters includes confidence, shape and pose parameters, and projection parameters;

[0068] Step S107: Based on the set of hand parameters, obtain the three-dimensional skeleton, mesh model and two-dimensional position of the hand, and based on the three-dimensional skeleton and the two-dimensional skeleton, construct a tracking sequence to reduce the jitter of the hand skeleton.

[0069] In this alternative embodiment, the monocular RGB two - hand pose real - time detection method further includes: before obtaining the hand parameter set, based on a backbone network for visual feature extraction, extracting an implicit feature vector of the hand in the bounding box, and connecting a decoder to the implicit feature vector to construct a two - hand pose regression network; obtaining hand synthetic data and multi - view data, combining them as training data, and training and optimizing the two - hand pose regression network based on the training data.

[0070] In this alternative embodiment, when successively extracting the human detection bounding box and the full - body two - dimensional key point coordinates in the standard - format image based on a pre - configured detection network, the standard - format image can be input into a human position detection network to output multiple candidate target detection bounding boxes. By calculating the non - maximum suppression of the candidate target detection bounding boxes, the human detection bounding box is obtained; the human detection bounding box is input into a human key point detection network to output the heat maps of each joint of the human body. By calculating the non - maximum suppression of the heat maps, the full - body two - dimensional key point coordinates are obtained.

[0071] In this alternative embodiment, when calculating the bounding boxes of the two hands based on the full - body two - dimensional key point coordinates, the bounding boxes of the left hand and the right hand can be calculated respectively based on the full - body two - dimensional key point coordinates. If the two hands overlap, the bounding boxes of the two hands are merged into one bounding box; based on the remaining full - body two - dimensional key point coordinates after calculating the bounding boxes of the two hands, the face bounding box, the human foot bounding box, and the human body bounding box are calculated as references and aids for human pose evaluation.

[0072] In this alternative embodiment, when obtaining the three - dimensional skeleton, mesh model, and two - dimensional position of the human hand based on the hand parameter set, and constructing a tracking sequence based on the three - dimensional skeleton and two - dimensional skeleton to reduce the jitter of the human hand skeleton, the human hand and its parameters with a confidence greater than a pre - set confidence threshold in the hand parameter set are extracted; the shape - pose parameters are input into a linear skinning model of the human hand to obtain the three - dimensional skeleton and mesh model of the hand; based on the projection parameters, the two - dimensional position of the human hand in the human body image is calculated; based on the three - dimensional skeleton and the two - dimensional skeleton of the hand joints projected onto the human body image, a tracking sequence is constructed. The initial human hand region and the skeleton of the joints are determined through the first frame, and quaternion filtering is performed on each skeleton of the hand joints in the subsequent frames to ensure the stability of the skeleton in the time dimension.

[0073] Figure 2 An embodiment of a monocular RGB two - hand pose real - time detection system of the present invention is shown.

[0074] In this alternative embodiment, the monocular RGB two - hand pose real - time detection system includes:

[0075] The image acquisition and processing module 201 is used to acquire a human body image through a monocular RGB camera, and obtain a standard format image through preprocessing the human body image;

[0076] The human body detection and positioning module 203 is used to sequentially extract the human body detection box and the full-body two-dimensional key point coordinates in the standard format image based on a pre-configured detection network;

[0077] The human hand parameter calculation module 205 is used to calculate the bounding boxes of both hands based on the full-body two-dimensional key point coordinates, and input the bounding boxes into a two-handed pose regression network to obtain a set of human hand parameters, where the set of human hand parameters includes confidence, shape and pose parameters, and projection parameters;

[0078] The human hand pose detection module 207 is used to obtain the three-dimensional skeleton, mesh model and two-dimensional position of the human hand based on the set of human hand parameters, and construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton to reduce the jitter of the human hand skeleton.

[0079] In this alternative embodiment, the monocular RGB two-handed pose real-time detection system further includes:

[0080] A model construction module (not shown in the figure) is used to extract an implicit feature vector of the hand in the bounding box based on a vision feature extraction backbone network before obtaining the set of human hand parameters, and connect a decoder to the implicit feature vector to construct and form a two-handed pose regression network;

[0081] A model training module (not shown in the figure) is used to obtain hand synthetic data and multi-view data, merge them as training data, and train and optimize the two-handed pose regression network based on the training data.

[0082] In this alternative embodiment, the human body detection and positioning module 203 includes: a position detection module (not shown in the figure) and a key point detection module (not shown in the figure), where the position detection module is used to input the standard format image into a human body position detection network, output multiple candidate target detection boxes, and obtain a human body detection box by calculating the non-maximum suppression of the candidate target detection boxes; the key point detection module is used to input the human body detection box into a human body key point detection network, output a heat map of each joint of the human body, and obtain the full-body two-dimensional key point coordinates by calculating the non-maximum suppression of the heat map.

[0083] In this alternative embodiment, the human hand parameter calculation module 205 includes: a bounding box detection module (not shown in the figure) and a reference assistance module (not shown in the figure). Among them, based on the two-dimensional key point coordinates of the whole body, the bounding boxes of the left hand and the right hand are calculated respectively. If the two hands overlap, the bounding boxes of the two hands are merged into one bounding box; based on the remaining two-dimensional key point coordinates of the whole body after calculating the bounding boxes of the two hands, the face bounding box, the human foot bounding box and the human body bounding box are calculated as a reference and assistance for human body pose evaluation.

[0084] In this alternative embodiment, the human hand pose detection module 207 includes: a threshold judgment module (not shown in the figure), a skeleton mesh module (not shown in the figure), a two-dimensional positioning module (not shown in the figure) and a tracking sequence module (not shown in the figure). Among them, the threshold judgment module is used to extract the human hands and their parameters in the human hand parameter set whose confidence level is greater than the preset confidence threshold; the skeleton mesh module is used to input the shape and pose parameters into the human hand linear skinning model to obtain the three-dimensional skeleton and mesh model of the hand; the two-dimensional positioning module is used to calculate the two-dimensional position of the human hand in the human body image based on the projection parameters; the tracking sequence module is used to construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton of the hand joints projected onto the human body image, determine the initial human hand area and the skeleton of the joints through the first frame, and perform quaternion filtering on each skeleton of the hand joints in the subsequent frames to ensure the stability of the skeleton in the time dimension.

[0085] In specific applications, such as Figure 3 shown, the construction and use of the monocular RGB two-handed pose real-time detection system can be divided into the following six process steps:

[0086] Step 1: Model training and optimization

[0087] Construct a two-handed pose regression network using an advanced visual feature extraction backbone network.

[0088] Perform pre-training using a large amount of hand synthetic data, and add multi-view high-quality data and daily scene images for fine-tuning and optimization.

[0089] Perform multiple rounds of training on the model to learn the human hand prior and projection prior, and improve the inference accuracy of the model in various scenarios.

[0090] Step 2: Device setup.

[0091] In a conventional indoor or outdoor environment, use a consumer-grade monocular RGB camera and fix it in an appropriate position. Ensure that the camera's line of sight is as parallel to the ground as possible.

[0092] Set the area of human activities within the camera's field of view to ensure complete capture of hand poses.

[0093] Connect the camera to a computer with sufficient performance (e.g., above RTX2060), and start the real-time detection system.

[0094] Step 3: Data Acquisition and Preprocessing

[0095] The data captured by the camera is transmitted to the computer.

[0096] Use OpenCV to transcode the captured image, scale it into a three-channel RGB matrix, and perform normalization processing.

[0097] Step 4: Human Detection and Key Point Localization

[0098] Input the normalized RGB matrix into the human body position detection network, which will output multiple candidate object detection boxes, including position, category, and confidence.

[0099] Obtain all human detection boxes in the original input image by calculating non-maximum suppression.

[0100] Resize the main human body box in the image, and then enter a human key point detection network. The network will output a heat map of each joint of the human body in the form of simcc. Similarly, by calculating non-maximum suppression, the two-dimensional key point coordinates of the whole body of the person in the original image can be finally obtained.

[0101] Step 5: Hand Bounding Box Extraction and Two-Hand Pose Regression

[0102] Calculate the bounding boxes of the left hand and the right hand using the two-dimensional joint points of the whole body. If the two hands overlap and interact, merge the bounding boxes of the two hands into a large bounding box; the remaining key points can also calculate the face bounding box, the human foot bounding box, the human body bounding box, etc., which can be used as a reference and assistance for other pose estimations.

[0103] Send the two-hand bounding boxes into the two-hand pose regression network to obtain the confidence, shape and pose parameters, and projection parameters of the hands.

[0104] Step 6: 3D Skeleton Reconstruction and Tracking Sequence Construction

[0105] Substitute the shape and pose parameters into the linear skinning model of the human hand to obtain the 3D skeleton and mesh model of the hand.

[0106] Use the projection parameters to calculate the two-dimensional position of the human hand in the image.

[0107] For the tracking sequence of the two-dimensional skeleton structure obtained by reprojection of the three-dimensional human hand skeleton and hand joints onto the camera image, the initial human hand region and skeleton can be determined in the first frame. In subsequent frames, quaternion filtering is performed on the rotation of each skeleton of the human hand joints based on the previous frame to ensure the stability of the skeleton in the time dimension and reduce the jitter caused by the inference error of the neural network. If necessary, one-euro filtering can also be performed on the three-dimensional and two-dimensional positions of the human hand skeleton nodes to further reduce jitter.

[0108] In addition, existing good open-source works (such as OpenPose, MediaPipe, HRNet, PiFPaf, etc.) can also be used for the human body detection model and human key point model. In the present invention, RTMPose is used and fine-tuned on the collected data set.

[0109] The two-handed pose regression network is a self-developed and self-trained network. In terms of architecture, the present invention adopts an advanced visual feature extraction backbone network in the industry. While ensuring real-time performance, the implicit feature vector of the hand in the image is extracted. Then, the present invention connects multiple decoders after the feature vector to regress the results required by the present invention, such as the confidence of the hand, shape and pose parameters, and projection parameters. In terms of training, the present invention first performs pre-training using a large batch of hand synthetic data, enabling the network to fully learn the human hand prior. Subsequently, the present invention adds a batch of high-quality multi-view data, which contains accurate hand fitting and projection parameters, enabling the network to further learn hand features and obtain projection priors. Since the multi-view high-quality data is expensive and the scene is limited, the present invention then collects a batch of images of daily scenes with slightly lower quality. Combining with the previous model with hand prior and projection prior, the present invention further optimizes the accuracy of the model inference results in difficult situations such as two-handed interaction, hand-object interaction, and motion blur, and eliminates the boundary between laboratory data and conventional scenes, and finally trains a two-handed pose regression network with good results.

[0110] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes the steps in the above method embodiments.

[0111] Those skilled in the art can understand that Figure 4 the structure shown in Figure 4 is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0112] In addition, the present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0113] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0114] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0115] The present invention is not limited to the structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A monocular RGB two-handed posture real-time detection method, characterized in that: include: A human body image is collected by using a monocular RGB camera, and a standard format image is obtained by preprocessing the human body image; Based on a pre-configured detection network, sequentially extracting a human body detection frame and coordinates of two-dimensional key points of the whole body in the standard format image; Based on the coordinates of the two-dimensional key points of the whole body, a bounding box of both hands is calculated, and the bounding box is input into a two-hand posture regression network to obtain a set of human hand parameters, wherein the set of human hand parameters includes confidence, shape posture parameters and projection parameters; Based on the set of human hand parameters, a three-dimensional skeleton, a mesh model and a two-dimensional position of the human hand are obtained, and a tracking sequence is constructed based on the three-dimensional skeleton and the two-dimensional skeleton to reduce the shaking of the human hand skeleton; The step of obtaining a three-dimensional skeleton, a mesh model, and a two-dimensional position of a hand based on the set of hand parameters, and constructing a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton to reduce the shaking of the hand skeleton includes: Extracting the human hands and their parameters whose confidence level is greater than a preset confidence threshold from the human hand parameter set; Input the shape and posture parameters into the linear skin model of the human hand to obtain the three-dimensional skeleton and mesh model of the hand; Based on the projection parameters, calculating the two-dimensional position of the human hand in the human body image; Based on the three-dimensional skeleton and the two-dimensional skeleton projected onto the human body image, a tracking sequence is constructed, the initial hand area and the skeleton of the joints are determined through the first frame, and quaternion filtering is performed on each skeleton of the hand joints in subsequent frames to ensure the stability of the skeleton in the time dimension.

2. The monocular RGB two-handed posture real-time detection method according to claim 1 is characterized in that: Also includes: Before obtaining the set of human hand parameters, based on the visual feature extraction backbone network, the implicit feature vector of the hand in the bounding box is extracted, and a decoder is connected to the implicit feature vector to construct a two-handed posture regression network; The hand synthesis data and the multi-view data are obtained and combined as training data, and based on the training data, the two-hand posture regression network is trained and optimized.

3. The monocular RGB two-handed posture real-time detection method according to claim 1 is characterized in that: The method of sequentially extracting the human body detection frame and the coordinates of the whole body two-dimensional key points in the standard format image based on the pre-configured detection network includes: Inputting the standard format image into a human position detection network, outputting a plurality of candidate target detection frames, and obtaining a human detection frame by calculating non-maximum suppression of the candidate target detection frames; The human body detection frame is input into the human body key point detection network, and the heat map of each joint of the human body is output. The two-dimensional key point coordinates of the whole body are obtained by calculating the non-maximum suppression of the heat map.

4. The monocular RGB two-handed posture real-time detection method according to claim 1 is characterized in that: The calculating the bounding boxes of both hands based on the coordinates of the two-dimensional key points of the whole body includes: Based on the coordinates of the two-dimensional key points of the whole body, respectively calculating the bounding boxes of the left hand and the right hand, and if the two hands overlap, merging the bounding boxes of the two hands into one bounding box; Based on the coordinates of the remaining two-dimensional key points of the whole body calculated by the hand bounding box, the face bounding box, the foot bounding box and the human body bounding box are calculated to serve as a reference and assistance for human posture assessment.

5. A monocular RGB two-handed posture real-time detection system, characterized in that: include: An image acquisition and processing module is used to acquire a human body image through a monocular RGB camera, and obtain a standard format image by preprocessing the human body image; A human body detection and positioning module, used to sequentially extract the human body detection frame and the coordinates of the two-dimensional key points of the whole body in the standard format image based on a pre-configured detection network; A hand parameter calculation module, used to calculate the bounding box of the hands based on the coordinates of the two-dimensional key points of the whole body, and input the bounding box into the two-hand posture regression network to obtain a set of hand parameters, wherein the set of hand parameters includes confidence, shape posture parameters and projection parameters; The hand posture detection module is used to obtain the three-dimensional skeleton, mesh model and two-dimensional position of the hand based on the hand parameter set, and to construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton to reduce the shaking of the hand skeleton.

6. The monocular RGB two-handed posture real-time detection system according to claim 5 is characterized in that: Also includes: A model building module is used to extract the implicit feature vector of the hand in the bounding box based on the visual feature extraction backbone network before obtaining the set of human hand parameters, and connect the decoder to the implicit feature vector to build a two-hand posture regression network; The model training module is used to obtain hand synthesis data and multi-view data, merge them as training data, and train and optimize the two-hand posture regression network based on the training data.

7. The monocular RGB two-handed posture real-time detection system according to claim 5, characterized in that: The human body detection and positioning module includes: a position detection module and a key point detection module, wherein: The position detection module is used to input the standard format image into a human position detection network, output a plurality of candidate target detection frames, and obtain a human detection frame by calculating non-maximum suppression of the candidate target detection frames; The key point detection module is used to input the human body detection frame into the human body key point detection network, output the heat map of each joint of the human body, and obtain the two-dimensional key point coordinates of the whole body by calculating the non-maximum suppression of the heat map.

8. The monocular RGB two-handed posture real-time detection system according to claim 5, characterized in that: The hand parameter calculation module includes: a bounding box detection module and a reference auxiliary module, wherein: Based on the coordinates of the two-dimensional key points of the whole body, respectively calculating the bounding boxes of the left hand and the right hand, and if the two hands overlap, merging the bounding boxes of the two hands into one bounding box; Based on the coordinates of the remaining two-dimensional key points of the whole body calculated by the hand bounding box, the face bounding box, the foot bounding box and the human body bounding box are calculated to serve as a reference and assistance for human posture assessment.

9. The monocular RGB two-handed posture real-time detection system according to claim 5, characterized in that: The hand gesture detection module includes: a threshold judgment module, a skeleton grid module, a two-dimensional positioning module and a tracking sequence module, wherein: The threshold judgment module is used to extract the human hands and their parameters whose confidence level is greater than a preset confidence threshold from the human hand parameter set; The skeleton mesh module is used to input shape and posture parameters into the linear skin model of the human hand to obtain a three-dimensional skeleton and mesh model of the hand; The two-dimensional positioning module is used to calculate the two-dimensional position of the human hand in the human body image based on the projection parameters; The tracking sequence module is used to construct a tracking sequence based on the three-dimensional skeleton and the two-dimensional skeleton projected onto the human body image by the hand joints, determine the initial hand area and the skeleton of the joints through the first frame, and perform quaternion filtering on each skeleton of the hand joints in subsequent frames to ensure the stability of the skeleton in the time dimension.

Citation Information

Patent Citations

  • Color image hand posture estimation method for shielding condition

    CN111027407A

  • Multi-person three-dimensional attitude estimation method and device and electronic equipment

    CN114550282A

  • Interactive two-hand three-dimensional reconstruction method and system based on single RGB image

    CN117333635A