Multi-view high-precision hand three-dimensional reconstruction method and system fusing two-dimensional completion and implicit optimization
The method addresses hand occlusion and motion blur in hand-based interaction systems by using VAE-GAN and implicit optimization with MANO parameterization and SeSDF for high-precision hand reconstruction, enhancing accuracy and reducing costs.
Patent Information
- Application Number
- CN202510404211.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
In the multi-view RGB data acquisition system, hand occlusion, motion blur and lighting inconsistent results in low single-view reconstruction accuracy, large error in multi-view triangle measurement, and high cost of obtaining multi-modal data, making it difficult to achieve high-precision three-dimensional reconstruction of high-precision hand.
The complete network, which combines the variational autoencoder and the generative adversarial network, completes the obstructed area at the two-dimensional image level, combines multi-view fusion and implicit representation technology, regresses the three-dimensional bone key points through deep learning networks, optimizes the hand surface with parametric hand models and implicit functions, and reconstructs a high-precision hand model.
In complex gestures and strong occlusion scenarios, high-precision hand reconstruction is achieved, which reduces system deployment costs, avoids dependence on high-precision calibration and multi-modal sensors, and improves occlusion robustness and reconstruction accuracy.
Smart Images

Figure CN120318290A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and three-dimensional reconstruction, and particularly to a multi-view high-precision hand three-dimensional reconstruction method and system that integrates two-dimensional completion and implicit optimization. Background Art
[0002] In the application of hand-object interaction, obtaining a three-dimensional model of the hand is of great significance for gesture recognition, pose tracking, and mechanical analysis. However, in a multi-view RGB data acquisition system, since the hand in the picture is often partially blocked by the object being held, and problems such as motion blur, sudden illumination change, and depth ambiguity may occur under the shooting conditions, the existing technologies still have the following deficiencies:
[0003] 1. Limited accuracy of single-view reconstruction
[0004] When there is significant occlusion of the hand or complex gestures, single-view two-dimensional detection usually easily produces phenomena such as key-point drift and shape loss, resulting in poor subsequent bone key-point prediction and three-dimensional reconstruction accuracy. At the same time, view motion blur and inconsistent illumination will further increase the uncertainty of single-view reconstruction.
[0005] 2. Amplification of multi-view triangulation error
[0006] Although multi-view methods can alleviate the defects of single-view to a certain extent, traditional triangulation methods are too sensitive to 2D detection results. If there is a deviation in single-view detection, this error will be further amplified in the three-dimensional space through triangulation. In addition, multi-view triangulation depends on the accuracy of camera calibration, such as focal length and distortion parameters. Once the calibration is inaccurate or the actual shooting environment is complex, the three-dimensional reconstruction accuracy will be greatly affected.
[0007] 3. High cost of multi-modal data acquisition
[0008] Some solutions propose to add multi-modal data such as depth maps and infrared images in addition to RGB to improve the robustness to problems such as occlusion and illumination. However, acquiring and synchronizing multi-modal sensors often requires higher costs and more complex hardware conditions, and it is difficult to be applicable to general low-cost scenarios.
[0009] Based on the above deficiencies, there is an urgent need for a new solution that can maintain high-precision hand three-dimensional reconstruction results even in the case of severe occlusion, does not rely on overly expensive hardware configurations, and can complete relatively refined hand reconstruction only relying on a small number of views of RGB images. Summary of the Invention
[0010] To achieve the above objects and other advantages of the present invention, the first object of the present invention is to provide a multi-view high-precision hand three-dimensional reconstruction method that integrates two-dimensional completion and implicit optimization, including the following steps:
[0011] Capture the hand-object interaction scene from multiple perspectives to obtain multiple images;
[0012] Process each image into a valid region containing the hand;
[0013] For each perspective image with hand occlusion, complete the occluded hand region at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture;
[0014] Under multiple perspectives, through hand 2D key point detection and multi-perspective fusion of the completed two-dimensional hand images, and combining rough or low-precision calibration of the camera poses of a small number of perspectives, obtain more accurate three-dimensional bone key points and an initial hand mesh;
[0015] Based on implicit representation, perform refined regression on the hand surface, taking into account parametric priors and data-driven features, and reconstruct the complex topology and details of the hand surface.
[0016] Furthermore, the step of completing the occluded hand region at the two-dimensional image level for each perspective image with hand occlusion to obtain a two-dimensional hand image with relatively complete shape and texture includes:
[0017] Based on a completion network that combines a variational autoencoder and a generative adversarial network, complete the occluded hand region at the two-dimensional image level.
[0018] Furthermore, the step of the completion network that combines a variational autoencoder and a generative adversarial network to complete the occluded hand region at the two-dimensional image level includes:
[0019] The encoder E of the variational autoencoder maps the input image to a latent vector z, and then a preliminary completion image I is generated through the decoder D of the variational autoencoder, which is expressed in mathematical form as:
[0020]
[0021] Among them, the input hand image is denoted as The occluded region is regarded as having a mask Ω;
[0022] Introduce a generative adversarial network discriminator Q to perform adversarial training with the decoder output I to ensure that the completion image is consistent with the original hand image in terms of local texture, lighting transition, and overall realism;
[0023] Design the loss function as a weighted sum of pixel-level reconstruction loss and adversarial loss, specifically:
[0024]
[0025] Among them, I refis the corresponding unoccluded reference figure, is the adversarial loss between the discriminator and the generator, λ rec and λ adv are the weight coefficients respectively.
[0026] Furthermore, the steps of obtaining more accurate 3D bone key points and the initial hand mesh under multiple perspectives by performing hand 2D key point detection and multi-perspective fusion on the completed 2D hand images and combining rough or low-precision calibration of the camera poses of a small number of perspectives include:
[0027] For each completed image, a deep learning-based hand key point detector locates the positions of 21 bone joints in the image coordinate system, denoted as:
[0028]
[0029] where N represents the number of perspectives;
[0030] Using a deep learning network, the 2D key points from multiple perspectives are used as inputs to directly regress the bone joints of the hand in the 3D coordinate system. Let this multi-perspective fusion regression network be denoted as F agg , then the 3D key point positions are represented as:
[0031]
[0032] where, are the coordinates of 21 bone joints of the hand with a dimension of 21×3;
[0033] After completing the prediction of , using the parameterized hand model as a prior, align the 3D mesh with in a way that minimizes the 3D key point deviation to obtain the preliminary 3D hand mesh shape.
[0034] Furthermore, the steps of using the parameterized hand model as a prior and aligning the 3D mesh with in a way that minimizes the 3D key point deviation to obtain the preliminary 3D hand mesh shape include:
[0035] Using the multi-perspective key point fusion result, perform parameter fitting on the parameterized hand model. Denote the joint and shape parameters of the parameterized hand model as {θ,β}, then the hand mesh generated by the model is represented as:
[0036]
[0037] where N v represents the number of mesh vertices.
[0038] Further, the step of performing refined regression on the hand surface based on the implicit representation includes:
[0039] Input the vertex coordinates of V or the sampled point cloud into the PointNet network to obtain the 3D spatial feature vector F of the entire hand 3d (X), where X represents any sampled point in the three-dimensional space;
[0040] For each completed image Use a convolutional neural network for encoding. After projecting the three-dimensional point X onto the corresponding coordinates of the image, obtain two-dimensional features through bilinear interpolation Then, combined with visibility weighting, fuse the two-dimensional features from different perspectives into F 2d (X):
[0041]
[0042] where, w i (X) is the weight representing visibility and confidence;
[0043] Based on the preliminary MANO mesh, calculate the signed distance field d′(X). For any point X = (x, y, z) in space, define:
[0044]
[0045] where, δ represents the distance to the hand surface;
[0046] Introduce a learnable correction function:
[0047] f sd (F 2d (X), F 3d (X), d′(X), n′(X)) → (d(X), n(X)),
[0048] where, n′(X) is the normal information given by the initial mesh, and d(X) and n(X) are the more accurate signed distance and normal vector after correction;
[0049] Perform high-frequency position encoding on d′(X):
[0050]
[0051] And input the encoding result together with the two-dimensional and three-dimensional features into the correction network;
[0052] The implicit function outputs the occupancy probability or signed distance, denoted as
[0053] f o (F 2d (X), F 3d(X), d(X), n(X)) → o(X)
[0054] Among them, o(X) ∈ [0, 1] represents whether point X is located inside the hand. When o(X) is greater than the threshold or d(X) is close to 0, it is determined to be near the hand surface.
[0055] Furthermore, the steps of reconciling parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface include:
[0056] By using the Marching Cubes algorithm on the implicit function in three-dimensional space, a complete three-dimensional hand mesh with fine fingertips, joints, and texture wrinkles is obtained.
[0057] Furthermore, the steps of reconciling parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface also include:
[0058] Use a rendering engine to map texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
[0059] The second object of the present invention is to provide a multi-view high-precision three-dimensional hand reconstruction system that integrates two-dimensional completion and implicit optimization. Applying the above method, it includes a data acquisition module, a data processing module, a hand bone key point regression module, a hand vertex reconstruction module, and a terminal visualization module; among them,
[0060] The data acquisition module is used to capture the hand-object interaction scene from multiple perspectives to obtain multiple images;
[0061] The data processing module is used to process each image into a valid area containing the hand;
[0062] The hand bone key point regression module is used to, for each perspective image with hand occlusion, complete the occluded hand area at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture; under multiple perspectives, through two-dimensional hand key point detection and multi-view fusion of the completed two-dimensional hand images, and combined with rough calibration or low-precision calibration of the camera poses of a small number of perspectives, more accurate three-dimensional bone key points and an initial hand mesh are obtained; based on implicit representation, fine regression is performed on the hand surface;
[0063] The hand vertex reconstruction module is used to reconcile parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface;
[0064] The terminal visualization module is used to use a rendering engine to map texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
[0065] The third object of the present invention is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.
[0066] The fourth object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0068] The present invention provides a multi-view high-precision hand three-dimensional reconstruction method and system integrating two-dimensional completion and implicit optimization. The occlusion robustness is improved: the VAE-GAN is used to complete the complementation of the occluded area of the hand at the two-dimensional image level, significantly reducing the feature loss caused by occlusion; combined with subsequent multi-view and implicit optimization, a relatively complete and accurate hand reconstruction effect can still be obtained in complex gesture and strong occlusion scenarios. High-precision hand model: The multi-view key point fusion and 3D bone point regression method effectively reduce the cumulative error caused by the single-view detection deviation; combined with the MANO parameter prior and the self-evolving SDF (SeSDF), it can perform fine regression on complex topologies such as the local surface of the hand, improving the accuracy of the 3D bone key point positions and the details of the hand mesh. Low-cost solution: Only a small amount of ordinary RGB camera data from a few viewpoints is required, without using a depth sensor or high-precision calibration; even if the calibration information is relatively rough, high accuracy can be achieved through the completion and data-driven 3D regression method, reducing the system deployment cost and the threshold for obtaining a hand model.
[0069] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it in accordance with the content of the specification, the following takes the preferred embodiments of the present invention and describes them in detail in conjunction with the accompanying drawings. The specific implementation manner of the present invention is given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0071] Figure 1 is the flow of a multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization Figure 1 ;
[0072] Figure 2 is the flow of a multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization Figure 2 ;
[0073] Figure 3 It is a flowchart for reconstruction based on a trained hand reconstruction network;
[0074] Figure 4 It is an architecture diagram of a multi-view high-precision hand 3D reconstruction system that integrates 2D completion and implicit optimization;
[0075] Figure 5 It is a schematic diagram of a computer device;
[0076] Figure 6 It is a schematic diagram of a computer-readable storage medium. Specific embodiments
[0077] Next, in combination with the accompanying drawings and specific embodiments, the present invention will be further described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. It should be noted that, on the premise of no conflict, the following-described embodiments or technical features can be arbitrarily combined to form new embodiments.
[0078] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0079] In the present application, the accompanying drawing numbers are only used to distinguish each step in the solution and are not used to limit the execution order of each step. The specific execution order shall be subject to the description in the specification.
[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0081] Embodiment 1
[0082] A multi-view high-precision hand 3D reconstruction method that integrates 2D completion and implicit optimization to achieve multi-view hand 3D reconstruction in a hand-object interaction scenario, as Figure 1 - Figure 2 shown, includes the following steps:
[0083] S100. Take pictures of the hand-object interaction scenario from multiple perspectives to obtain multiple images;
[0084] For example, take pictures of the hand-object interaction scenario from 2 to 4 perspectives (or more) to obtain multiple RGB images. Compared with the traditional multi-view scheme that requires precise calibration, the present invention can be completed only under relatively rough calibration or through weak calibration.
[0085] S110. Process each image into an effective area containing the hand;
[0086] Optionally, using a hand segmentation network or other foreground segmentation methods, each image is cropped / normalized to an effective area containing the hand to prepare for subsequent feature extraction and completion.
[0087] S120. For each perspective image with hand occlusion, the occluded hand area is completed at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture;
[0088] In some embodiments, the step of completing the occluded hand area of each perspective image at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture includes:
[0089] Based on a completion network combining a variational autoencoder and a generative adversarial network, the occluded hand area is completed at the two-dimensional image level.
[0090] In this embodiment, for each perspective image with hand occlusion, a method combining a variational autoencoder (VQ-VAE) or VAE and an adversarial network (GAN) is used to first complete the occluded hand area at the two-dimensional image level to obtain a 2D hand image with relatively complete shape and texture.
[0091] Furthermore, the step of the completion network combining a variational autoencoder and a generative adversarial network to complete the occluded hand area at the two-dimensional image level includes:
[0092] The encoder E of the variational autoencoder (VAE) maps the input image to a latent vector z, and then a preliminary completed image I is generated through the decoder D of the variational autoencoder, which is represented in mathematical form as:
[0093]
[0094] Among them, the input hand image is denoted as The occluded area is regarded as having a mask Ω;
[0095] The discriminator Q of the generative adversarial network (GAN) is introduced to perform adversarial training with the decoder output I to ensure that the completed image conforms to the original hand image in terms of local texture, lighting transition, and overall realism;
[0096] To ensure that the occluded part is consistent with the visible hand area in style, the loss function is designed as a weighted sum of the pixel-level reconstruction loss and the adversarial loss, specifically:
[0097]
[0098] Among them, I ref is the corresponding unoccluded reference image (which can be from data annotation or synthesis), is the adversarial loss between the discriminator and the generator, λ rec and λ adv are weight coefficients respectively.
[0099] Through this VAE - GAN framework, the network can learn the reasonable texture and structure that the occluded area should contain, and obtain a complete image consistent with the real hand appearance style.
[0100] Using VAE - GAN to first complete the occluded part in the perspective image avoids the over - reliance of traditional hand detection / triangulation on the complete image. In hand - object interaction, even if some finger or palm areas are blocked by an object, reasonable texture and shape inference can be performed at the 2D level.
[0101] The image after the above - mentioned completion can effectively reduce the 2D information loss caused by object occlusion, thereby providing more accurate input for subsequent key - point detection and 3D reconstruction.
[0102] S130. Under multiple perspectives, by performing hand 2D key - point detection and multi - perspective fusion on the completed 2D hand image, the random error of single - perspective detection can be effectively reduced, and combined with rough calibration or low - precision calibration of the camera poses of a small number of perspectives, more accurate 3D bone key - points and the initial hand mesh (fitted using a parametric hand model such as MANO) can be obtained;
[0103] In some embodiments, the step of obtaining more accurate 3D bone key - points and the initial hand mesh by performing hand 2D key - point detection and multi - perspective fusion on the completed 2D hand image under multiple perspectives and combining rough calibration or low - precision calibration of the camera poses of a small number of perspectives includes:
[0104] For each completed image, a deep - learning - based hand key - point detector locates the positions of 21 bone joints in the image coordinate system, denoted as:
[0105]
[0106] where N represents the number of perspectives;
[0107] Since VAE - GAN has completed the missing finger or palm areas in the image, the detector can more accurately locate the 2D key - points of the occluded parts.
[0108] To obtain high - precision 3D bone key - points, the present invention does not directly use traditional geometric triangulation, but takes the 2D key - points of multiple perspectives as input through a deep - learning network, and directly regresses the bone joints of the hand in the 3D coordinate system. Let this multi - perspective fusion regression network (multi - layer perceptron MLP) be denoted as F agg , then the 3D key - point position is expressed as:
[0109]
[0110] Among them, is the coordinates of 21 bone joint points of the hand with a dimension of 21×3;
[0111] This regression network can not only make full use of multi-view information to reduce the influence of misdetection or noise in a single view, but also effectively improve the 3D bone positioning accuracy when the precise camera calibration is lacking.
[0112] After completing the prediction of , using the parametric hand model (MANO) as a prior, align the 3D mesh with in a way that minimizes the 3D key point deviation, so as to obtain the preliminary 3D mesh shape of the hand.
[0113] Furthermore, the step of using the parametric hand model as a prior and aligning the 3D mesh with to obtain the preliminary 3D mesh shape of the hand includes:
[0114] Using the multi-view key point fusion result, perform parameter fitting on the parametric hand model (MANO). Denote the joint and shape parameters of the parametric hand model (MANO) as {θ,β}, then the hand mesh generated by the model is expressed as:
[0115]
[0116] where N v represents the number of mesh vertices, usually 799. By minimizing the key point projection error in the image plane or the 3D key point error, a mesh V that is as close as possible to the actual hand shape and pose can be obtained. This mesh can reflect the hand bones and basic deformations, and subsequent further fine reconstruction is still carried out on local details such as fingertips, finger seams, and wrinkles.
[0117] Using the implicit representation (self-evolving signed distance field, i.e., the SeSDF module) to perform refined regression on the hand surface, taking into account the parametric prior (bones, shape) and data-driven features (2D completion results and 3D mesh features), the complex topology and details of the hand surface can be robustly reconstructed.
[0118] This embodiment introduces the self-evolving SDF (SeSDF) module, combines the prior of the parametric hand model MANO and the 2D / 3D features extracted from multi-view data, and realizes the precise regression of local details on the hand surface. This implicit function can correct the geometric error on the basis of the preliminary MANO mesh and retain details such as fingertips, finger seams, and surface textures in the real scene.
[0119] S140. Perform refined regression on the hand surface based on implicit representation, taking into account parametric priors and data-driven features, and reconstruct the complex topology and details of the hand surface.
[0120] In some embodiments, the step of performing refined regression on the hand surface based on implicit representation includes:
[0121] After obtaining the initial 3D mesh V, to improve the regression accuracy of the subsequent implicit representation, it is necessary to extract multi-source features and fuse them:
[0122] 3D feature extraction: Input the vertex coordinates or sampled point cloud of V into a network such as PointNet to obtain the 3D spatial feature vector F 3d (X) of the whole hand, where X represents any sampled point (near the vertex or the entire hand bounding box) in 3D space;
[0123] Multi-view 2D feature extraction: For each completed image Use a convolutional neural network (CNN) for encoding. After projecting the 3D point X to the corresponding coordinates of the image, obtain the 2D feature through bilinear interpolation Then, combined with visibility weighting, fuse the 2D features from different views into F 2d (X):
[0124]
[0125] where, w i (X) is the weight representing visibility and confidence;
[0126] To make up for the deficiencies of the MANO mesh in local details, referring to multi-view human 3D reconstruction, introduce an implicit representation based on the signed distance field (SDF), and perform self-evolution (Self-Evolved SDF, SeSDF) on this basis. This process is generally divided into two parts:
[0127] SDF initialization: Calculate the signed distance field d′(X) based on the initial MANO mesh. For any point X = (x, y, z) in space, define:
[0128]
[0129] where, δ represents the distance to the hand surface;
[0130] Since the MANO mesh lacks complex shapes, d ′ (X) is often only a rough approximation.
[0131] Self-evolution correction (SeSDF): Introduce a learnable correction function:
[0132] f sd(F 2d (X), F 3d (X), d′(X), n′(X)) → (d(X), n(X)),
[0133] Among them, n′(X) is the normal information given by the initial grid, and d(X) and n(X) are the more accurate signed distance and normal vector after correction;
[0134] This network integrates multi-view 2D features, 3D features, and MANO prior information, and can adaptively correct the distance field and refine the local surface of the hand
[0135] Distance Encoding: To enhance the sensitivity to local geometry, high-frequency position encoding is performed on d′(X):
[0136]
[0137] And the encoding result is jointly input into the correction network together with the two-dimensional and three-dimensional features;
[0138] Finally, the implicit function outputs the occupancy probability or signed distance, denoted as
[0139] f o (F 2d (X), F 3d (X), d(X), n(X)) → o(X)
[0140] Among them, o(X) ∈ [0, 1] represents whether the point X is inside the hand. When o(X) is greater than the threshold or d(X) is close to 0, it can be determined that it is near the hand surface.
[0141] In some embodiments, the steps of taking into account parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface include:
[0142] By using the Marching Cubes algorithm on the implicit function in three-dimensional space, a complete three-dimensional hand mesh with fine fingertips, joints, and texture wrinkles can be obtained.
[0143] Compared with complex systems that rely on multi-modalities (such as RGB + depth), the method of the present invention only requires RGB images from a small number of viewpoints, and improves the three-dimensional reconstruction accuracy through a combination of completion and implicit optimization, reducing the hardware system and data acquisition costs.
[0144] Furthermore, the steps of taking into account parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface further include:
[0145] For further visualization, a rendering engine can be used to map the texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
[0146] As Figure 3 shown, after the trained hand reconstruction model, obtain the multi-view RGB images of the hand to be reconstructed, and use the multi-view RGB images as the input of the trained hand reconstruction model to obtain a high-precision hand reconstruction model.
[0147] This embodiment provides a high-precision three-dimensional hand reconstruction method for occlusion scenarios, which can be applied to scenarios that require real-time and accurate hand modeling, such as virtual reality interaction systems, intelligent gesture control devices, and medical simulation training platforms. By combining an image generation and completion network with implicit representation technology, high-quality completion and three-dimensional reconstruction of the occluded area of the hand can still be achieved under the premise of only using a small number of perspective RGB images.
[0148] Embodiment 2
[0149] A multi-view high-precision three-dimensional hand reconstruction system that integrates two-dimensional completion and implicit optimization applies the method provided in Embodiment 1. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be elaborated here. As Figure 4 shown, it includes a data acquisition module, a data processing module, a hand bone key point regression module, a hand vertex reconstruction module, and a terminal visualization module; among them,
[0150] The data acquisition module is used to capture the hand-object interaction scene from multiple perspectives to obtain multiple images;
[0151] The data processing module is used to process each image into a valid area containing the hand;
[0152] The hand bone key point regression module is used to, for each perspective image with hand occlusion, complete the occluded hand area at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture; under multiple perspectives, through hand 2D key point detection and multi-view fusion of the completed two-dimensional hand images, and combining rough calibration or low-precision calibration of the camera poses of a small number of perspectives, more accurate three-dimensional bone key points and an initial hand mesh are obtained; and fine regression of the hand surface is performed based on implicit representation;
[0153] The hand vertex reconstruction module is used to reconstruct the complex topology and details of the hand surface while taking into account parametric priors and data-driven features;
[0154] The terminal visualization module is used to use a rendering engine to map the texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
[0155] Based on the technical solution of the above embodiment, optionally, the step of complementing the occluded hand region at the two-dimensional image level for each perspective image with hand occlusion to obtain a two-dimensional hand image with relatively complete shape and texture includes:
[0156] Based on a complementation network that combines a variational autoencoder and a generative adversarial network, complement the occluded hand region at the two-dimensional image level.
[0157] Based on the technical solution of the above embodiment, optionally, the step of the complementation network that combines a variational autoencoder and a generative adversarial network to complement the occluded hand region at the two-dimensional image level includes:
[0158] The encoder E of the variational autoencoder maps the input image to a latent vector z, and then a preliminary complemented image I is generated through the decoder D of the variational autoencoder, which is represented in mathematical form as:
[0159]
[0160] Among them, the input hand image is denoted as The occluded region is regarded as having a mask Ω;
[0161] Introduce a generative adversarial network discriminator Q to perform adversarial training with the output I of the decoder to ensure that the complemented image is consistent with the original hand image in terms of local texture, illumination transition, and overall realism;
[0162] The loss function is designed as the weighted sum of the pixel-level reconstruction loss and the adversarial loss, specifically:
[0163]
[0164] Among them, I ref is the corresponding unoccluded reference image, is the adversarial loss between the discriminator and the generator, and λ rec and λ adv are weight coefficients respectively.
[0165] Based on the technical solution of the above embodiment, optionally, the step of, under multiple perspectives, obtaining more accurate three-dimensional bone key points and an initial hand mesh by performing hand 2D key point detection and multi-perspective fusion on the complemented two-dimensional hand image and combining rough calibration or low-precision calibration of the camera poses of a small number of perspectives includes:
[0166] For each complemented image, a deep learning-based hand key point detector locates the positions of 21 bone joints in the image coordinate system, denoted as:
[0167]
[0168] Wherein, N represents the number of viewing angles;
[0169] The multi-view 2D key points are used as inputs to a deep learning network to directly regress the bone joint points of the hand in a 3D coordinate system, and let this multi-view fusion regression network be denoted as F agg , then the 3D key point position is expressed as:
[0170]
[0171] Wherein, is the coordinate of 21 bone joint points of the hand with a dimension of 21×3;
[0172] After completing the prediction of , using a parametric hand model as a prior, align the 3D mesh with in a way that minimizes the 3D key point deviation to obtain a preliminary 3D mesh shape of the hand.
[0173] Based on the technical solution of the above embodiment, optionally, the step of using a parametric hand model as a prior and aligning the 3D mesh with in a way that minimizes the 3D key point deviation to obtain a preliminary 3D mesh shape of the hand includes:
[0174] Using the multi-view key point fusion result, perform parameter fitting on the parametric hand model. Denote the joint and shape parameters of the parametric hand model as {θ,β}, then the hand mesh generated by the model is expressed as:
[0175]
[0176] Wherein, N v represents the number of mesh vertices.
[0177] Based on the technical solution of the above embodiment, optionally, the step of performing refined regression on the hand surface based on implicit representation includes:
[0178] Input the vertex coordinates of V or the sampled point cloud into the PointNet network to obtain the 3D spatial feature vector F 3d (X) of the whole hand, where X represents any sampled point in 3D space;
[0179] For each completed image Encode it using a convolutional neural network. After projecting the 3D point X to the corresponding coordinates of the image, obtain the 2D feature through bilinear interpolation Then, combined with visibility weighting, fuse the 2D features from different views into F 2d (X):
[0180]
[0181] where w i (X) is the weight representing visibility and confidence;
[0182] Based on the preliminary MANO mesh, calculate the signed distance field d′(X). For any point X = (x, y, z) in space, define:
[0183]
[0184] where δ represents the distance to the hand surface;
[0185] Introduce a learnable correction function:
[0186] f sd (F 2d (X), F 3d (X), d′(X), n′(X)) → (d(X), n(X)),
[0187] where n′(X) is the normal information given by the initial mesh, and d(X) and n(X) are the more accurate signed distance and normal vector after correction;
[0188] Perform high-frequency position encoding on d′(X):
[0189]
[0190] And input the encoding result together with two-dimensional and three-dimensional features into the correction network;
[0191] The implicit function outputs the occupancy probability or signed distance, denoted as
[0192] f o (F 2d (X), F 3d (X), d(X), n(X)) → o(X)
[0193] where o(X) ∈ [0, 1] characterizes whether the point X is inside the hand. When o(X) is greater than the threshold or d(X) is close to 0, it is determined to be near the hand surface.
[0194] Based on the technical solution of the above embodiment, optionally, the steps of reconciling parametric priors and data-driven features to reconstruct the complex topology and details of the hand surface include:
[0195] By using the Marching Cubes algorithm on the implicit function in three-dimensional space, obtain a complete three-dimensional hand mesh with fine fingertips, joints, and texture wrinkles.
[0196] This embodiment provides a high-precision three-dimensional hand reconstruction system for occluded scenarios, which can be applied to scenarios that require real-time and accurate hand modeling, such as virtual reality interaction systems, intelligent gesture control devices, and medical simulation training platforms. By combining an image generation and completion network with implicit representation technology, high-quality completion and three-dimensional reconstruction of the occluded area of the hand can still be achieved under the premise of only using a small number of perspective RGB images.
[0197] Embodiment 3
[0198] A computer device 200, as Figure 5 shown, includes a memory 210, a processor 220, and a computer program 230 stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a multi-view high-precision three-dimensional hand reconstruction method that combines two-dimensional completion and implicit optimization. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be repeated here.
[0199] Embodiment 4
[0200] A computer-readable storage medium, as Figure 6 shown, stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of a multi-view high-precision three-dimensional hand reconstruction method that combines two-dimensional completion and implicit optimization. For a detailed description of the method, reference can be made to the corresponding description in the above method embodiments, which will not be repeated here.
[0201] The number of devices and the scale of processing described here are used to simplify the description of the present invention. Applications, modifications, and variations of the present invention will be obvious to those skilled in the art.
[0202] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.
[0203] The device, computer device, non-volatile computer storage medium, and method provided in the embodiments of this specification are corresponding. Therefore, the device, computer device, and non-volatile computer storage medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding device, computer device, and non-volatile computer storage medium will not be repeated here.
[0204] Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps so that the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software units for implementing the method or the structures within the hardware component.
[0205] The systems, devices, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various units according to functions and described separately. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0206] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0207] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0208] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so as to cause a series of operational steps to be performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.
[0210] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element qualified by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising said element.
[0211] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program units. Generally, program units include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The specification may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program units may be located in both local and remote computer storage media including storage devices.
[0212] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments may be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and the relevant parts may be referred to the partial description of method embodiments.
[0213] The above is only for the embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of one or more embodiments of this specification.
Claims
1. A multi-view high-precision 3D hand reconstruction method integrating 2D completion and implicit optimization, characterized in that, Including the following steps: Shoot the hand-object interaction scene from multiple perspectives to obtain multiple images; Process each image into a valid region containing the hand; For each perspective image with hand occlusion, complement the occluded hand region at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture; Under multiple perspectives, through hand 2D key point detection and multi-perspective fusion of the complemented two-dimensional hand images, and combined with rough calibration or low-precision calibration of the camera poses of a small number of perspectives, obtain more accurate three-dimensional bone key points and the initial hand mesh; Based on implicit representation, perform refined regression on the hand surface, taking into account parametric priors and data-driven features, and reconstruct the complex topology and details of the hand surface.
2. A multi-view high-precision 3D hand reconstruction method integrating two-dimensional completion and implicit optimization as claimed in claim 1, characterized in that, The step of complementing the occluded hand region at the two-dimensional image level for each perspective image with hand occlusion to obtain a two-dimensional hand image with relatively complete shape and texture includes: Based on a complementation network combining a variational autoencoder and a generative adversarial network, complement the occluded hand region at the two-dimensional image level.
3. A multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization as described in claim 2, characterized in that, The step of the complementation network based on the combination of a variational autoencoder and a generative adversarial network complementing the occluded hand region at the two-dimensional image level includes: The encoder E of the variational autoencoder maps the input image to a latent vector z, and then generates a preliminary complemented image I through the decoder D of the variational autoencoder, which is expressed in mathematical form as: wherein, the input hand image is denoted as the occluded region is regarded as having a mask Ω; Introduce a generative adversarial network discriminator Q to perform adversarial training with the decoder output I to ensure that the complemented image is consistent with the original hand image in terms of local texture, lighting transition, and overall realism; Design the loss function as a weighted sum of pixel-level reconstruction loss and adversarial loss, specifically: Among them, I ref is the corresponding unoccluded reference diagram, is the adversarial loss between the discriminator and the generator, and λ rec and λ adv are weight coefficients respectively.
4. A multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization according to claim 3, characterized in that The step of under multiple perspectives, through hand 2D key point detection and multi-perspective fusion of the complemented two-dimensional hand images, and combined with rough calibration or low-precision calibration of the camera poses of a small number of perspectives, to obtain more accurate three-dimensional bone key points and the initial hand mesh includes: For each complemented image, based on a deep learning-based hand key point detector, locate the positions of 21 bone joint points in the image coordinate system, denoted as: where N represents the number of perspectives; Taking the two-dimensional key points from multiple perspectives as input through a deep learning network, directly regress the bone joint points of the hand in the three-dimensional coordinate system, and denote this multi-perspective fusion regression network as F agg , then the position of the three-dimensional key points is expressed as: Among them, are the coordinates of 21 bone joint points of the hand with a dimension of 21×3; After completing the prediction of , using the parameterized hand model as a prior, align the 3D mesh with in a way that minimizes the 3D keypoint deviation to obtain the initial 3D hand mesh shape.
5. A multi-view high-precision hand 3D reconstruction method integrating two-dimensional completion and implicit optimization as claimed in claim 4, characterized in that Using the parametric hand model as a prior, aligning the 3D mesh with in a way that minimizes the 3D key point deviation, and the steps for obtaining the preliminary 3D hand mesh shape include: aligning, getting the preliminary 3D hand mesh shape steps include: Using the multi-perspective key point fusion result, perform parameter fitting on the parametric hand model. Denote the joint and shape parameters of the parametric hand model as {θ,β}, then the hand mesh generated by the model is expressed as: Among them, N v represents the number of grid vertices.
6. The multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization as claimed in claim 5, wherein, The step of performing refined regression on the hand surface based on implicit representation includes: Input the vertex coordinates of V or the sampled point cloud into the PointNet network to obtain the 3D spatial feature vector F of the entire hand 3d (X), where X represents any sampled point in the three-dimensional space; For each completed image It is encoded using a convolutional neural network. After projecting the 3D point X onto the corresponding coordinates of the image, a 2D feature is obtained through bilinear interpolation Then, combined with visibility weighting, the 2D features from different perspectives are fused into F 2d (X): where w i (X) is the weight representing visibility and confidence; Calculate the signed distance field d′(X) based on the preliminary MANO mesh. For any point X=(x,y,z) in space, define: where δ represents the distance to the hand surface; Introduce a learnable correction function: f sd (F 2d (X), F 3d (X), d′(X), n′(X)) → (d(X), n(X)), where n′(X) is the normal information given by the initial mesh, and d(X) and n(X) are the more accurate signed distance and normal vector after correction; Perform high-frequency position encoding on d′(X); And input the encoded result together with two-dimensional and three-dimensional features into the correction network; The implicit function outputs the occupancy probability or signed distance, denoted as f o (F 2d (X),F 3d (X),d(X),n(X))→o(X) Among them, o(X) ∈ [0, 1] represents whether point X is located inside the hand. When o(X) is greater than the threshold or d(X) is close to 0, it is determined to be near the hand surface.
7. A multi-view high-precision 3D hand reconstruction method integrating two-dimensional completion and implicit optimization as claimed in claim 6, characterized in that The steps of reconstructing the complex topology and details of the hand surface by taking into account the parametric prior and data-driven features include: By using the Marching Cubes algorithm on the implicit function in three-dimensional space, a complete three-dimensional hand mesh with fine fingertips, joints, and texture folds is obtained.
8. The multi-view high-precision hand three-dimensional reconstruction method integrating two-dimensional completion and implicit optimization according to claim 7, characterized in that, The steps of reconstructing the complex topology and details of the hand surface by taking into account the parametric prior and data-driven features also include: Using a rendering engine to map texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
9. A multi-view high-precision hand three-dimensional reconstruction system integrating two-dimensional completion and implicit optimization, which applies the method according to any one of claims 1 to 8, characterized in that: It includes a data acquisition module, a data processing module, a hand bone key point regression module, a hand vertex reconstruction module, and a terminal visualization module; among them, The data acquisition module is used to capture the hand-object interaction scene from multiple perspectives to obtain multiple images. The data processing module is used to process each image into an effective area containing the hand. The hand bone key point regression module is used to, for each perspective image with hand occlusion, complement the occluded hand area at the two-dimensional image level to obtain a two-dimensional hand image with relatively complete shape and texture; under multiple perspectives, through hand 2D key point detection and multi-perspective fusion of the complemented two-dimensional hand images, and combined with rough calibration or low-precision calibration of the camera poses of a small number of perspectives, more accurate three-dimensional bone key points and an initial hand mesh are obtained; and refined regression of the hand surface is performed based on implicit representation. The hand vertex reconstruction module is used to reconstruct the complex topology and details of the hand surface by taking into account the parametric prior and data-driven features. The terminal visualization module is used to use a rendering engine to map texture maps fused from multiple perspectives onto the reconstructed mesh, thereby presenting a realistic hand model.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Cited By
Clothing care flatness monitoring method, electronic equipment and fabric cleaning equipment
CN120543551A
Deep foundation pit crack feature coding detection method based on image recognition
CN121169994A