Monocular hand-object interaction three-dimensional reconstruction method and device, terminal and storage medium

CN122636873BActive Publication Date: 2026-09-25PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611082421.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-25
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

该方法在对手的三维结果和物体的三维结果进行组合的时候,容易出现错位的情况,导致最终的手-物交互重建结果的质量达不到预期效果

Benefits of technology

本申请实施例在手-物交互三维重建过程中引入了粗配准,以避免手的三维结果和物体的三维结果在组合的时候出现错位的情况。具体的,本申请实施例获取目标图像对应的初始物体网格、初始手部网格和目标手-物交互网格;对所述目标手-物交互网格进行手部区域和物体区域提取处理,获得目标手部粗区域和目标物体粗区域;采用物体配准算法将所述初始物体网格粗配准到所述目标物体粗区域,获得优化物体网格;采用手部配准算法将所述初始手部网格粗配准到所述目标手部粗区域,获得优化手部网格;采用联合优化算法对所述优化物体网格和所述优化手部网格进行参数调整处理,获得所述目标图像对应的手-物交互重建结果。与现有技术相比,本申请实施例对初始物体网格和初始手部网格进行粗配准处理,以获得符合整体交互关系的优化值,也即优化物体网格和优化手部网格,接着,对优化物体网格和优化手部网格进行联合优化以获得手-物交互重建结果,由此,本申请实施例能够避免手的三维结果和物体的三维结果在组合的时候出现错位的情况,从而提高手-物交互重建结果的质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636873B_ABST
    Figure CN122636873B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image three-dimensional reconstruction. The application discloses a monocular hand-object interaction three-dimensional reconstruction method and device, a terminal and a storage medium, which can improve the quality of a hand-object interaction reconstruction result. The method comprises the following steps: acquiring an initial object grid, an initial hand grid and a target hand-object interaction grid corresponding to a target image; performing hand region and object region extraction processing on the target hand-object interaction grid to obtain a target hand rough region and a target object rough region; performing coarse registration of the initial object grid to the target object rough region by using an object registration algorithm to obtain an optimized object grid; performing coarse registration of the initial hand grid to the target hand rough region by using a hand registration algorithm to obtain an optimized hand grid; and performing parameter adjustment processing on the optimized object grid and the optimized hand grid by using a joint optimization algorithm to obtain a hand-object interaction reconstruction result corresponding to the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image 3D reconstruction technology. More specifically, this application relates to a monocular hand-object interactive 3D reconstruction method, apparatus, terminal, and storage medium. Background Technology

[0002] Traditional hand-object interaction 3D reconstruction methods typically begin by detecting, segmenting, or locating the hand and object regions in the input image, then using different models to reconstruct the hand and object in 3D. The hand is generally reconstructed using a parametric 3D hand pose estimation model to obtain its position, pose, and parametric representation. The object is usually recovered using methods such as object detection, pose estimation, category prior matching, or 3D shape reconstruction to obtain its 3D shape, spatial position, and orientation. Then, through coordinate transformation, contact constraints, or joint optimization, the hand's position, pose, and parametric representation, along with the object's 3D shape, spatial position, and orientation, are combined into a unified 3D space to form the final hand-object interaction reconstruction result. In other words, this method first independently constructs the 3D results of the hand and object, then directly combines them to form the final hand-object interaction reconstruction result. However, this method is prone to misalignment when combining the 3D results of the hand and object, leading to a final hand-object interaction reconstruction result that does not meet expectations. Summary of the Invention

[0003] The purpose of this application is to provide a monocular hand-object interactive 3D reconstruction method, apparatus, terminal, and storage medium, which can improve the quality of hand-object interactive reconstruction results. This application is mainly achieved through the following technical solutions: A first aspect of this application provides a monocular hand-object interaction 3D reconstruction method, comprising: Obtain the initial object mesh, initial hand mesh, and target hand-object interaction mesh corresponding to the target image; The target hand-object interaction mesh is processed to extract the hand region and object region to obtain the target hand coarse region and the target object coarse region; An object registration algorithm is used to coarsely register the initial object mesh to the coarse region of the target object to obtain an optimized object mesh. The initial hand mesh is coarsely registered to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh; A joint optimization algorithm is used to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image.

[0004] According to one embodiment of this application, the step of extracting the hand region and object region of the target hand-object interaction mesh to obtain the target hand coarse region and target object coarse region includes: The target hand-object interaction mesh is converted into a point cloud representation; A point cloud segmentation network is used to perform point-by-point semantic classification on the point cloud representation to obtain the coarse region of the target hand and the coarse region of the target object.

[0005] According to one embodiment of this application, the step of obtaining the target hand-object interaction mesh includes: Obtain the hand-object region cropped image corresponding to the target image; A three-dimensional reconstruction model is used to extract the hand-object interaction mesh from the cropped image of the hand-object region, thereby obtaining the target hand-object interaction mesh.

[0006] According to one embodiment of this application, the step of obtaining the hand-object region cropped image corresponding to the target image includes: A visual language model is used to perform object category and appearance description recognition processing on the target image to obtain the target object category and target object description. The target image, the target object category, and the target object description are input into the segmentation model to perform mask extraction processing on the interaction region between the hand and the object, thereby obtaining the target mask; The target image is cropped based on the target mask to obtain a cropped image of the hand-object region corresponding to the target image.

[0007] According to one embodiment of this application, the step of obtaining the initial object mesh includes: The object region cropping image of the hand-object region is subjected to object region occlusion and completion processing to obtain the object completion image corresponding to the target image; The object mesh is extracted from the object completion image using the 3D reconstruction model to obtain the initial object mesh.

[0008] According to one embodiment of this application, the step of obtaining the initial hand mesh includes: The target image is processed by a hand pose estimation model or a hand reconstruction model to extract the hand mesh, thereby obtaining the initial hand mesh.

[0009] According to an embodiment of this application, the step of using a joint optimization algorithm to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image includes: The total loss function of the joint optimization algorithm is used to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image.

[0010] A second aspect of this application provides a monocular hand-object interactive three-dimensional reconstruction device, comprising: The acquisition module is used to acquire the initial object mesh, initial hand mesh, and target hand-object interaction mesh corresponding to the target image; The region extraction module is used to extract the hand region and object region of the target hand-object interaction mesh to obtain the target hand coarse region and the target object coarse region. The first coarse registration module is used to coarsely register the initial object mesh to the coarse region of the target object using an object registration algorithm to obtain an optimized object mesh. The second coarse registration module is used to coarsely register the initial hand mesh to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh; The joint optimization module is used to perform parameter adjustment processing on the optimized object mesh and the optimized hand mesh using a joint optimization algorithm to obtain the hand-object interaction reconstruction result corresponding to the target image.

[0011] A third aspect of this application provides a terminal device, including a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to execute the steps of the monocular hand-object interactive three-dimensional reconstruction method provided in the first aspect of this application.

[0012] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the monocular hand-object interactive three-dimensional reconstruction method provided in the first aspect of this application.

[0013] The beneficial effects of the embodiments of this application include: This application introduces coarse registration in the hand-object interaction 3D reconstruction process to avoid misalignment when combining the 3D results of the hand and the object. Specifically, this application obtains an initial object mesh, an initial hand mesh, and a target hand-object interaction mesh corresponding to the target image; it extracts the hand and object regions from the target hand-object interaction mesh to obtain a coarse target hand region and a coarse target object region; it uses an object registration algorithm to coarsely register the initial object mesh to the target object region to obtain an optimized object mesh; it uses a hand registration algorithm to coarsely register the initial hand mesh to the target hand region to obtain an optimized hand mesh; and it uses a joint optimization algorithm to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image. Compared with the prior art, the embodiments of this application perform coarse registration processing on the initial object mesh and the initial hand mesh to obtain optimized values ​​that conform to the overall interaction relationship, that is, optimize the object mesh and optimize the hand mesh. Then, the optimized object mesh and the optimized hand mesh are jointly optimized to obtain the hand-object interaction reconstruction result. Thus, the embodiments of this application can avoid misalignment when the three-dimensional results of the hand and the three-dimensional results of the object are combined, thereby improving the quality of the hand-object interaction reconstruction result. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 The flowcharts for some embodiments of the monocular hand-object interaction 3D reconstruction method of this application are shown below. Figure 2 The flowcharts for some other embodiments of the monocular hand-object interaction 3D reconstruction method of this application are shown below; Figure 3 Flowcharts of some embodiments of the joint optimization of this application; Figure 4 This is a schematic diagram of the monocular hand-object interactive 3D reconstruction device of this application in some embodiments; Figure 5 This is a schematic block diagram of the terminal device of this application in some embodiments. Detailed Implementation

[0016] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0017] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0018] The terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0019] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or apparatus.

[0020] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.

[0021] Hand-Object Interaction (HOI) mesh reconstruction is an important research direction in computer vision and 3D vision, with significant application value in virtual reality, human-computer interaction, motion understanding, dexterity hand manipulation learning, and digital human modeling. Hand-Object Interaction mesh reconstruction refers to recovering the 3D geometry of the hand and the interacting object from a single image, as well as their relative positional relationship in 3D space. Compared to simply recovering the hand or only the object, hand-Object Interaction reconstruction not only requires ensuring the reasonable geometric structure and accurate spatial position of both the hand and the object, but also needs to maintain the realistic contact state, occlusion relationship, and relative pose relationship between them as much as possible. Therefore, this task has greater complexity and greater technical challenges.

[0022] Traditional hand-object interaction 3D reconstruction methods typically begin by detecting, segmenting, or locating the hand and object regions in the input image, then using different models to reconstruct the hand and object in 3D. The hand is generally reconstructed using a parametric 3D hand pose estimation model to obtain its position, pose, and parametric representation. The object is usually recovered using methods such as object detection, pose estimation, category prior matching, or 3D shape reconstruction to obtain its 3D shape, spatial position, and orientation. Then, through coordinate transformation, contact constraints, or joint optimization, the hand's position, pose, and parametric representation, along with the object's 3D shape, spatial position, and orientation, are combined into a unified 3D space to form the final hand-object interaction reconstruction result. In other words, this method first independently constructs the 3D results of the hand and object, then directly combines them to form the final hand-object interaction reconstruction result. However, this method is prone to misalignment when combining the 3D results of the hand and object, leading to a final hand-object interaction reconstruction result that does not meet expectations. To overcome the aforementioned technical problems, this application provides a monocular hand-object interactive 3D reconstruction method. The specific embodiments of this application will be further described below with reference to the accompanying drawings.

[0023] refer to Figure 1 The diagram shown is a flowchart of a monocular hand-object interactive 3D reconstruction method provided in the first aspect of an embodiment of this application. Figure 1 The monocular hand-object interactive 3D reconstruction method includes the following steps S1, S2, S3, S4 and S5.

[0024] S1. Obtain the initial object mesh, initial hand mesh, and target hand-object interaction mesh corresponding to the target image.

[0025] The target image is an RGB image (R represents red, G represents green, and B represents blue).

[0026] Further, the initial object mesh acquisition step includes: performing object region occlusion completion processing on the hand-object region cropped image to obtain the object completion image corresponding to the target image; and using the 3D reconstruction model to perform object mesh extraction processing on the object completion image to obtain the initial object mesh (see reference). Figure 2 The steps include "object completion image", "3D reconstruction of large model" and "individual reconstruction of object mesh".

[0027] The three-dimensional reconstruction model can be Hunyuan3D-2 (also known as Hunyuan 3D-2). In other embodiments, the three-dimensional reconstruction model can also be other models, which can be set by those skilled in the art according to actual needs.

[0028] The initial object mesh can be expressed as: .

[0029] Furthermore, the step of performing object region occlusion completion processing on the hand-object region cropped image to obtain the object completion image corresponding to the target image includes: using an image editing tool or an image completion model to complete the object region occluded by the hand in the hand-object region cropped image according to the target object category and target object description, thereby obtaining the object completion image corresponding to the target image.

[0030] The image editing tool can be any commercially available image processing tool used to analyze, repair, enhance, complete, and synthesize digital images.

[0031] The image completion model can be Qwen-Image-Edit or FLUX.1. FLUX.1 is an open-source text-to-image model. In other embodiments, the image completion model can also be other models, which can be set by those skilled in the art according to actual needs.

[0032] The object image in the object completion image is an approximately complete object image.

[0033] Further, the step of using a 3D reconstruction model to extract object meshes from the object completion image to obtain the initial object mesh includes: extracting the first visual semantic features of the object region in the object completion image through the visual encoder of the 3D reconstruction model; using the first visual semantic features as a conditionally guided backbone network (generally based on a conditional diffusion Transformer architecture, DiT) to perform multi-step denoising in the 3D latent space, so that the latent vectors gradually evolve from random noise into a first 3D representation that conforms to the structure of the object region in the object completion image; and using the pre-trained decoder in the 3D reconstruction model to decode the first 3D representation into the initial object mesh.

[0034] The initial object mesh is a complete 3D mesh.

[0035] Further, the initial hand mesh acquisition step includes: performing hand mesh extraction processing on the target image using a hand pose estimation model or a hand reconstruction model to obtain the initial hand mesh (refer to...). Figure 2 The steps are: "target image", "hand pose estimation model" and "MANO parameterized hand network".

[0036] The hand pose estimation model is existing technology.

[0037] The hand reconstruction model can be HaMeR (Hand Mesh Recovery). HaMeR is a monocular 3D hand mesh reconstruction model with a full Transformer architecture. In other embodiments, the hand reconstruction model can also be other models, which can be set by those skilled in the art according to actual needs.

[0038] Further, the step of using a hand reconstruction model to extract hand mesh from the target image to obtain the initial hand mesh includes: mapping the target image into a visual token representation (or visual token representation) using the visual encoder of the hand reconstruction model; using the Transformer-based regression head in the hand reconstruction model to predict the MANO hand pose and shape parameters of the visual token representation to obtain the target MANO hand pose and target shape parameters; and using the target MANO hand pose and target shape parameters as the initial hand mesh.

[0039] The visual encoder of the hand reconstruction model can be a Vision Transformer. In other embodiments, the visual encoder of the hand reconstruction model can also be other encoders, which can be set by those skilled in the art according to actual needs.

[0040] The initial hand mesh is a hand mesh that conforms to MANO parametric constraints, and the initial hand mesh is expressed as follows: ,in, It is a MANO (hand Model with Articulated and Non-rigid defOrmations, a parametric hand model with joint mobility and non-rigid deformation capabilities) model; It is the shape of the MANO model (i.e., the shape parameters of the initial object mesh). , It is a 10-dimensional real vector; These are the pose parameters of the MANO model (i.e., the hand pose parameters of the initial hand mesh). , It is a 48-dimensional real vector.

[0041] The MANO model is a parametric hand model that can generate a structured 3D hand mesh using posture and shape parameters to ensure that hand geometry and joint movements conform to physiological laws.

[0042] Further, the step of obtaining the target hand-object interaction mesh includes: obtaining a cropped image of the hand-object region corresponding to the target image; and using a 3D reconstruction model to perform hand-object interaction mesh extraction processing on the cropped image of the hand-object region to obtain the target hand-object interaction mesh (refer to...). Figure 2 The steps include "image cropping of hand-object region", "3D reconstruction of large model" and "overall reconstruction of hand-object mesh".

[0043] Further, the step of obtaining the hand-object region cropped image corresponding to the target image includes: using a visual language model to perform object category and appearance description recognition processing on the target image to obtain the target object category and target object description; inputting the target image, the target object category, and the target object description into a segmentation model to perform mask extraction processing of the hand and object interaction area to obtain a target mask; and cropping the target image based on the target mask to obtain the hand-object region cropped image corresponding to the target image.

[0044] For example, when the prompt word of the visual language model is "Please answer what object the hand in the picture is interacting with? Only answer the category of the object and a brief description of its appearance", the category of the target object can be "muzzle" or "remote control", and the description of the target object can be "gray" or "black".

[0045] The visual language model can be Qwen-3.5 (also known as Qianwen 3.5). In other embodiments, the visual language model can also be other models, which can be set by those skilled in the art according to actual needs.

[0046] Further, the step of inputting the target image, the target object category, and the target object description into a segmentation model to perform mask extraction processing on the interaction region between the hand and the object to obtain the target mask includes: the segmentation model performing hand mask segmentation processing on the target image to obtain a target hand mask; the segmentation model performing object mask segmentation processing on the object region in the target image based on the target object category and the target object description to obtain a target object mask; and the segmentation model merging the target hand mask and the target object mask to obtain the target mask.

[0047] The segmentation model can be a series of Segment Anything models. In other embodiments, the segmentation model can also be other models, which can be set by those skilled in the art according to actual needs.

[0048] Further, the step of cropping the target image based on the target mask to obtain the hand-object region cropped image corresponding to the target image includes: cropping the hand region and object region in the target image based on the target mask, and using the cropped hand region and object region as the hand-object region cropped image corresponding to the target image.

[0049] Furthermore, the step of using a 3D reconstruction model to extract the hand-object interaction mesh from the cropped hand-object region image to obtain the target hand-object interaction mesh includes: using the visual encoder of the 3D reconstruction model to extract the second visual semantic features of the object region in the cropped hand-object region image; using the second visual semantic features as a condition to guide the backbone network to perform multi-step denoising in the 3D latent space, so that the latent vector gradually evolves from random noise into a second 3D representation that conforms to the structure of the hand region and the object region in the cropped hand-object region image; and using the pre-trained decoder in the 3D reconstruction model to decode the second 3D representation into the target hand-object interaction mesh.

[0050] The target hand-object interaction mesh can be represented as: The target hand-object interaction mesh can maintain the relative spatial relationship between the hand and the object well, especially with high geometric fidelity from a visible viewpoint.

[0051] S2. Extract the hand region and object region from the target hand-object interaction mesh to obtain the target hand coarse region and the target object coarse region.

[0052] Further, step S2 includes: converting the target hand-object interaction mesh into a point cloud representation; and performing point-by-point semantic classification processing on the point cloud representation using a point cloud segmentation network to obtain the coarse region of the target hand and the coarse region of the target object. The steps of obtaining the coarse regions of the target hand and the target object can determine the approximate spatial range of the hand and object within the target hand-object interaction mesh, thereby constraining the search space during subsequent registration of the hand mesh and object mesh.

[0053] The point cloud is referred to as an unordered point cloud.

[0054] The point cloud segmentation network can be a Point Transformer series model. In other embodiments, the point cloud segmentation network can also be other models, which can be set by those skilled in the art according to actual needs.

[0055] Further, the step of using a point cloud segmentation network to perform point-by-point semantic classification processing on the point cloud representation to obtain the coarse region of the target hand and the coarse region of the target object includes: the point cloud segmentation network serializes the point cloud representation using a space-filling curve to obtain a point cloud sequence; the point cloud segmentation network divides the point cloud sequence into multiple local point blocks; the point cloud segmentation network extracts the geometric and semantic features of each point within each local point block using multi-head self-attention combined with conditional position encoding to obtain the target geometric and semantic features of each point within each local point block; all target geometric and semantic features are passed across layers along with the point cloud sequence, and each network layer in the point cloud segmentation network iteratively calculates based on all the input target geometric and semantic features to generate multi-level, multi-scale features with different receptive fields; the multi-level, multi-scale features are gradually fused using a hierarchical encoding-decoding structure in the point cloud segmentation network, and in the decoding stage, the high-order semantic features in the multi-level, multi-scale features are back-transmitted to the single-point fusion features corresponding to the original point positions; based on the multi-layer perceptron (Multi-Layer) in the point cloud segmentation network... The Perceptron (MLP) classification head maps the single-point fusion features corresponding to each point location to generate a semantic label for each point. Among all semantic labels, all points belonging to the semantic label of the hand are aggregated to form a continuous coarse region of the target hand. Among all semantic labels, all points belonging to the semantic label of the object are aggregated to form a continuous coarse region of the target object.

[0056] Based on the coarse region of the target hand and the coarse region of the target object, the coarse spatial regions corresponding to the hand and the object can be located on the overall hand-object interaction grid (i.e., the target hand-object interaction grid), which can be used as the region prior for subsequent independent registration of the hand and the object.

[0057] Furthermore, after obtaining the initial segmentation results (i.e., the coarse regions of the target hand and the target object), the initial segmentation results can be optimized through the following steps: First, neighborhood consistency constraints are used to correct isolated pixels or small misclassified regions, making local labels more consistent; second, connected component filtering is performed to retain the main hand and object regions and remove small noise regions; subsequently, outlier detection is used to remove abnormal pixels scattered outside the background or boundaries; finally, local smoothing is performed on the remaining regions to make the segmentation boundaries smoother and more natural. After these operations, a smoother, more compact, and spatially consistent coarse hand and object regions can be obtained.

[0058] S3. Use an object registration algorithm to coarsely register the initial object mesh to the coarse region of the target object to obtain an optimized object mesh.

[0059] Furthermore, the calculation formula for step S3 is as follows: ; in, These are the rotation parameters of the objects in the initial object mesh; These are the translation parameters of the objects in the initial object mesh; yes The registration weights of points are used to reflect visibility priors; yes point; yes The nearest distance from the point to the initial object mesh; It is the coarse region of the target object; It is a function for finding the minimum value.

[0060] Used to represent visibility priors: when A point is assigned a higher weight when it is closer to the surface corresponding to the visible region in the target image; while a point is assigned a lower weight when it is located in an occluded region or has high geometric uncertainty. The impact.

[0061] S4. The initial hand mesh is coarsely registered to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh.

[0062] Furthermore, the calculation formula for step S4 is as follows: ; in, These are the rotation parameters of the hand in the initial hand mesh; These are the translation parameters of the hand in the initial hand mesh; It refers to the coarse area of ​​the target hand.

[0063] By implementing steps S3 and S4, the coarse registration can be improved to respond to reliable observation areas, thereby reducing the interference of local noise and incomplete reconstruction on the overall pose estimation.

[0064] After coarse registration, the hand mesh and the object mesh have been initialized to a reasonable relative spatial position. Further joint optimization is then performed, using the pose parameters of the hand and object, as well as the MANO parameters of the hand, as optimization variables. Differentiable rendering and backpropagation iterative optimization are employed, and an optimizer is used to minimize the overall objective function. The optimizer can be Adam (Adaptive Moment Estimation).

[0065] S5. A joint optimization algorithm is used to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image.

[0066] The optimized object mesh and the optimized hand mesh have reasonable structures and consistent relative poses.

[0067] Further, step S5 includes: applying the total loss function of the joint optimization algorithm to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image. Step S5 can be referenced... Figure 3 The "joint optimization" step, in which the hand-object interaction reconstruction results are referred to, can be used as a reference. Figure 3 The results show that the hand-object grid estimation is "more reasonable".

[0068] Furthermore, the formula for calculating the total loss function is as follows: ; in, It is the total loss function; These are the two-dimensional supervision loss weighting coefficients; It is a two-dimensional supervised loss term; It is the penetration penalty loss weight coefficient; It is a penalty for penetration; These are the regularization loss weight coefficients; It is the regularization loss term.

[0069] Furthermore, the two-dimensional supervised loss term The calculation formula is: ; ; ; in, It is the hand mask loss function; It is the object mask loss function; It is the two-dimensional mask obtained by optimizing the hand mesh projection; It is the hand mask corresponding to the target image; It is the two-dimensional mask obtained by optimizing the object mesh projection; It is the object mask corresponding to the target image; It is an L1 norm.

[0070] The embodiments of this application introduce two-dimensional projection consistency constraints, which can ensure that the optimized hand and object remain consistent with the observation of the original image (i.e., the target image) under the input viewpoint.

[0071] Furthermore, penetrating the penalty loss item The calculation formula is: ; in, It is a function for finding the maximum value; It is the apex of the hand The signed distance to the object's surface. When When the value is less than 0, it indicates that the vertex is inside the object, and a penalty is applied to that vertex; when... No penalty is incurred when the value is greater than or equal to 0.

[0072] The embodiments of this application introduce a penetration penalty loss term, which can avoid unreasonable geometric penetration between the hand mesh and the object mesh.

[0073] Furthermore, the regularization loss term The calculation formula is: ; in, These are the regularization weights of the MANO model parameters; These are the hand pose parameters of the optimized hand mesh; These are the hand posture parameters of the initial hand mesh; These are the shape parameters of the optimized object mesh; These are the shape parameters of the initial object mesh; These are the pose regularization weights of the hand and the object; These are the initial hand rotation parameters obtained during the coarse registration stage; These are the initial hand translation parameters obtained during the coarse registration stage; These are the initial object rotation parameters obtained during the coarse registration stage; These are the initial object translation parameters obtained during the coarse registration stage; It is the square of the L2 norm of the vector; It is the squared Frobenius norm of the matrix.

[0074] The goal of the joint optimization stage in this application embodiment is to refine the coarse registration result, rather than drastically altering the spatial positions of the hand and object. Therefore, it is necessary to constrain the offset of the optimization variables relative to the initial values ​​to avoid excessive drift during the optimization process. Specifically, on the one hand, it is necessary to limit the MANO parameters of the hand to only undergo local fine-tuning around the initial estimate; on the other hand, it is also necessary to limit the rigid body poses of the hand and object from deviating too much from the initial results obtained in the coarse registration stage. For this purpose, both can be uniformly expressed as the regularization loss term.

[0075] The core objective of the joint optimization phase is to refine the local shape of the hand, contact relationship, and image projection consistency while maintaining the overall relative stability of the hand-object relationship.

[0076] The embodiments of this application further refine the spatial relationship between the hand and the object and the shape of the hand through joint optimization, so that the final result (i.e. the hand-object interaction reconstruction result) can maintain the consistency of the overall hand-object interaction relationship and have good hand geometric rationality and image observation consistency.

[0077] Through the above implementation methods, the embodiments of this application perform coarse registration processing on the initial object mesh and the initial hand mesh to obtain optimized values ​​that conform to the overall interaction relationship, that is, optimize the object mesh and optimize the hand mesh. Then, the optimized object mesh and the optimized hand mesh are jointly optimized to obtain the hand-object interaction reconstruction result. Thus, the embodiments of this application can avoid misalignment when combining the three-dimensional results of the hand and the three-dimensional results of the object, thereby improving the quality of the hand-object interaction reconstruction result.

[0078] The core of this application's embodiments lies in employing a two-stage strategy of "coarse registration followed by joint optimization." In the coarse registration stage, the spatial priors provided by the overall hand-object interaction mesh are fully utilized. ICP (Iterative ClosestPoint) registration (i.e., steps S3 and S4) establishes a reasonable initial relative pose for the hand and object, providing a stable and high-quality initialization for subsequent optimization. Subsequently, joint optimization further refines the relative pose relationship between the hand and object, as well as the local shape of the hand, ensuring that the final reconstruction result remains consistent with image observation from the input viewpoint, reducing hand-object penetration, and improving overall geometric rationality and optimization stability.

[0079] refer to Figure 4 The diagram shown is a schematic block diagram of a monocular hand-object interactive three-dimensional reconstruction device provided in the second aspect of an embodiment of this application. Figure 4 The monocular hand-object interactive 3D reconstruction device 100 includes: The acquisition module 101 is used to acquire the initial object mesh, the initial hand mesh, and the target hand-object interaction mesh corresponding to the target image; The region extraction module 102 is used to extract the hand region and object region of the target hand-object interaction mesh to obtain the target hand coarse region and the target object coarse region. The first coarse registration module 103 is used to coarsely register the initial object mesh to the coarse region of the target object using an object registration algorithm to obtain an optimized object mesh. The second coarse registration module 104 is used to coarsely register the initial hand mesh to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh; The joint optimization module 105 is used to perform parameter adjustment processing on the optimized object mesh and the optimized hand mesh using a joint optimization algorithm to obtain the hand-object interaction reconstruction result corresponding to the target image.

[0080] A third aspect of this application provides a terminal device, the schematic diagram of which is as follows: Figure 5 As shown. The terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the terminal device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a monocular hand-object interactive 3D reconstruction method. The display screen can be a liquid crystal display screen or an e-ink display screen, and the temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.

[0081] Those skilled in the art will understand that Figure 5 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0082] In some embodiments, this application provides a terminal device, which includes a processor and a memory. The memory stores a computer program, and the processor calls and runs the computer program stored in the memory to perform the steps of the monocular hand-object interactive three-dimensional reconstruction method provided in the first aspect of this application.

[0083] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the monocular hand-object interactive three-dimensional reconstruction method provided in the first aspect of this application.

[0084] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0085] The technical features of the above embodiments can be combined without changing the basic principles of this application. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0086] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the patent protection scope of this application should be determined by the appended claims.

Claims

1. A monocular hand-object interactive 3D reconstruction method, characterized in that, include: Obtain the initial object mesh, initial hand mesh, and target hand-object interaction mesh corresponding to the target image; The target hand-object interaction mesh is processed to extract the hand region and object region to obtain the target hand coarse region and the target object coarse region; An object registration algorithm is used to coarsely register the initial object mesh to the coarse region of the target object to obtain an optimized object mesh. The initial hand mesh is coarsely registered to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh; A joint optimization algorithm is used to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image; The initial object mesh is coarsely registered to the coarse region of the target object using an object registration algorithm. The calculation formula for the optimized object mesh step is as follows: ; in, These are the rotation parameters of the objects in the initial object mesh; These are the translation parameters of the objects in the initial object mesh; yes The registration weights of points are used to reflect visibility priors; yes point; yes The nearest distance from the point to the initial object mesh; It is the coarse region of the target object; It is a function for finding the minimum value; It is the initial object mesh.

2. The monocular hand-object interactive 3D reconstruction method according to claim 1, characterized in that, The steps for extracting the hand and object regions from the target hand-object interaction mesh to obtain the coarse regions of the target hand and the target object include: The target hand-object interaction mesh is converted into a point cloud representation; A point cloud segmentation network is used to perform point-by-point semantic classification on the point cloud representation to obtain the coarse region of the target hand and the coarse region of the target object.

3. The monocular hand-object interactive 3D reconstruction method according to claim 1, characterized in that, The steps for obtaining the target hand-object interaction mesh include: Obtain the hand-object region cropped image corresponding to the target image; A three-dimensional reconstruction model is used to extract the hand-object interaction mesh from the cropped image of the hand-object region, thereby obtaining the target hand-object interaction mesh.

4. The monocular hand-object interactive 3D reconstruction method according to claim 3, characterized in that, The steps for obtaining the hand-object region cropped image corresponding to the target image include: A visual language model is used to perform object category and appearance description recognition processing on the target image to obtain the target object category and target object description. The target image, the target object category, and the target object description are input into the segmentation model to perform mask extraction processing on the interaction region between the hand and the object, thereby obtaining the target mask. The target image is cropped based on the target mask to obtain a cropped image of the hand-object region corresponding to the target image.

5. The monocular hand-object interactive three-dimensional reconstruction method according to claim 3, characterized in that, The steps for obtaining the initial object mesh include: The object region cropping image of the hand-object region is subjected to object region occlusion and completion processing to obtain the object completion image corresponding to the target image; The object mesh is extracted from the object completion image using the 3D reconstruction model to obtain the initial object mesh.

6. The monocular hand-object interactive 3D reconstruction method according to claim 1, characterized in that, The steps for obtaining the initial hand mesh include: The target image is processed by a hand pose estimation model or a hand reconstruction model to extract the hand mesh, thereby obtaining the initial hand mesh.

7. The monocular hand-object interactive 3D reconstruction method according to claim 1, characterized in that, The steps of using a joint optimization algorithm to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image include: The total loss function of the joint optimization algorithm is used to adjust the parameters of the optimized object mesh and the optimized hand mesh to obtain the hand-object interaction reconstruction result corresponding to the target image.

8. A monocular hand-object interactive three-dimensional reconstruction device, characterized in that, include: The acquisition module is used to acquire the initial object mesh, initial hand mesh, and target hand-object interaction mesh corresponding to the target image; The region extraction module is used to extract the hand region and object region of the target hand-object interaction mesh to obtain the target hand coarse region and the target object coarse region. The first coarse registration module is used to coarsely register the initial object mesh to the coarse region of the target object using an object registration algorithm to obtain an optimized object mesh. The second coarse registration module is used to coarsely register the initial hand mesh to the target hand coarse region using a hand registration algorithm to obtain an optimized hand mesh; The joint optimization module is used to perform parameter adjustment processing on the optimized object mesh and the optimized hand mesh using a joint optimization algorithm to obtain the hand-object interaction reconstruction result corresponding to the target image; The initial object mesh is coarsely registered to the coarse region of the target object using an object registration algorithm. The calculation formula for the optimized object mesh step is as follows: ; in, These are the rotation parameters of the objects in the initial object mesh; These are the translation parameters of the objects in the initial object mesh; yes The registration weights of points are used to reflect visibility priors; yes point; yes The nearest distance from the point to the initial object mesh; It is the coarse region of the target object; It is a function for finding the minimum value; It is the initial object mesh.

9. A terminal device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to perform the steps of the monocular hand-object interactive three-dimensional reconstruction method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the steps of the monocular hand-object interactive three-dimensional reconstruction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interactive two-hand three-dimensional reconstruction method and system based on single RGB image

    CN117333635A

  • Hand and hinged object 4D interaction reconstruction method based on monocular video input

    CN122024325A