Handheld object three-dimensional reconstruction method based on three-dimensional generation prior and semantic consistency

By generating 3D priors and semantic consistency methods, a 3D candidate set is generated using RGB image sequences. Combined with semantic segmentation and geometric priors, the object pose is optimized, solving the problem of inaccurate pose estimation in 3D reconstruction of handheld objects. This achieves high-quality 3D reconstruction and improves the practicality and generalization ability of the method.

CN120876780APending Publication Date: 2025-10-31ZHEJIANG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510985090.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing 3D reconstruction technologies for handheld objects, inaccurate object pose estimation leads to low reconstruction quality, and the technology relies on high-precision CAD models and datasets, making it difficult to generalize to diverse handheld objects.

Method used

By generating 3D priors and semantic consistency methods, a 3D candidate set is generated using RGB image sequences. Combined with semantic segmentation and geometric priors, the object pose is optimized, and a neural radiation field network for object reconstruction is constructed. The pose is iteratively optimized to improve accuracy.

Benefits of technology

This method achieves high-quality 3D reconstruction without the need for CAD models and dataset supervision, improves the practicality and generalization ability of the method, mitigates the impact of pose estimation errors, and significantly improves the accuracy and stability of the reconstruction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876780A_ABST
    Figure CN120876780A_ABST
Patent Text Reader

Abstract

The invention discloses a handheld object three-dimensional reconstruction method based on three-dimensional generation prior and semantic consistency, and the method comprises the steps: obtaining the three-dimensional prior from an input RGB image sequence in combination with a pre-trained multi-modal model and a three-dimensional generation model; performing semantic alignment on the three-dimensional priori and an input image based on feature similarity measurement of a visual basic model, and preliminarily estimating the pose of an object; completing initial three-dimensional reconstruction of the object according to the rough pose by using a neural radiation field; fine adjustment is carried out on the object pose based on the initial reconstruction result; and generating a high-precision three-dimensional reconstruction result through a nerve radiation field by using the optimized pose. The method has the advantages that the inherent problem of pose estimation in a hand-held object scene is solved through semantic consistency constraint; high-precision object three-dimensional reconstruction can be realized only by depending on RGB video data easy to obtain; dependence of a traditional method on complex sensor data or manual annotation is remarkably reduced, and reconstruction efficiency and practicability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and more particularly to a method for 3D reconstruction of handheld objects based on 3D generation priors and semantic consistency. Background Technology

[0002] In recent years, 3D object shape reconstruction technology has become an important research direction in the fields of computer vision and graphics, showing broad application prospects in scenarios such as virtual reality, augmented reality, robot manipulation, and human-computer interaction. Compared with static object 3D reconstruction, handheld object 3D reconstruction can dynamically rotate the object with the hand, providing richer observation angles and obtaining more detailed object information. Research on this technology will stimulate its application potential in various industries and help promote the development of various sensing and intelligent interaction systems.

[0003] Early methods for 3D reconstruction of handheld objects primarily employed model-fitting approaches. Due to the diverse categories and complex shapes of objects, it was difficult to define a universal 3D representation. Therefore, these methods typically assumed that the CAD model of the target object was known, transforming the reconstruction task into a 6-DOF pose estimation problem, and combining optimization methods to jointly estimate the hand parameters and object pose from the input image. However, in practical applications, obtaining a high-quality CAD model of the target object beforehand presents significant challenges, requiring high-precision instruments for scanning, which limits the generalization and practicality of such methods.

[0004] In recent years, the development of deep learning has driven significant progress in learning-based 3D reconstruction methods for handheld objects. These methods primarily utilize implicit representations as 3D expressions of objects, eliminating reliance on CAD models and directly learning object shape representations from images. However, these methods still depend on training datasets with 3D supervision, and their applicability is typically limited to a finite number of predefined object categories within the dataset, making it difficult to generalize to the diverse handheld objects encountered in daily life.

[0005] To address this issue, a significant breakthrough has been achieved by combining object neural radiation fields with volume rendering techniques. The input is continuous 5D coordinates (spatial position and viewing angle), and the output is the SDF value and color corresponding to the spatial position. Classical volume rendering techniques are used to project the output color and SDF value onto the image, requiring only multi-view 2D RGB images as supervisory information for 3D object reconstruction. This type of method can reconstruct objects without requiring a CAD model of the object or 3D supervisory information from a dataset. However, such methods typically require accurate object pose as a prerequisite. In practical applications, directly estimating accurate object pose from RGB image sequences remains a highly challenging problem. Previous methods often rely on traditional motion structure recovery or hand motion estimation for pose estimation, but with limited accuracy. Inaccurate pose significantly degrades the quality of 3D object reconstruction. Summary of the Invention

[0006] This invention addresses the problem of inaccurate object pose estimation in existing handheld object 3D reconstruction technologies, which leads to low reconstruction quality. It proposes a handheld object 3D reconstruction method based on 3D generation priors and semantic consistency.

[0007] The objective of this invention is achieved through the following technical solution: a method for 3D reconstruction of handheld objects based on 3D generation priors and semantic consistency, comprising:

[0008] S1. Generate a three-dimensional prior candidate set for the first frame of the input RGB image sequence, including a semantic-based three-dimensional prior candidate set and an image-generated three-dimensional candidate set.

[0009] S2. Select the three-dimensional prior with the highest cross-modal similarity to the input RGB image as the final three-dimensional prior model;

[0010] S3. Obtain the semantic segmentation results, monocular normal vectors, and cross-frame corresponding point relationships of the input image sequence as geometric priors;

[0011] S4. Render the final 3D prior model in 3D space, perform semantic alignment based on the feature similarity between the input image and the rendered image, and make a preliminary estimate of the object pose.

[0012] S5. Construct and optimize an object neural radiation field network that includes only SDF branches. The object neural radiation field network is input with the coarse pose obtained in step S4 and optimized with the geometric prior obtained in S3 to complete the initial three-dimensional reconstruction of the object.

[0013] S6: Based on the initial reconstruction results, the pose of the object is finely adjusted using the reprojection relationship of the object segmentation results through differentiable rendering;

[0014] S7: Based on the adjusted object pose, the object neural radiation field network containing color branches and SDF branches is reconstructed by combining the input RGB image. The neural radiation field network is trained until convergence, and finally a high-fidelity 3D reconstruction result is output.

[0015] On the other hand, the present invention also provides a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the handheld object 3D reconstruction method based on 3D generation prior and semantic consistency.

[0016] The beneficial effects of this invention are as follows:

[0017] 1. A method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency is proposed.

[0018] 2. The proposed framework is used to achieve accurate reconstruction of 3D object models.

[0019] 3. Overcoming the dependence of traditional methods on known CAD models, this invention can complete reconstruction by inputting only RGB image sequences, significantly improving the practicality of the method.

[0020] 4. Overcomes the problem that learning-based methods rely on the dataset and can only reconstruct a small number of object categories in the dataset, significantly improving the generalization ability of the reconstruction method on different object categories.

[0021] 5. Effectively mitigates the impact of pose estimation errors in dynamic interactive scenarios by utilizing the semantic consistency relationship between the generated 3D prior model and the image to improve the accuracy and stability of pose estimation.

[0022] 6. A strategy based on reconstruction mesh and pose iteration optimization is adopted to further improve the quality of the final 3D reconstruction results. Attached Figure Description

[0023] Figure 1 This is an architecture diagram of a handheld object 3D reconstruction method based on 3D generation prior and semantic consistency in an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram of a model for the neural radiation field of an object in an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram showing the results of comparing the framework with other methods in the 3D reconstruction of handheld objects in this embodiment of the invention;

[0026] Figure 4A schematic diagram of a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency provided in an embodiment of the present invention. Detailed Implementation

[0027] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0028] like Figure 1 As shown, in this embodiment, a method for 3D reconstruction of a handheld object based on 3D generation prior and semantic consistency is provided, including the following steps:

[0029] S1, Obtain the three-dimensional prior candidate set, specifically:

[0030] The first frame image I1 is input into the multimodal large model GPT-4V, along with the prompt "Describe the object interacting with the hand in the image," to obtain a textual semantic description of the object described in the image. This textual description is then input into the text-driven 3D generative model Genie to obtain a semantic-based 3D prior candidate set. Simultaneously, image I1 is directly input into the image-driven 3D generative model Hunyuan3D to obtain a 3D candidate set generated from the image. The two sets are then merged to form a complete 3D prior candidate set.

[0031] S2, using a retrieval method to obtain the three-dimensional prior with the highest matching degree, specifically:

[0032] Given a 3D prior candidate set and an input RGB image sequence {I} k} k=0,…,N The OpenShape multimodal model is used to calculate the relationship between each candidate 3D prior and the input image sequence {I}. k} k=0,…,N The cross-modal similarity is calculated, and the candidate 3D model with the highest similarity is selected as the final 3D prior V.

[0033] S3, obtain the semantic segmentation results, monocular normal vectors, and cross-frame corresponding point relationships of the input image sequence as geometric priors, specifically:

[0034] For the input RGB image sequence {I k} k=0,…,N Using Segment-Anything, the semantic segmentation results of the hand and object in each frame are obtained {M}. k} k=0,…,N Use StableNormal to obtain the monocular surface normal vector in each frame. Using DKM to estimate the dense correspondence between 5 adjacent frames, p j and p i These are the predicted two-dimensional coordinates of the corresponding point.

[0035] S4, using semantic consistency to initially estimate the object's pose, specifically:

[0036] The 3D prior V obtained in step S2 is rendered in 3D space with 3000 uniform viewpoints to obtain a set of rendered images. The image is cropped using the object segmentation results obtained in step S3, and then features are obtained using the visual base model DINO. The input image I is then calculated. k With rendering images Feature similarity, the specific calculation method is as follows:

[0037]

[0038] Where F DINO (.) represents the features extracted by DINO, and <.> represents the vector inner product operation. As the most suitable perspective, ξ k For the corresponding rendering pose, according to F DINO (I k )and Construct a cyclic nearest neighbor matching relationship between two feature maps, calculate the similarity between each position and its cyclic corresponding point, generate a semantic matching score map, and select the positions with the highest top-K scores as matching point pairs to obtain image I. k With rendering images The high-confidence semantic correspondence between them. Then, using ξ k Perform a backprojection operation to obtain the vertex v in the 3D prior V. k and Image I k medium two-dimensional pixel p k The correspondence between them is established, and the object poses of the 3D prior are initially optimized by minimizing the reprojection loss of semantic perception and the intersection-union ratio loss of the object segmentation results. The specific optimization objective is as follows:

[0039]

[0040] Where π(.) represents the projection operation, and DR(.) represents the differentiable rendering operation implemented based on PyTorch3D. To ensure the continuity of the pose sequence, if the rotation angle between adjacent frames exceeds a preset threshold, the pose of the previous frame is used as the initialization of the current frame.

[0041] S5, object initialization and reconstruction, specifically:

[0042] like Figure 2As shown, the object neural radiation field network includes a multi-resolution hashing module for mapping the input spatial coordinates into a high-dimensional feature representation; and a multilayer perceptron (MLP) for predicting the color value and SDF attribute associated with the spatial location based on the feature representation. In the initial reconstruction stage, the network only includes an SDF branch, excluding the color branch, and the object pose is treated as an optimizable variable. The optimization process first transforms the image from a two-dimensional pixel coordinate system to a three-dimensional camera coordinate system based on camera intrinsic parameters. For each pixel, random sampling is performed along the origin of the camera coordinate system and its ray direction. Then, based on the object pose, the sampled points are transformed to the object coordinate system. The sampled points are input into the network to obtain the corresponding SDF value. Using volume rendering technology, the segmentation result value of the predicted pixel can be obtained. normal vector and depth value Therefore, the geometric prior obtained in step S3 is used to supervise the model. The model's loss function is as follows:

[0043] L=λ mask L mask +λ eik L eik +λ normal L normal +λ corres L corres

[0044] Where λ mask , λ eik , λ normal , λ corres These are all constant coefficients used to balance the loss function. Each term of the loss function is defined as follows:

[0045]

[0046] These two loss functions are used to regularize the predicted SDF value of the sampling points, where N p This represents the number of sampling points needed to predict each pixel. M represents the gradient of the output SDF value relative to the input sampling point location. r It is the semantic segmentation result corresponding to the pixels.

[0047]

[0048]

[0049] in p is the monocular normal vector obtained in step S3. j and p i Let π(.) and π be the two-dimensional coordinates of the corresponding point predicted in step S3. -1 (.) represents projection and back projection operations, respectively. This represents the cross-frame transformation relationship of the corresponding point pose.

[0050] The training process is based on the PyTorch framework, using tiny-cuda-nn and nerfacc for network acceleration, and is executed on an RTX 4080 GPU, with a total of 5000 optimization epochs. After completion, the Marching-Cubes algorithm is used to obtain the 3D mesh of the initial reconstruction results.

[0051] S6 refines the pose of the object, specifically:

[0052] The initialized reconstructed mesh obtained in step S5 is used as the new 3D prior V. The object's pose is then finely adjusted again by minimizing the intersection-union ratio loss of the object segmentation results.

[0053]

[0054] During the optimization process, a smoothness loss is also incorporated to prevent abrupt changes in motion, which means constraining the difference in object vertex coordinates between two frames.

[0055] S7, obtain the final reconstruction result, specifically:

[0056] Similar to step S5, the object neural radiation field network is reconstructed, the difference being the addition of a color prediction branch network. The structure of the color prediction branch network is the same as the SDF branch in step S5, but the output attribute dimension is changed to color value. The optimization process is also similar to the loss function, but color supervision is added using the input two-dimensional image, specifically:

[0057] L = L color +λ mask L mask +λ eik L eik +λ normal L normal +λ corres L corres

[0058] Where L mask L eik L normal and L corres The definition remains the same as before, L color for:

[0059]

[0060] Where N r This represents the number of pixels in the input image. C rThese represent the estimated and true color values ​​of the pixels, respectively. A total of 20,000 optimization rounds are performed in this stage. After optimization, the Marching-Cubes algorithm is used to obtain high-fidelity 3D reconstruction results.

[0061] Figure 3 The reconstruction results of this method and other hand-object interaction 3D reconstruction methods are presented. This method significantly outperforms other methods in terms of reconstruction accuracy.

[0062] Corresponding to the aforementioned embodiment of a handheld object 3D reconstruction method based on 3D generation prior and semantic consistency, the present invention also provides an embodiment of a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency.

[0063] See Figure 4 The present invention provides a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a handheld object 3D reconstruction method based on 3D generation prior and semantic consistency in the above embodiment.

[0064] The present invention provides an embodiment of a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency. This embodiment can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 4 The diagram shown is a hardware structure diagram of any data processing-capable device, including a handheld object 3D reconstruction device based on 3D generation prior and semantic consistency provided by the present invention. (Except for...) Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0065] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0066] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0067] This invention also provides a computer-readable storage medium storing a program that, when executed by a processor, implements a handheld object 3D reconstruction method based on 3D generation prior and semantic consistency as described in the above embodiments.

[0068] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0069] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for 3D reconstruction of handheld objects based on 3D generation priors and semantic consistency.

[0070] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0071] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. This application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for 3D reconstruction of handheld objects based on 3D generation priors and semantic consistency, characterized in that, include: S1. Generate a three-dimensional prior candidate set for the first frame of the input RGB image sequence, including a semantic-based three-dimensional prior candidate set and an image-generated three-dimensional candidate set. S2. Select the three-dimensional prior with the highest cross-modal similarity to the input RGB image as the final three-dimensional prior model; S3. Obtain the semantic segmentation results, monocular normal vectors, and cross-frame corresponding point relationships of the input image sequence as geometric priors; S4. Render the final 3D prior model in 3D space, perform semantic alignment based on the feature similarity between the input image and the rendered image, and make a preliminary estimate of the object pose. S5. Construct and optimize an object neural radiation field network that includes only SDF branches. The object neural radiation field network is input with the coarse pose obtained in step S4 and optimized with the geometric prior obtained in S3 to complete the initial three-dimensional reconstruction of the object. S6: Based on the initial reconstruction results, the object pose is finely adjusted by using the reprojection relationship of the object segmentation results obtained by S3 through differentiable rendering. S7: Based on the adjusted object pose, the object neural radiation field network containing color branches and SDF branches is reconstructed by combining the input RGB image. The neural radiation field network is trained until convergence, and finally a high-fidelity 3D reconstruction result is output.

2. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 1, characterized in that, The specific steps for generating the three-dimensional prior candidate set are as follows: For the first frame image I1 of the input RGB image sequence, GPT-4V is used to perform text description on the image. The text description is then input into the text 3D generation model Genie to obtain the text generation 3D prior candidate set. At the same time, the first frame image I1 is input into the image 3D generation model Hunyuan3D to obtain the image generation 3D prior candidate set. The text generates a 3D prior candidate set, and the image generates a 3D candidate set, which are then used as the 3D prior candidate set.

3. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 1, characterized in that, Specifically, S2 is: Given a 3D prior candidate set and an input RGB image sequence {I} k } k=0,…,N The OpenShape multimodal model is used to calculate the relationship between each candidate 3D prior and the input image sequence {I}. k } k=0,…,N The cross-modal similarity is calculated, and the candidate 3D model with the highest similarity is selected as the final 3D prior V.

4. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 1, characterized in that, Specifically, S3 is: For the input RGB image sequence {I k } k=0,…,N Using Segment-Anything, the semantic segmentation results of the hand and object in each frame are obtained {M}. k } k=0,…,N Use StableNormal to obtain the monocular surface normal vector in each frame. DKM is used to estimate the dense correspondence between 5 adjacent frames.

5. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 1, characterized in that, Specifically, S4 is: The 3D prior V obtained in S2 is uniformly sampled in 3D space to render the viewpoint and obtain the rendered image. The object segmentation results obtained using S3 are used to crop the image, and then the visual base model DINO is used to obtain features and calculate the input image I. k With rendering images Feature similarity, the specific calculation method is as follows: Where F DINO (.) represents the features extracted by DINO, and <.> represents the vector inner product operation. As the most suitable perspective, ξ k For the corresponding rendering pose, according to F DINO (I k )and Construct a cyclic nearest neighbor matching relationship between two feature maps, calculate the similarity between each position and its cyclic corresponding point, generate a semantic matching score map, and select the positions with the highest top-K scores as matching point pairs to obtain image I. k With rendering images The high-confidence semantic correspondence between them, and then using ξ k Perform a back projection operation to obtain the vertex v in the 3D prior V. k and Image I k medium two-dimensional pixel p k The correspondence between them is determined by minimizing the reprojection loss of semantic perception and the intersection-union ratio loss of the object segmentation results obtained by S3. This allows for the initial optimization of the 3D prior object poses. The specific optimization objective is as follows: Where π(.) represents the projection operation, DR(.) represents the differentiable rendering operation, and M k This is the semantic segmentation result.

6. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 1, characterized in that, The construction of the object neural radiation field network in S5 is specifically as follows: The object neural radiation field network includes a multi-resolution hashing module for mapping input spatial coordinates into a high-dimensional feature representation; a multilayer perceptron for predicting the SDF value associated with the spatial location based on the high-dimensional feature representation; and volume rendering technology for predicting the segmentation results of pixels. normal vector and depth value Meanwhile, the object pose obtained in step S4 is used as an optimizable variable to reduce the impact of pose noise caused by motion estimation error.

7. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 6, characterized in that, The specific loss function for the S5 object neural radiation field network optimization is as follows: L=λ mask L mask +λ eik L eik +λ normal L normal +λ corres L corres Where λ mask , λ eik , λ normal , λ corres All are constant coefficients used to balance the loss function. Each term of the loss function is defined as follows: These two loss functions are used to regularize the predicted SDF value of the sampling points, where N p This represents the number of sampling points needed to predict each pixel. M represents the gradient of the output SDF value relative to the input sampling point location. r It is the semantic segmentation result corresponding to the pixels. in p is the monocular normal vector obtained in S3. j and p i Let π(.) and π be the two-dimensional coordinates of the corresponding point predicted in S3. -1 (.) represents projection and back projection operations, respectively. Represents the cross-frame transformation relationship of pose.

8. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 7, characterized in that, Specifically, S6 is: The object neural radiation field network, optimized using S5, is used to obtain an initial 3D reconstruction mesh via the Marching-Cubes algorithm. This mesh is then used as a new 3D prior, V, and the object's pose is further refined by minimizing the cross-union ratio loss of the object segmentation results. DR(.) represents a differentiable rendering operation, M k This is the semantic segmentation result.

9. The method for 3D reconstruction of handheld objects based on 3D generation prior and semantic consistency according to claim 7, characterized in that, S7 is: The object neural radiation field network was reconstructed, and a color prediction branch network was added. The structure of the color prediction branch network is the same as that of the SDF branch. The input includes the spatial location of the current sampling point and the observation viewpoint. The output is the SDF value and color value of the sampling point. The optimization process and loss were supervised by color, specifically as follows: L=L color +λ mask L mask +λ eik L eik +λ normal L normal +λ corres L corres Where N r This represents the number of pixels in the input image. C r These represent the estimated and true color values ​​of the pixels, respectively. After optimization, the Marching-Cubes algorithm is used to obtain high-fidelity 3D reconstruction results.

10. A handheld object 3D reconstruction device based on 3D generation prior and semantic consistency, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it implements a handheld object 3D reconstruction method based on 3D generation prior and semantic consistency as described in any one of claims 1-9.

Citation Information

Cited By

  • Building engineering component part normalization judgment method based on image processing

    CN122156078A