Three-dimensional reconstruction method and device, equipment, storage medium and computer program product
By determining the initial 3D mesh and skeleton point cloud in a single-view image and combining local and global scale fusion, a fully accurate scaled target 3D mesh is generated. This solves the problems of dependence on specific reference objects and incomplete reconstruction in complex scenes in single-view multi-object 3D reconstruction, and achieves higher accuracy and completeness in 3D reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing single-view multi-object 3D reconstruction methods are highly dependent on specific explicit reference objects, and the integrity and accuracy of 3D mesh model reconstruction are limited in complex multi-object scenes, especially in the ability to reconstruct details in occluded areas.
By acquiring single-view images, the initial 3D mesh and skeleton point cloud are determined. Combining multi-stage fine fusion of local and global scales, a fully accurate scaled target 3D mesh is generated using a 3D generation diffusion model, a monocular depth estimation model, and a six-degree-of-freedom pose estimation model.
It improves the accuracy of real-scale estimation and the integrity and precision of 3D mesh reconstruction in complex multi-object scenes, reduces the dependence on specific explicit references, and enhances the ability to reconstruct details in occluded areas.
Smart Images

Figure CN121767556A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] Accurate geometric reconstruction is a key long-term research problem in computer graphics and computer vision, and its application in object reconstruction can be widely used in games, film and television, education and many other fields. At present, some explorations have been made in the 3D (3-Dimensional) reconstruction of multiple objects in a single view.
[0003] Single-view multi-object 3D reconstruction methods in related technologies are highly dependent on specific, explicit physical references (hereinafter referred to as references), and require high integrity, standardization, and clarity of these references. However, in practical applications, if these references are missing, occluded, irregular in shape, have blurred boundaries, or inaccurate prior size information, the accuracy and stability of scale estimation will be greatly reduced, making it difficult to generalize to highly diverse real-world scenes. Furthermore, the generated 3D mesh model may still have geometrical omissions, deformations, or inaccuracies, especially with insufficient detail reconstruction capabilities in occluded areas. This limits the integrity and accuracy of 3D mesh model reconstruction in complex multi-object scenes. Summary of the Invention
[0004] To address the technical problems existing in related technologies, embodiments of this application provide a three-dimensional reconstruction method, apparatus, device, storage medium, and computer program product.
[0005] To achieve the above objectives, the technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a three-dimensional reconstruction method, the method comprising: Acquire a single-view image; the single-view image contains one or more target objects; Based on the single-view image, an initial 3D mesh and skeleton point cloud are determined; each sub-mesh in the initial 3D mesh corresponds to each target object. Based on the initial 3D mesh and skeleton point cloud, a first scale factor is determined; the first scale factor characterizes the local scale factor of the target. The first scale factor and the second scale factor are fused to obtain the target scale factor, and the target three-dimensional mesh is determined based on the target scale factor; the second scale factor represents the global scale factor.
[0006] Secondly, embodiments of this application also provide a three-dimensional reconstruction apparatus, the apparatus comprising: An acquisition unit is used to acquire a single-view image; the single-view image contains one or more target objects. The first determining unit is used to determine an initial three-dimensional mesh based on the single-view image; each sub-mesh in the initial three-dimensional mesh corresponds to each target object. The second determining unit is used to determine the skeleton point cloud based on the single-view image; The third determining unit is used to determine a first scale factor based on the initial three-dimensional mesh and skeleton point cloud; the first scale factor characterizes the local scale factor of the target. The fusion unit is used to fuse the first scale factor and the second scale factor to obtain the target scale factor; the second scale factor represents the global scale factor. The fourth determining unit is used to determine the target three-dimensional mesh based on the target scale factor.
[0007] Thirdly, embodiments of this application also provide a three-dimensional reconstruction device, including: a processor and a memory for storing a computer program capable of running on the processor; When the processor runs the computer program, it executes the steps of the three-dimensional reconstruction method described in the embodiments of this application.
[0008] Fourthly, embodiments of this application also provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the three-dimensional reconstruction method described in embodiments of this application.
[0009] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the three-dimensional reconstruction method described in embodiments of this application.
[0010] The 3D reconstruction method, apparatus, device, storage medium, and computer program product provided in this application embodiment acquire a single-view image; the single-view image contains one or more target objects; based on the single-view image, an initial 3D mesh and skeleton point cloud are determined; each sub-mesh in the initial 3D mesh corresponds to each target object; based on the initial 3D mesh and skeleton point cloud, a first scale factor is determined; the first scale factor represents the target local scale factor; the first scale factor and a second scale factor are fused to obtain a target scale factor, and based on the target scale factor, a target 3D mesh is determined; the second scale factor represents the global scale factor. Using the technical solution of this application embodiment, taking a single-view image containing one or more target objects as input, an initial 3D mesh and skeleton point cloud are determined. Through joint generation of multiple target objects, the spatial relationships and occlusion areas between target objects can be better understood and reconstructed. By introducing multi-stage refined fusion of local and global scales, a fully accurately scaled target 3D mesh is obtained. This avoids high dependence on specific explicit reference objects, improves the accuracy of real scale estimation, and enhances the completeness and accuracy of 3D mesh reconstruction in complex multi-object scenes. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 1 ; Figure 2 This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 2 ; Figure 3 This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 3 ; Figure 4 This is a comparative diagram of the 3D reconstruction process; Figure 5 This is a diagram illustrating the comparison of 3D reconstruction results; Figure 6 This is a schematic diagram of the composition structure of the three-dimensional reconstruction device according to an embodiment of this application; Figure 7 This is a schematic diagram of the hardware structure of the three-dimensional reconstruction device according to an embodiment of this application. Detailed Implementation
[0012] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0014] Accurate geometric reconstruction is a key long-term research problem in the fields of computer graphics and computer vision. Its application in object reconstruction can be widely used in games, film and television, education, and many other fields. Currently, some explorations have been made in 3D reconstruction and volume estimation of multiple objects in a single view.
[0015] One proposed technical solution involves a method for single-view multi-object 3D reconstruction using target segmentation and 3D mesh generation. Specifically, this solution first uses the FoodMem model (a method applicable to specific object segmentation scenarios) to segment the input image, obtaining a mask for the target object. Then, the target object mask is applied to the image and input into the Hunyuan3D model to generate a 3D mesh of the target object. However, 3D meshes generated from monocular images often lack real-world scale information. To address this issue, the aforementioned technical solution relies on reference objects of known size (i.e., reference objects) existing in the image. A global scaling factor is calculated by estimating the size of these reference objects and applied to the generated 3D mesh to calculate the volume of the target object. While simple, this method's accuracy is highly dependent on the accurate identification of reference objects, the precision of prior dimensions, and the sharpness of the reference objects in the image. If reference objects are missing, inaccurate in size, or occluded, the robustness of scale estimation is severely limited, and the completeness and accuracy of the 3D mesh reconstruction are constrained.
[0016] Related technical solution two proposes a fully automated solution aimed at generating a realistic-scale 3D mesh model of a target object using only an image and name. This solution first utilizes a Large Language Model (LLM) to convert the target object name into detailed hints, and then combines this with a Segment Anything Model (SAM) for hint-driven target segmentation, effectively handling multi-object scenes. To correct potential target classification errors during segmentation, this solution further introduces a Vision Language Model (VLM) for image-driven reclassification. Finally, using key reference objects, a perspective correction algorithm is employed to calculate the true lengths of the key reference objects and the target object, and this is used as the basis for scale estimation of the 3D mesh model. Although the scheme improves object segmentation and classification and focuses on common objects as references, its scale estimation still highly depends on the completeness and standardization of the references. When the references are occluded, have irregular shapes, or have complex perspective relationships, the accuracy of the true scale estimation may still be significantly affected. Furthermore, its applicability is limited in cases where there are no references or the references are not obvious. In addition, for complex occlusion, overlap, irregular shapes, and blurred boundaries between target objects, the generated 3D mesh model may still have geometrical defects, deformations, or inaccuracies, especially in terms of insufficient detail reconstruction capability in occluded areas.
[0017] It is evident that the single-view multi-object 3D reconstruction methods in related technologies suffer from insufficient generalization and robustness in scale estimation, as well as limited completeness and accuracy in 3D mesh model reconstruction under complex multi-object scenes.
[0018] Based on this, this application proposes a three-dimensional reconstruction method. In various embodiments of this application, a single-view image containing one or more target objects is used as input to determine an initial three-dimensional mesh and skeleton point cloud. Through joint generation of multiple target objects, the spatial relationships and occlusion areas between target objects can be better understood and reconstructed. By introducing multi-stage fine fusion of local and global scales, a fully accurate scaled target 3D mesh is obtained. In this way, the high dependence on specific explicit reference objects can be avoided, the accuracy of real scale estimation can be improved, and the integrity and accuracy of 3D mesh reconstruction in complex multi-target object scenes can be enhanced.
[0019] This application provides a three-dimensional reconstruction method, which is applied to a three-dimensional reconstruction device. Figure 1 This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 1 ,like Figure 1 As shown, the three-dimensional reconstruction method includes: Step 101: Obtain a single-view image.
[0020] In this embodiment of the application, the single-view image includes one or more target objects. These one or more target objects may also be referred to as at least one target object.
[0021] It should be noted that the single-view image here can also be called a monocular image or a single-view image. The embodiments of this application do not limit the name of the single-view image.
[0022] Here, the 3D reconstruction device can acquire a single-view image through a camera used for depth estimation, wherein the single-view image is a single input 2D (2-Dimension) image.
[0023] Step 102: Based on the single-view image, determine the initial 3D mesh and skeleton point cloud.
[0024] In this embodiment, each sub-mesh in the initial 3D mesh corresponds to a target object. The initial 3D mesh is an initial, scale-free 3D mesh, which may also be referred to as a joint 3D mesh.
[0025] Here, the 3D reconstruction device can generate an initial 3D mesh (i.e., a joint 3D mesh) containing all target objects based on a single input 2D image using a single-view multi-object 3D reconstruction module (i.e., a single-view multi-object 3D reconstruction module), and generate a skeleton point cloud for subsequent alignment.
[0026] In practical applications, the 3D reconstruction device can generate an initial 3D mesh based on multiple target objects.
[0027] Based on this, in one embodiment, determining the initial 3D mesh based on the single-view image includes: The single-view image is input into the first model to obtain the initial three-dimensional mesh output by the first model; The first model represents a three-dimensional generation diffusion model, and the initial three-dimensional mesh contains one or more target objects in the real scene.
[0028] Here, a first model is set in the single-view multi-object 3D reconstruction module. This first model is a 3D generation diffusion model, that is, a pre-trained diffusion-based 3D model. For example, the first model can be the Hunyuan3D model.
[0029] Specifically, the 3D reconstruction device inputs a single 2D image (corresponding to the aforementioned single-view image) into the 3D generative diffusion model to jointly generate an initial 3D mesh; that is, the 3D generative diffusion model can generate an initial 3D mesh containing one or more target objects in the real scene based on the input single 2D image, which can be represented as follows: This joint generation method utilizes the implicit understanding of the relative positions and occlusion relationships between target objects in a 2D image by a 3D generative diffusion model, thereby generating a geometrically more complete and relationally accurate initial 3D mesh.
[0030] Here, due to the generated initial 3D mesh Since it is a whole mesh, in order to obtain an independent 3D mesh model of an object, this embodiment of the application can also use a clustering algorithm to cluster all vertices of the initial 3D mesh, obtain clustering results, and based on the clustering results, divide the initial 3D mesh into multiple sub-mesh, wherein each sub-mesh corresponds to each target object.
[0031] Here, the clustering algorithm can be, for example, the Density Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm. Specifically, for the initial 3D mesh... All vertices are clustered using the DBSCAN algorithm to obtain clustering results. These clustering results are then used to refine the initial 3D mesh. Divided into An independent subgrid, for example ,in, Indicates the number of clusters, and It is a positive integer greater than or equal to 2. The target object can be determined by methods such as image segmentation or category recognition, with each sub-grid corresponding to a specific target object.
[0032] In one embodiment, determining the skeleton point cloud based on the single-view image includes: The single-view image is input into the second model to obtain a pseudo-depth map output by the second model; the second model represents the monocular depth estimation model. The skeleton point cloud is determined based on the pseudo-depth map.
[0033] Here, the 3D reconstruction device can utilize a monocular depth estimation (DE) model, such as the DepthPro model, to generate a pseudo-depth map based on a single input 2D image (corresponding to the aforementioned single-view image), which can be represented as: .
[0034] In practical applications, after determining the pseudo-depth map, the 3D reconstruction device can combine the camera intrinsic parameter matrix of the single-view image to determine the skeleton point cloud.
[0035] Based on this, in one embodiment, determining the skeleton point cloud based on the pseudo-depth map includes: Obtain the first matrix; the first matrix represents the camera intrinsic parameter matrix of the single-view image; The skeleton point cloud is determined based on the pseudo-depth map and the first matrix.
[0036] Here, obtaining the first matrix includes: obtaining the first matrix from the metadata of the single-view image. Specifically, the first matrix can be obtained from the Exchangeable Image File Format (EXIF) metadata of the single-view image, wherein the EXIF metadata is used to record the attribute information and shooting parameters of the single-view image. This first matrix is the camera intrinsic parameter matrix of the single-view image, which can be represented as follows: .
[0037] Here, determining the skeleton point cloud based on the pseudo-depth map and the first matrix includes: performing a back-projection operation on the pseudo-depth map and the first matrix to obtain the skeleton point cloud.
[0038] Here, the camera intrinsic matrix is combined with the single-view image. The pseudo-depth map is transformed through back-projection. Converting it into a scaffold point cloud can be represented as follows: Specifically, the skeleton point cloud can be calculated by back-projection using the following formula (1): (1); in, Represents the 3D points that make up the skeleton point cloud; Represents pixel coordinates; This represents the depth value of the corresponding pixel.
[0039] It should be noted that when obtaining 3D points using the above formula (1) Then, it can be directly based on 3D points Construct a skeleton point cloud .
[0040] Here, skeletal point clouds Although it lacks an absolute scale, it provides a 3D spatial structure that is geometrically consistent with a single input 2D image, providing an important geometric reference for subsequent skeleton point cloud mesh alignment operations.
[0041] Step 103: Determine the first scale factor based on the initial 3D mesh and skeleton point cloud.
[0042] In this embodiment of the application, the first scale factor represents the local scale factor of the target, and can be expressed as: .
[0043] In practical applications, the initial 3D mesh generated by the single-view multi-object 3D reconstruction module lacks local scale and standardized coordinates. To solve this problem, this application embodiment uses a skeleton point cloud mesh alignment module to align the initial 3D mesh with the skeleton point cloud to restore the accuracy of the local scale of the initial 3D mesh.
[0044] Based on this, in one embodiment, determining the first scale factor based on the initial 3D mesh and skeleton point cloud includes: Align the initial 3D mesh with the skeleton point cloud to obtain the aligned 3D mesh and skeleton point cloud; The first scale factor is determined based on the aligned 3D mesh and skeleton point cloud.
[0045] In practical applications, in one embodiment, aligning the initial 3D mesh with the skeleton point cloud to obtain the aligned 3D mesh and skeleton point cloud includes: The initial 3D mesh and the skeleton point cloud are input into the third model, and the initial 3D mesh and the skeleton point cloud are aligned using the third model to obtain the aligned 3D mesh and skeleton point cloud. The third model represents a six-degree-of-freedom attitude estimation model.
[0046] Here, a third model is set in the skeleton point cloud mesh alignment module. This third model is a six-degrees-of-freedom (6D) pose estimation model. That is, this third model is a pre-trained basic model that can perform 6D pose estimation. The 6D pose includes rotation, translation, and scale.
[0047] Specifically, the initial 3D mesh and skeleton point cloud are input into the 6D pose estimation model, and the 6D pose estimation model is used to perform pose estimation on the initial 3D mesh (which contains multiple independent sub-meshes). , The value ranges from 1 to ) and skeleton point cloud (containing each individual submesh) The corresponding subset of point cloud in the skeleton point cloud , The value ranges from 1 to Alignment is performed to obtain the aligned 3D mesh and skeleton point cloud. It should be noted that this 6D pose estimation model may also output a more accurate scale factor, namely the fourth scale factor. .
[0048] In practical applications, the geometric distance between the aligned 3D mesh and the skeleton point cloud can be optimized using a differentiable loss function. Through this optimization process, the optimal chamfer scale factor of each sub-mesh can be obtained, thereby determining the target local scale factor.
[0049] Based on this, in one embodiment, determining the first scale factor based on the aligned 3D mesh and skeleton point cloud includes: Construct the loss function; The geometric distance between the aligned 3D mesh and the skeleton point cloud is minimized using the loss function to obtain the third scale factor; the third scale factor represents the optimal chamfer scale factor. The first scale factor is determined based on the third scale factor.
[0050] Here, the 3D reconstruction device defines a differentiable loss function and uses this loss function to minimize the geometric distance between the aligned 3D mesh and the skeleton point cloud. Its optimization objective is to find the optimal chamfer scale factor (denoted as...). And optimal attitude parameters, so that the submesh Transformed and a subset of point clouds The geometric distance between them is the closest. The optimal pose parameters include the rotation matrix (represented as...). Translation vector (represented as) ).
[0051] Here, the point cloud chamfer distance (also known as Chamfer distance) can be used as the loss function, which can be expressed as: Chamfer distance is a method for measuring the distance between two subsets of a point cloud; that is, it is used to quantify the difference between two subsets of a point cloud.
[0052] Specifically, the third scale factor (i.e., the optimal chamfer scale factor) can be determined using the following formula (2): (2); in, This represents multiple independent sub-mesh in the initial 3D mesh; This represents the optimal chamfer scale factor to be determined; and Indicates the optimal attitude parameters; Represents each individual subgrid The corresponding subset of point cloud in the skeleton point cloud; Indicates volume; Represents the loss function; This represents the local scale factor of the target.
[0053] It should be noted that, through the optimization process of the above formula (2), in addition to obtaining the optimal chamfer scale factor for each subgrid, It can also obtain a standardized aligned 3D mesh (represented as...) ).
[0054] Here, in one embodiment, determining the first scale factor based on the third scale factor includes: determining a fourth scale factor and a fifth scale factor; and determining the first scale factor based on the product of the third scale factor, the fourth scale factor, and the fifth scale factor.
[0055] Specifically, the first scale factor (i.e., the target local scale factor) can be calculated using the following formula (3): (3); in, Indicates the first scale factor; Indicates the fifth scale factor; Indicates the fourth scale factor; This represents the third scale factor.
[0056] In practical applications, in one embodiment, before determining the first scale factor based on the third scale factor, the method further includes: determining a fifth scale factor; the fifth scale factor characterizing an initial coarse scale factor.
[0057] Here, determining the fifth scale factor includes: determining a first oriented bounding box and a second oriented bounding box; the first oriented bounding box characterizes each sub-mesh. The corresponding oriented bounding box, the second oriented bounding box represents a subset of the point cloud. The corresponding oriented bounding box; based on the first oriented bounding box and the second oriented bounding box, the fifth scale factor is determined.
[0058] It should be noted that the Oriented Bounding Box (OBB) can also be called an Oriented Bounding Box or an Oriented Bounding Frame. The name of the Oriented Bounding Box is not limited in the embodiments of this application.
[0059] Using the cloud subset For example, Principal Component Analysis (PCA) can be used to construct a bounding box based on the principal axis direction of the point cloud, making the direction of the bounding box closer to the geometry of the object, thereby improving space utilization efficiency. Specifically, firstly, PCA is used to obtain the three principal directions of the point cloud, obtain the centroid, calculate the covariance, obtain the covariance matrix, and find the eigenvalues and eigenvectors of the covariance matrix, where the eigenvectors are the principal directions; then, using the principal directions and centroid, the input point cloud is transformed to the origin, with the principal directions coinciding with the coordinate system direction, establishing the bounding box of the point cloud transformed to the origin; finally, the principal directions and bounding box are set for the input point cloud, achieved through the inverse transformation from the input point cloud to the origin point cloud.
[0060] Here, after obtaining the first and second oriented bounding boxes, the fifth scale factor (i.e., the initial coarse scale factor) is calculated by comparing the lengths of the principal axes of the two oriented bounding boxes (which can be simply referred to as the principal axis lengths). Specifically, the fifth scale factor can be calculated using the following formula (4): (4); in, This represents the initial coarse scaling factor; This indicates the main axis length of the oriented bounding box.
[0061] Step 104: Fuse the first scale factor and the second scale factor to obtain the target scale factor, and determine the target three-dimensional mesh based on the target scale factor.
[0062] In this embodiment of the application, the second scale factor represents the global scale factor, and can be expressed as: .
[0063] Here, in one embodiment, before fusing the first scale factor and the second scale factor to obtain the target scale factor, the method further includes: determining the second scale factor.
[0064] In practical applications, 3D reconstruction devices can estimate the global true scale of the entire real scene by combining external large-scale knowledge bases and implicit size cues in images.
[0065] Based on this, in one embodiment, determining the second scale factor includes: Determine the first size information and the second size information; the first size information represents the true size information of the reference image, and the second size information represents the oriented bounding box size information of the normalized aligned 3D mesh; The second scale factor is determined based on the first size information and the second size information.
[0066] Here, the external large-scale knowledge base may also be referred to as an external knowledge base or a reference database. This application embodiment does not limit the name of the external large-scale knowledge base.
[0067] Here, the reference database contains multiple reference images and multiple real-world size information corresponding to the multiple reference images. For example, the reference database could be Amazon's product image database.
[0068] In one embodiment, determining the first size information includes one of the following: When the image features of the target object meet the first condition, the first size information is determined based on the real size information contained in an external reference database; When the target object image features do not meet the first condition, the first size information is determined based on the average size of the reference objects in the single-view image; The first condition includes the successful matching of the target object image features with the reference image features in the reference database.
[0069] Here, for each segmented and cropped target object image Using a visual language model (such as the Contrastive Language-Image Pre-training (CLIP) model), the target object image is processed. The feature embedding space is matched with a reference database containing a large number of reference images with real size information. If the target object image The features of the target object image were successfully matched with the features of the reference image in the reference database (i.e., a match was found between the target object image and the reference image). If the reference image is the most similar in features, then the true size information of the reference image with the most similar features is retrieved from an external reference database, and the retrieved true size information is determined as the first size information. This first size information can be represented as... If the target object image If the features of a single-view image do not match the features of a reference image in the reference database (i.e., no exact matching reference image can be found in the reference database), then the average size of common reference objects in the single-view image can be used as the first size information. .
[0070] It should be noted that the aforementioned PCA technique can be used to calculate the normalized and aligned 3D mesh. The oriented bounding box size information, i.e., the second size information, can be represented as: .
[0071] In one embodiment, determining the second scale factor based on the first size information and the second size information includes: The second scale factor is determined based on the ratio of the first size information to the second size information.
[0072] Specifically, the global scale factor can be determined using the following formula (5): (5); in, This represents the global scaling factor (i.e., the second scaling factor). This represents the actual size information of the reference image (i.e., the first size information); The oriented bounding box size information (i.e., the second size information) represents the oriented bounding box size information of the standardized aligned 3D mesh.
[0073] It should be noted that this global scale factor can map the size units of the scaleless initial 3D mesh to real-world size units (e.g., centimeters).
[0074] Here, the target scale factor can be understood as the final real-world scale factor. For each target object in a single-view image, its final real-world scale factor... Based on the target local scale factor and global scale factor It was obtained through fusion.
[0075] Here, in one embodiment, fusing the first scale factor and the second scale factor to obtain the target scale factor includes: determining the target scale factor based on the product of the first scale factor and the second scale factor.
[0076] Specifically, the target scale factor can be determined using the following formula (6): (6); in, Indicates the target scale factor; This represents the target's local scale factor (i.e., the first scale factor). This represents the global scale factor (i.e., the second scale factor).
[0077] In one embodiment, the method further includes: determining a standardized aligned three-dimensional mesh.
[0078] Here, the normalized and aligned 3D mesh can be represented as Specifically, the geometric distance between the aligned 3D mesh and the skeleton point cloud can be optimized using a differentiable loss function, and through this optimization process, a standardized aligned 3D mesh can be obtained.
[0079] In one embodiment, determining the target 3D mesh based on the target scale factor includes: The target scale factor is applied to the standardized and aligned 3D mesh to obtain the target 3D mesh.
[0080] Here, the target 3D mesh can be understood as the final precisely scaled 3D mesh. Specifically, the target 3D mesh can be calculated using the following formula (7): (7); in, Represents the target's three-dimensional mesh; Indicates the target scale factor; Represents a standardized, aligned 3D mesh.
[0081] In one embodiment, after determining the target 3D mesh based on the target scale factor, the method further includes: Based on the target 3D mesh, the target volume is determined; the target volume represents the actual volume of the target object.
[0082] Here, the 3D reconstruction device determines the target 3D mesh. Then, standard 3D geometric calculation methods can be used to directly extract the data from the target 3D mesh. The volume enclosed by it is calculated, which is the target volume; among them, standard 3D geometric calculation methods, such as 3D geometric calculation methods based on divergence theorem or tetrahedral decomposition, can be used.
[0083] Specifically, the target volume can be calculated using the following formula (8): (8); in, Indicates the target volume.
[0084] It should be noted that this target volume can be understood as the actual volume of the target object, that is, the volume of the target object in the real world. Therefore, this target volume can also be called the real-world volume. This target volume can be used for downstream related quantitative analysis.
[0085] This application also provides another three-dimensional reconstruction method, which is applied to a three-dimensional reconstruction device. Figure 2This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 2 ,like Figure 2 As shown, the three-dimensional reconstruction method includes: Step 201: Obtain a single-view image.
[0086] In this embodiment of the application, the single-view image contains one or more target objects.
[0087] Step 202: Based on the single-view image, determine the initial 3D mesh and skeleton point cloud.
[0088] In this embodiment of the application, each sub-mesh in the initial three-dimensional mesh corresponds to each target object.
[0089] In one embodiment, determining the initial 3D mesh based on the single-view image includes: The single-view image is input into the first model to obtain the initial three-dimensional mesh output by the first model; The first model represents a three-dimensional generation diffusion model, and the initial three-dimensional mesh contains one or more target objects in the real scene.
[0090] In one embodiment, determining the skeleton point cloud based on the single-view image includes: The single-view image is input into the second model to obtain a pseudo-depth map output by the second model; the second model represents the monocular depth estimation model. The skeleton point cloud is determined based on the pseudo-depth map.
[0091] In one embodiment, determining the skeleton point cloud based on the pseudo-depth map includes: Obtain the first matrix; the first matrix represents the camera intrinsic parameter matrix of the single-view image; The skeleton point cloud is determined based on the pseudo-depth map and the first matrix.
[0092] Step 203: Align the initial 3D mesh with the skeleton point cloud to obtain the aligned 3D mesh and skeleton point cloud.
[0093] In one embodiment, aligning the initial 3D mesh with the skeleton point cloud to obtain an aligned 3D mesh and skeleton point cloud includes: The initial 3D mesh and the skeleton point cloud are input into the third model, and the initial 3D mesh and the skeleton point cloud are aligned using the third model to obtain the aligned 3D mesh and skeleton point cloud. The third model represents a six-degree-of-freedom attitude estimation model.
[0094] Step 204: Determine the first scale factor based on the aligned 3D mesh and skeleton point cloud.
[0095] In this embodiment of the application, the first scale factor represents the local scale factor of the target.
[0096] In one embodiment, determining the first scale factor based on the aligned 3D mesh and skeleton point cloud includes: Construct the loss function; The geometric distance between the aligned 3D mesh and the skeleton point cloud is minimized using the loss function to obtain the third scale factor; the third scale factor represents the optimal chamfer scale factor. The first scale factor is determined based on the third scale factor.
[0097] Step 205: Determine the second scale factor.
[0098] In this embodiment of the application, the second scale factor represents the global scale factor.
[0099] In one embodiment, determining the second scale factor includes: Determine the first size information and the second size information; the first size information represents the true size information of the reference image, and the second size information represents the oriented bounding box size information of the normalized aligned 3D mesh; The second scale factor is determined based on the first size information and the second size information.
[0100] In one embodiment, determining the first size information includes one of the following: When the image features of the target object meet the first condition, the first size information is determined based on the real size information contained in an external reference database; When the target object image features do not meet the first condition, the first size information is determined based on the average size of the reference objects in the single-view image; The first condition includes the successful matching of the target object image features with the reference image features in the reference database.
[0101] In one embodiment, determining the second scale factor based on the first size information and the second size information includes: The second scale factor is determined based on the ratio of the first size information to the second size information.
[0102] It should be noted that the execution order of the first scale factor determination step and the second scale factor determination step is not limited in the embodiments of this application. That is, the three-dimensional reconstruction device can execute step 203 and step 204 first and then execute step 205; or it can execute step 205 first and then execute step 203 and step 204; or it can execute step 203, step 204, and step 205 at the same time.
[0103] Step 206: Fuse the first scale factor and the second scale factor to obtain the target scale factor.
[0104] Step 207: Based on the target scale factor, determine the target three-dimensional mesh, and based on the target three-dimensional mesh, determine the target volume.
[0105] In this embodiment of the application, the target volume represents the actual volume of the target object.
[0106] In one embodiment, the method further includes: Determine the standardized and aligned 3D mesh; Determining the target 3D mesh based on the target scale factor includes: The target scale factor is applied to the standardized and aligned 3D mesh to obtain the target 3D mesh.
[0107] It should be noted that the specific processing steps of the 3D reconstruction device to complete the 3D reconstruction have been described in detail above, and will not be repeated here.
[0108] The technical solution of this application takes a single-view image containing one or more target objects as input to determine the initial 3D mesh and skeleton point cloud. Through joint generation of multiple target objects, it is possible to better understand and reconstruct the spatial relationship and occlusion area between target objects. By introducing multi-stage fine fusion of local and global scales, a fully accurate scaled target 3D mesh is obtained. In this way, it is possible to avoid high dependence on specific explicit reference objects, improve the accuracy of real scale estimation, and improve the integrity and accuracy of 3D mesh reconstruction in complex multi-target object scenes.
[0109] The present application will be described below with reference to application examples.
[0110] The 3D reconstruction methods in related technologies have the following problems in terms of 3D reconstruction and real-scale estimation: 1. Insufficient generalization and robustness of scale estimation: Related technologies are highly dependent on specific, explicit reference objects, and have high requirements for the completeness, standardization and clarity of the reference objects. Once these reference objects are missing, occluded, non-standard or their prior size information is inaccurate, the accuracy and stability of scale estimation will be greatly reduced, making it difficult to generalize to highly diverse real-world scenarios.
[0111] 2. Limited completeness and accuracy of 3D mesh model reconstruction in complex multi-object scenes: Although some related technologies attempt to handle multi-object scenes, the generated 3D mesh model may still have geometrical defects, deformations or inaccuracies due to complex occlusion, overlap, irregular shapes and blurred boundaries between target objects, especially the insufficient ability to reconstruct details in occluded areas.
[0112] 3. The overall volume estimation accuracy still needs to be improved to meet the high requirements of applications: The volume estimation schemes in related technologies perform reasonably well on some simple target objects, but for target objects with complex shapes, blurred boundaries or severe occlusion, the error of volume estimation will increase significantly, making it difficult to meet the stringent requirements of data accuracy for applications such as high-precision metrology.
[0113] This application aims to address the shortcomings of existing technologies in single-view multi-object 3D reconstruction, including insufficient accuracy, poor robustness, and inadequate handling of complex scenes. To resolve these issues, this application introduces a multi-stage, refined processing workflow, combining local and global scale estimation, and fully utilizing multi-source information (including 3D geometric cues and large-scale external knowledge bases). This approach aims to overcome the high dependence of traditional methods on specific explicit reference objects, significantly improving the completeness of 3D reconstruction and the accuracy of true scale estimation in complex multi-object scenes, thereby enabling more reliable applications such as object volume estimation and related quantization.
[0114] This application provides a scale-aware single-view multi-object 3D reconstruction (SVMOR) and true scale estimation method. This scheme is implemented through a phased, multi-module collaborative pipeline, which aims to overcome the shortcomings of related technologies in terms of scale accuracy, reconstruction integrity and generalization ability.
[0115] 1. Overall Process Description of the Solution Figure 3 This is a flowchart illustrating the three-dimensional reconstruction method according to an embodiment of this application. Figure 3 ,like Figure 3As shown, the overall implementation architecture of this scheme consists of three core modules: a single-view multi-object 3D reconstruction module, a skeleton point cloud mesh alignment module, and a metric depth refinement module. This method takes a monocular image containing one or more target objects (corresponding to the aforementioned single-view image) as input. First, the single-view multi-object 3D reconstruction module generates an initial, scale-free 3D mesh (corresponding to the aforementioned initial 3D mesh) and a skeleton point cloud. Then, the skeleton point cloud mesh alignment module utilizes this generated geometric information to recover the local scale of each target object (corresponding to the aforementioned target local scale factor) through refined multi-stage alignment and optimization, and aligns it with the skeleton point cloud. Next, the metric depth refinement module calculates the global scale of the entire real scene (corresponding to the aforementioned global scale factor) by combining an external large-scale knowledge base and implicit size cues in the image. Finally, the local scale and global scale are fused to obtain a fully accurately scaled 3D mesh model (corresponding to the aforementioned target 3D mesh), and the real-world volume of the target object (corresponding to the aforementioned target volume) is directly calculated from this 3D mesh model.
[0116] 2. Functional Description of Each Module 2.1 Single-view multi-object 3D reconstruction module This module is designed to generate a joint 3D mesh containing all target objects from a single input 2D image and to generate a skeleton point cloud for subsequent alignment.
[0117] 1) Joint generation of multiple objects A pre-trained, diffusion-based 3D model (such as the Hunyuan 3D model, corresponding to the first model mentioned above) is used as the core generator. This model can generate a 3D mesh containing all target objects in the real scene from a single input 2D image in one go. (Corresponding to the aforementioned initial three-dimensional mesh).
[0118] This joint generation method leverages the model's implicit understanding of the relative positions and occlusion relationships between objects in a 2D image, thereby generating an initial 3D mesh that is geometrically more complete and has more accurate relationships.
[0119] 2) Object separation This module generates It is a monolithic mesh; to obtain individual 3D sub-mesh models of objects, this module... The DBSCAN clustering algorithm is applied to all vertices.
[0120] Number of clusters This can be determined through methods such as image segmentation or category recognition. The clustering results will... Divided into An independent 3D submesh Each subgrid corresponds to a target object.
[0121] 3) Skeleton point cloud generation Using a monocular depth estimation model (e.g., the DepthPro model), based on the input single-view image Generate a pseudo depth map .
[0122] Camera intrinsic parameter matrix combined with image (Corresponding to the first matrix mentioned above, which can be obtained from the EXIF metadata of the image), the pseudo-depth map is converted into a 3D skeleton point cloud using the back-projection formula. .
[0123] The back projection formula is: ,in, These are the 3D points that make up the skeleton point cloud. These are pixel coordinates; It is the depth value of the corresponding pixel.
[0124] It should be noted that this skeleton point cloud While lacking an absolute scale, it provides a 3D spatial structure consistent with the geometry of 2D images, offering an important geometric reference for subsequent skeleton point cloud mesh alignment.
[0125] 2.2 Skeleton Point Cloud Mesh Alignment Module This module aims to address the issue of lack of local scale and standardized coordinates in the 3D mesh generated by the single-view multi-object 3D reconstruction module. By aligning the 3D mesh with the skeleton point cloud, it restores the accuracy of its local scale.
[0126] 1) Coarse scale estimation For each individual 3D submesh and its corresponding point cloud subset in the skeleton point cloud. Calculate their respective oriented bounding boxes (OBBs).
[0127] An initial coarse scale factor is calculated by comparing the lengths of the main axes of the OBB. (Corresponding to the aforementioned fifth scale factor): .
[0128] 2) Scale-aware 6D pose estimation Using a pre-trained base model capable of 6D pose estimation (including rotation, translation, and scaling) (corresponding to the third model mentioned above), for each individual 3D sub-mesh... and its corresponding point cloud subset in the skeleton point cloud. Alignment is performed. The model will then output a more accurate scaling factor. (corresponding to the aforementioned fourth scale factor) and alignment pose (i.e., the aligned 3D mesh and skeleton point cloud) are used as the initial values for optimization.
[0129] 3) Differentiable scale refinement This is the core optimization step of this module. A differentiable loss function is defined to minimize the geometric distance between the aligned 3D mesh and the skeleton point cloud. This application uses the Chamfer distance as the loss function. .
[0130] Its optimization objective is to find the optimal chamfer scale factor. Rotation matrix Translation vector , making After transformation and The geometric distance between them is the closest.
[0131] Here, the optimized formula is: (2); in, It is the original 3D sub-mesh. The optimal chamfer scale factor is to be determined. and These are attitude parameters. It is the corresponding point cloud subset. It refers to volume. Through this optimization process, the optimal chamfer scale factor for each subgrid can be obtained. 3D mesh aligned with standardization .
[0132] The final local scale factor can be expressed by the following formula (3): (3); in, It is the final local scale factor.
[0133] 2.3 Measurement Depth Refinement Module This module aims to address the problem of inaccurate pseudo-depth information generated by single-view multi-object 3D reconstruction modules in the real world. It estimates the global real scale of the entire real scene by combining external knowledge bases.
[0134] 1) External knowledge base matching For each segmented and cropped target object image This involves using a visual language model (e.g., CLIP) to match its feature embedding space with a reference database (e.g., Amazon's product image database) containing a large number of reference images with real-world size information. The goal is to find... The reference image with the most similar features is selected, and its real-world size information is retrieved, denoted as . .
[0135] 2) Global scale calculation Once a matching reference size is found This module uses this information to calculate the global scale factor. .
[0136] First, calculate the aligned 3D mesh obtained in the single-view multi-object 3D reconstruction module or the skeleton point cloud mesh alignment module. The size of the oriented bounding box is denoted as Then, calculate the global scale factor. for: This global scale factor maps the size units of the scale-free 3D mesh to real-world size units (e.g., centimeters).
[0137] Implicit cue supplementation: When an exact match cannot be found in an external database, this module can also degenerate into using the average size of common reference objects in the image as a clue. The estimates are used to provide backup plans and improve the robustness of the system.
[0138] 2.4 Final Scale Fusion and Volume Calculation 1) Final scale fusion For each target object, its final real-world scale factor By local scale factor and global scale factor It is obtained by fusion.
[0139] The fusion formula is as follows: .
[0140] Apply this final scale factor to the normalized aligned 3D mesh This yields the final, precisely scaled 3D mesh. The implementation method is as follows: .
[0141] 2) Real-world volume calculation Using standard 3D geometry calculation methods (such as 3D geometry calculation methods based on divergence theorem or tetrahedral decomposition) directly from Calculate the volume it encloses. (Corresponding to the aforementioned target volume): .
[0142] This volume value represents the real-world volume of the target object and can be used for downstream quantitative analysis.
[0143] The process of 3D reconstruction of a single-view image using 3D reconstruction methods in related technologies and the 3D reconstruction method of this application is compared as follows: Figure 4 As shown, different 3D reconstruction processes yield the following results: Figure 5 The different final 3D reconstruction results shown.
[0144] Key points of this application include: 1) Phased collaborative 3D reconstruction and scale estimation pipeline: An integrated pipeline is proposed, which includes a single-view multi-object 3D reconstruction module, a skeleton point cloud mesh alignment module, and a metric depth refinement module, to ensure the collaborative contribution of each module in 3D mesh model generation, local scale accurate correction, and global scale recovery.
[0145] 2) Monocular multi-target object joint generation and separation mechanism based on basic 3D model: A coarse 3D mesh of multiple target objects is generated from a single image at one time using a pre-trained basic 3D model (such as the Hunyuan 3D model), which effectively captures the relative pose and spatial relationship between target objects; and the DBSCAN clustering algorithm is used to separate the jointly generated mesh to obtain independent target object 3D models.
[0146] 3) Global scale recovery based on external knowledge and implicit size cues: A scale-based monocular depth estimation model is introduced. By matching the cropped target object in the image with CLIP features from a large-scale external reference database (e.g., using the CLIP model and the Amazon product image database) and fusing implicit size cues of common reference objects in the real world, the inherent limitations of monocular depth estimation are overcome, and the global scale of the image is accurately estimated.
[0147] 4) Multi-stage refined local scale correction: By combining coarse scale estimation of oriented bounding boxes, scale-aware 6D pose estimation based on basic pose, and differentiable scale refinement optimized by Chamfer distance loss, the local scale information of each target object's 3D mesh is gradually recovered and refined, and standardized coordinate alignment is performed.
[0148] Compared with the solutions of related technologies, the solution of this application has the following advantages: 1) Significantly improves the robustness and accuracy of scale estimation: Related techniques rely excessively on explicit and standard references, while this application introduces multi-stage fine-grained fusion of local and global scales, especially by using implicit size cues from external knowledge bases (such as Amazon databases) and common references for global scale correction, which greatly reduces the dependence on specific, perfect references. Thus, it can still achieve high-accuracy real-world scale estimation in a wider range of more complex real-world scenarios (including situations where there are no references, references are occluded, or are non-standard).
[0149] 2) More complete and accurate 3D reconstruction of multi-object objects: This application can better understand and reconstruct the spatial relationship and occlusion area between target objects by jointly generating multiple target objects. Combined with the fine alignment of point cloud mesh, it effectively reduces the common problems of missing 3D models, geometric distortion or boundary blurring in related technologies, so that high-quality 3D meshes can be obtained even in complex and overlapping multi-object scenes.
[0150] 3) Significantly improved volume estimation accuracy for complex target objects: Related technologies have large volume estimation errors for target objects with complex shapes, blurred boundaries, or severe occlusion. However, this application optimizes the accuracy and integrity of 3D mesh by combining multi-source information (such as multi-stage scale correction and external databases), thereby providing more accurate real-world volume estimation for various target objects and meeting the needs of applications such as high-precision quantization.
[0151] 4) Greater automation and adaptability: This application reduces the strict requirements on specific reference objects in the image and the reliance on manual dimensioning. By automatically utilizing implicit cues and big data matching for scale correction, it improves the automation level of the entire system and its adaptability to unknown scenarios, making it more feasible and convenient in practical applications.
[0152] To implement the three-dimensional reconstruction method of this application embodiment, this application embodiment also provides a three-dimensional reconstruction apparatus. Figure 6 This is a schematic diagram of the composition structure of the three-dimensional reconstruction device according to an embodiment of this application, such as... Figure 6 As shown, the three-dimensional reconstruction device includes: Acquisition unit 61 is used to acquire a single-view image; the single-view image contains one or more target objects; The first determining unit 62 is used to determine an initial three-dimensional mesh based on the single-view image; each sub-mesh in the initial three-dimensional mesh corresponds to each target object. The second determining unit 63 is used to determine the skeleton point cloud based on the single-view image; The third determining unit 64 is used to determine a first scale factor based on the initial three-dimensional mesh and skeleton point cloud; the first scale factor characterizes the local scale factor of the target. The fusion unit 65 is used to fuse the first scale factor and the second scale factor to obtain the target scale factor; the second scale factor represents the global scale factor. The fourth determining unit 66 is used to determine the target three-dimensional mesh based on the target scale factor.
[0153] In one embodiment, the first determining unit 62 is specifically used for: The single-view image is input into the first model to obtain the initial three-dimensional mesh output by the first model; The first model represents a three-dimensional generation diffusion model, and the initial three-dimensional mesh contains one or more target objects in the real scene.
[0154] In one embodiment, the second determining unit 63 includes: a first determining subunit and a second determining subunit; wherein, The first determining subunit is used to input the single-view image into the second model to obtain a pseudo-depth map output by the second model; the second model represents a monocular depth estimation model. The second determining subunit is used to determine the skeleton point cloud based on the pseudo-depth map.
[0155] In one embodiment, the second determining subunit is specifically used for: Obtain the first matrix; the first matrix represents the camera intrinsic parameter matrix of the single-view image; The skeleton point cloud is determined based on the pseudo-depth map and the first matrix.
[0156] In one embodiment, the third determining unit 64 includes: an alignment unit and a third determining subunit; wherein, The alignment unit is used to align the initial 3D mesh with the skeleton point cloud to obtain an aligned 3D mesh and skeleton point cloud. The third determining subunit is used to determine the first scale factor based on the aligned 3D mesh and skeleton point cloud.
[0157] In one embodiment, the alignment unit is specifically used for: The initial 3D mesh and the skeleton point cloud are input into the third model, and the initial 3D mesh and the skeleton point cloud are aligned using the third model to obtain the aligned 3D mesh and skeleton point cloud. The third model represents a six-degree-of-freedom attitude estimation model.
[0158] In one embodiment, the third determining subunit is specifically used for: Construct the loss function; The geometric distance between the aligned 3D mesh and the skeleton point cloud is minimized using the loss function to obtain the third scale factor; the third scale factor represents the optimal chamfer scale factor. The first scale factor is determined based on the third scale factor.
[0159] In one embodiment, the three-dimensional reconstruction apparatus further includes: a fifth determining unit; wherein, The fifth determining unit is used to determine the second scale factor.
[0160] In one embodiment, the fifth determining unit includes: a fourth determining subunit, a fifth determining subunit, and a sixth determining subunit; wherein, The fourth determining subunit is used to determine the first size information; the first size information represents the true size information of the reference image; The fifth determining subunit is used to determine the second size information; the second size information represents the oriented bounding box size information of the standardized aligned three-dimensional mesh; The sixth determining subunit is used to determine the second scale factor based on the first size information and the second size information.
[0161] In one embodiment, the fourth determining subunit is specifically used for one of the following: When the image features of the target object meet the first condition, the first size information is determined based on the real size information contained in an external reference database; When the target object image features do not meet the first condition, the first size information is determined based on the average size of the reference objects in the single-view image; The first condition includes the successful matching of the target object image features with the reference image features in the reference database.
[0162] In one embodiment, the sixth determining subunit is specifically used for: The second scale factor is determined based on the ratio of the first size information to the second size information.
[0163] In one embodiment, the three-dimensional reconstruction apparatus further includes: a sixth determining unit; wherein, The sixth determining unit is used to determine the standardized aligned three-dimensional mesh.
[0164] In one embodiment, the fourth determining unit 66 is specifically used for: The target scale factor is applied to the standardized and aligned 3D mesh to obtain the target 3D mesh.
[0165] In one embodiment, the three-dimensional reconstruction apparatus further includes: a seventh determining unit; wherein, The seventh determining unit is used to determine the target volume based on the target three-dimensional mesh; the target volume represents the actual volume of the target object.
[0166] In practical applications, the acquisition unit 61 can be implemented by the communication interface in the three-dimensional reconstruction device, and the first determination unit 62, the second determination unit 63, the third determination unit 64, the fusion unit 65 and the fourth determination unit 66 can be implemented by the processor in the three-dimensional reconstruction device.
[0167] It should be noted that the 3D reconstruction device provided in the above embodiments is only illustrated by the division of the above-described program modules. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the 3D reconstruction device and the 3D reconstruction method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the 3D reconstruction method embodiments, which will not be repeated here.
[0168] Based on the hardware implementation of the above program modules, and in order to implement the three-dimensional reconstruction method of this application embodiment, this application embodiment also provides a three-dimensional reconstruction device. Figure 7 This is a schematic diagram of the hardware composition structure of the three-dimensional reconstruction device according to an embodiment of this application, such as... Figure 7 As shown, the three-dimensional reconstruction device 70 includes: The communication interface 71 enables information exchange with other devices; The processor 72 is connected to the communication interface 71 to enable information interaction with other devices. When running a computer program, it executes the three-dimensional reconstruction method provided above, and the computer program is stored in the memory 73.
[0169] It should be noted that the specific processing procedures of communication interface 71 and processor 72 can be understood by referring to the above-mentioned three-dimensional reconstruction method.
[0170] Of course, in practical applications, the various components in the 3D reconstruction device 70 are coupled together via a bus system 74. It is understood that the bus system 74 is used to achieve communication between these components. In addition to a data bus, the bus system 74 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 7 The general labeled all buses as Bus System 74.
[0171] The memory 73 in this embodiment is used to store various types of data to support the operation of the 3D reconstruction device 70. Examples of such data include any computer program used to operate on the 3D reconstruction device 70.
[0172] The three-dimensional reconstruction method disclosed in the above embodiments of this application can be applied to the processor 72, or implemented by the processor 72. The processor 72 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above three-dimensional reconstruction method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the processor 72. The processor 72 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 72 can implement or execute the three-dimensional reconstruction methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the three-dimensional reconstruction method disclosed in the embodiments of this application can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the memory 73. The processor 72 reads the information in the memory 73 and combines its hardware to complete the steps of the aforementioned three-dimensional reconstruction method.
[0173] In an exemplary embodiment, the 3D reconstruction device 70 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned 3D reconstruction method.
[0174] It is understood that the memory 73 in this embodiment can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 73 described in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0175] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 73 storing a computer program. This computer program can be executed by a processor 72 in the 3D reconstruction device 70 to complete the steps of the 3D reconstruction method described in the foregoing embodiments of this application. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0176] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a processor 72 in a 3D reconstruction device 70 to complete the steps of the 3D reconstruction method described in the foregoing embodiments of this application.
[0177] It should be noted that terms such as "first," "second," and "third" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0178] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0179] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A three-dimensional reconstruction method, characterized by, The method comprises: acquiring a single-view image; the single-view image contains one or more target objects; determining an initial three-dimensional mesh and a skeleton point cloud based on the single-view image; each sub-mesh in the initial three-dimensional mesh corresponds to each target object respectively; determining a first scale factor based on the initial three-dimensional mesh and the skeleton point cloud; the first scale factor represents a target local scale factor; fusing the first scale factor and a second scale factor to obtain a target scale factor, and determining a target three-dimensional mesh based on the target scale factor; the second scale factor represents a global scale factor.
2. The method of claim 1, wherein, The method comprises: inputting the single-view image into a first model to obtain the initial three-dimensional mesh output by the first model; wherein the first model represents a three-dimensional generation diffusion model, and the initial three-dimensional mesh contains one or more target objects in a real scene.
3. The method of claim 1, wherein, The method comprises: inputting the single-view image into a second model to obtain a pseudo-depth map output by the second model; the second model represents a monocular depth estimation model; determining the skeleton point cloud based on the pseudo-depth map.
4. The method of claim 3, wherein, The method comprises: acquiring a first matrix; the first matrix represents a camera intrinsic matrix of the single-view image; determining the skeleton point cloud based on the pseudo-depth map and the first matrix.
5. The method of claim 1, wherein, The method comprises: aligning the initial three-dimensional mesh and the skeleton point cloud to obtain an aligned three-dimensional mesh and an aligned skeleton point cloud; determining the first scale factor based on the aligned three-dimensional mesh and the aligned skeleton point cloud.
6. The method of claim 5, wherein, The method comprises: inputting the initial three-dimensional mesh and the skeleton point cloud into a third model to align the initial three-dimensional mesh and the skeleton point cloud by using the third model, so as to obtain the aligned three-dimensional mesh and the aligned skeleton point cloud; wherein the third model represents a six-degree-of-freedom pose estimation model.
7. The method of claim 5, wherein, The method comprises: constructing a loss function; minimizing the geometric distance between the aligned three-dimensional mesh and the aligned skeleton point cloud by using the loss function to obtain a third scale factor; the third scale factor represents an optimal chamfer scale factor; determining the first scale factor based on the third scale factor.
8. The method of claim 1, wherein, The method further comprises: determining the second scale factor; The method comprises: determining first size information and second size information; the first size information represents real size information of a reference image, and the second size information represents directional bounding box size information of a standardized aligned three-dimensional mesh; determining the second scale factor based on the first size information and the second size information.
9. The method of claim 8, wherein, The method comprises one of the following: determining the first size information based on real size information contained in an external reference database when the target object image feature satisfies a first condition; determining the first size information based on average size of a reference object in the single-view image when the target object image feature does not satisfy the first condition; wherein the first condition comprises that the target object image feature successfully matches a reference image feature in the reference database.
10. The method of claim 8, wherein, determining the second scale factor based on a ratio of the first size information to the second size information. The method further comprises:
11. The method of claim 1, wherein, determining a standardized aligned three-dimensional mesh; applying the target scale factor to the standardized aligned three-dimensional mesh to obtain the target three-dimensional mesh. The method further comprises: determining a target volume based on the target three-dimensional mesh; the target volume representing a real volume of the target object.
12. The method according to any one of claims 1 to 11, characterized in that, The apparatus comprises: an acquisition unit configured to acquire a single-view image; the single-view image containing one or more target objects; 13. A three-dimensional reconstruction apparatus, characterized by comprising: a first determination unit configured to determine an initial three-dimensional mesh based on the single-view image; each sub-mesh in the initial three-dimensional mesh corresponding to each target object respectively; a second determination unit configured to determine a skeleton point cloud based on the single-view image; a third determination unit configured to determine a first scale factor based on the initial three-dimensional mesh and the skeleton point cloud; the first scale factor representing a target local scale factor; a fusion unit configured to fuse the first scale factor and a second scale factor to obtain a target scale factor; the second scale factor representing a global scale factor; a fourth determination unit configured to determine a target three-dimensional mesh based on the target scale factor. comprise: a processor and a memory for storing a computer program capable of running on the processor; 14. A three-dimensional reconstruction device, characterized by wherein the processor is configured to execute the computer program to perform the steps of the method of any one of claims 1 to 12. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 12. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 12.
15. A storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 12.
16. A computer program product comprising a computer program, characterized in that,