A single-view 3D reconstruction method fusing structure perception and geometric features
Patent Information
- Application Number
- CN202311462326.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-06
AI Technical Summary
Smart Images

Figure CN117409062B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision 3D reconstruction, specifically relating to a single-view method that integrates structural perception and geometric features. Figure 3 Dimensional reconstruction method. Background Technology
[0002] single vision Figure 3 3D reconstruction refers to the process of reconstructing a three-dimensional object using a single two-dimensional image. Figure 3 3D shape reconstruction has always been a hot topic in the field of computer vision. Due to the complexity and irregularity of 3D shape data, effectively representing 3D shapes remains a challenging problem. Furthermore, 3D reconstruction requires understanding the shape and appearance of objects from 2D images, which is also a daunting task for computers. In recent years, deep learning has become a powerful tool for solving various computer vision tasks, and it has proven effective for 3D reconstruction from 2D images. In early research on 3D shape representation, 3D objects were typically modeled using global methods, such as constructing solid geometry and deforming hyperquadratic surfaces. However, when dealing with local details, the reconstruction results may suffer from inaccurate or missing details due to the lack of global features. To compensate for the shortcomings of global methods, researchers have begun to integrate local feature reconstruction methods into the reconstruction process. Methods based on local features can better describe the details and shape of object surfaces.
[0003] However, 3D reconstruction methods based on the fusion of global and local information, such as DISN, struggle to perceive the structural information of the model. The reconstruction process does not consider connections and symmetry, including structural relationships such as rotation, translation, and mirror symmetry, and cannot effectively handle the discontinuities in the reconstruction results of occluded parts and thin or fine components. Summary of the Invention
[0004] To address the problems mentioned above, this invention proposes a single-view approach that integrates structure awareness and geometric features. Figure 3 The 3D reconstruction method, under the joint constraints of structural features, global geometric features, and local geometric features, achieves faster model training convergence and significantly improves reconstruction accuracy.
[0005] The technical solution adopted in this invention is: a single-view method that integrates structure perception and geometric features. Figure 3 The 3D reconstruction method includes the following steps:
[0006] Step 1.1 Train a single-view algorithm that fuses structure awareness and geometric features using 3D models and 2D rendered images from the ShapeNet Core dataset. Figure 3 The network is reconstructed; specifically, the structural feature extraction network is trained separately first. Then, the parameters of the structural feature extraction network are frozen, and the SDF prediction network is trained using the dataset with a symbolic distance function.
[0007] Step 2.1 Input the single view to be reconstructed into the network, extract structural features, and encode the structural features into structural feature vectors; extract global geometric features, and extract local features from the global geometric features based on spatial sampling point information and camera parameters; specifically, the structural feature acquisition first uses a ResNet-50 network to extract structural features from the input single view, and then encodes the structural features into implicit vectors representing the structural features through an encoder. A recursive neural network for structural reconstruction is used to recursively classify nodes using a node classifier, and then decodes the nodes using different node decoders according to the category until no further decoding is possible, i.e., reaching the leaf nodes of the tree structure. The leaf nodes represent the parameters of the directed bounding box of each component. Similarly, global geometric feature vector encoding is performed using a ResNet-50 network. For a spatial sampling point p, it is projected onto a single view image point q according to the camera parameters. The index q is found in the corresponding position in the global geometric feature sub-graph, with dimensions of 256, 512, 1024, and 2048, respectively, and then connected to obtain the local feature vector.
[0008] Step 2.2 After fusing the global geometric feature vector, structural feature vector, and spatial sampling point information, a coarse representation of the 3D model to be reconstructed is obtained by decoding. Specifically, a multilayer perceptron is used to map the given spatial sampling point p to a higher-dimensional feature space. Then, the higher-dimensional feature is fused with the implicit vectors of the global feature and structural feature respectively, and then decoded to obtain the signed distance function (SDF) value of the spatial sampling point, which serves as the coarse representation of the 3D model of the single view to be reconstructed.
[0009] Step 2.3: After fusing local features and spatial sampling point information, a refined representation of the 3D model of the single view to be reconstructed is obtained;
[0010] Step 2.4 fuses the coarse and fine representations of the 3D model of the single view to be reconstructed to obtain the signed distance function (SDF) representation of the 3D model. To generate the implicit plane, a dense 3D mesh with a resolution of 256×256×256 is first defined, the point cloud of the sampling points is placed into it, and the SDF value is predicted for each point in the mesh. After obtaining the SDF value of each point in the dense mesh, moving cubes are used to generate the corresponding plane on the isosurface where the SDF value is 0.
[0011] Compared with the prior art, the present invention has the following advantages:
[0012] (1) The present invention proposes a single-view method that integrates structural perception and geometric features. Figure 3 The 3D reconstruction method generates a model that is significantly different from the previous single-view model. Figure 3 The 3D reconstruction method has a more accurate reconstruction surface.
[0013] (2) Under the constraints of global features, structural features and local features, SDF is more accurate, reduces outliers, and better restores details.
[0014] (3) In the reconstruction process of this invention, the structural information of the perception model is considered, including connection and symmetry, such as rotation, translation, and mirror symmetry, so as to better handle the discontinuity of the reconstruction results of occluded parts and thin and fine parts.
[0015] The advantages of this invention provide a novel technical method for 3D reconstruction based on a single view, and have significant practical value. Attached Figure Description
[0016] Figure 1 This invention proposes a single-view method that integrates structural perception and geometric features. Figure 3 Reconstruction flowchart.
[0017] Figure 2 This invention proposes a single-view method that integrates structural perception and geometric features. Figure 3 Reconstruction network framework diagram.
[0018] Figure 3 The present invention provides a three-dimensional structure reconstruction network framework.
[0019] Figure 4 The present invention provides a feature extraction network framework.
[0020] Figure 5 Examples of the three-dimensional structure reconstruction results of this invention.
[0021] Figure 6 The reconstruction results of this invention are compared with those of other methods. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0023] like Figure 1 As shown, this invention discloses a single-view method that integrates structure perception and geometric features. Figure 3The method for 3D model reconstruction extracts hierarchical structural features from the input single-view image to be reconstructed, encoding these features into structural feature vectors to represent structural information such as rotation, symmetry, and connectivity. It also extracts global features from the input single-view image, extracting local features from the global feature blocks based on sampling point information and camera parameters. The structural and global geometric features are then fused and decoded to obtain a coarse but reasonable implicit surface function representation of the 3D model of the single-view image to be reconstructed. The local geometric features are then decoded to obtain a fine representation of the 3D model. The coarse and fine representations of the 3D model of the single-view image to be reconstructed are further fused to obtain the signed distance function (SDF) representation of the 3D model. Finally, the surfaces with an isosurface value of 0 in the SDF representation are used as the surfaces of the 3D model.
[0024] like Figure 2 As shown, the network of this invention consists of two networks. The first network extracts hierarchical features from the input single-view, i.e., two-dimensional image, encoding them into a fixed-length implicit vector (root code). This implicit vector can be decoded into the composition structure of a three-dimensional shape corresponding to the input single-view. The second network first extracts the global features of the single-view, i.e., two-dimensional image, fuses them with the implicit vector reflecting hierarchical structure information and the coordinates of the sampling point, and predicts the SDF of the sampling point. global The value is then calculated; secondly, multi-scale local features of the sampling points are extracted and fused with the sampling point location coordinates to predict the SDF of the sampling point. local Value, finally, SDF global With SDF local The final SDF value is determined by fusion, and the surfaces with an SDF value of 0 are extracted as the final surface of the generated 3D model.
[0025] This invention is a single-view Figure 3 The training process of the hierarchical reconstruction network is divided into two stages: training of the hierarchical structure reconstruction network and training of the SDF prediction network.
[0026] Hierarchical structure reconstruction network training: This stage requires extracting hierarchical structure features from single-view image features and needs to be trained separately. For example... Figure 3The single-view image shown is processed by a ResNet-50 network to extract hierarchical structure features. These features are then encoded into root node features using a Multilayer Perceptron (MLP). The root node features are then used to reconstruct the structure via a Recurrent Neural Network (RvNN). Each root node is first classified into leaf node, symmetric node, and connection node types. Leaf node types are regressed to their directed bounding box values, symmetric node types to their symmetry parameters and to generate child nodes, and connection node types to generate two child nodes. During decoding, the node classification network is recursively called to determine which decoder to use for each root node until the corresponding directed bounding box of the leaf node can be decoded, thus reconstructing the complete hierarchical structure. As the structure reconstruction continues to optimize, the root node features are gradually decoupled from the single-view image features to contain only hierarchical structure features, which are then used for subsequent 3D model reconstruction. The structure reconstruction network uses stochastic gradient descent (SGD) to optimize the structure reconstruction network and trains the recurrent neural network decoder through backpropagation through time (BPTT).
[0027] SDF Prediction Network Training: This stage of training freezes the parameters from the single-view image to the root node structural features, ensuring the stability of the root node hierarchical structural features obtained from the single-view image feature encoding. The root node hierarchical structural features are concatenated with the image features to reconstruct a coarse 3D model from the single-view image. For example... Figure 4 As shown, the feature maps of blocks in the ResNet-50 network are upsampled to the size of the original single-view image. Camera parameter transformation is used to find the position of each sampling point on the feature map, resulting in 1×1 local features with 256, 512, 1024, and 2048 channels respectively. After concatenation of the local features, a 1×1 Conv layer is applied to reduce the dimensionality to 1024. Camera parameters can be predicted from the input image. To better demonstrate the effectiveness of the SDF prediction network that fuses structural features, ground-based camera parameters are used to train the network. The position information p(x,y,z) of the sampling points is mapped to a 512-dimensional high-dimensional feature through an MLP, and then concatenated with both global and local features. Two different MLP networks are used to decode the global shape and local detail shape respectively, and the final sum of the decoding results is the final predicted SDF value.
[0028] The purpose of the structural reconstruction network in this invention is solely to obtain the hierarchical structure encoding of the 3D model. During the training of the SDF prediction network, the loss is not fed back to this network; the two networks are trained independently, and therefore their loss functions are also independent. When reconstructing the hierarchical structure from a single-view image, the reconstruction effect is constrained by three parts of the loss function, designed as follows: The first part is the mean squared error loss (MSE) between the reconstructed directed bounding boxes; the second part is the mean squared error loss between symmetric information in the hierarchical structure; and the third part is the cross-entropy loss (MAE) during node classification.
[0029]
[0030] Where, n b n s n c These represent the number of bounding boxes, the number of symmetric nodes, and the number of connected nodes in the hierarchical structure of each 3D model, respectively. These represent the bounding box location, symmetry parameter, and node classification predicted by the network, respectively. Instead of establishing a binary classification problem (e.g., inside or outside the shape) as in Shape Model Generative Implicit Field Networks (IM-NET), we regress continuous SDF values. Furthermore, to ensure the network focuses on recovering details near and inside the isosurface S0, we use a weighted loss function. The loss is defined as:
[0031]
[0032] Where f(I,p) represents the input of a single-view image I and three-dimensional sampling points p into this network, while SDF I (p) represents the true SDF value on the ground, m1 and m2 represent different weights, and σ is an empirical threshold.
[0033] We chose m1 = 1, m2 = 4, and σ = 0.1 as the hyperparameters of the loss function. The second stage of training used the Adam optimizer to update the weights, with a learning rate of 1 × 10⁻⁶. -4 The batch size is 48.
[0034] During the testing phase, a 2D image is input. First, the camera parameters for the image are determined; in this invention, groundtruth camera parameters are used. To output the SDF value of the spatial location, 2563 sampling points with uniform density are preset in a standard space. By transforming the spatial location using the image and camera parameters, hierarchical structure features, global features, and local features are obtained. The hierarchical structure features are concatenated with the global features to reconstruct a coarse global shape, while the local features are used to reconstruct local detailed shapes. Their sum gives the final predicted SDF value. Finally, the Marchingcubes algorithm is used to extract the zero-value surface and reconstruct the 3D model.
[0035] For the hierarchical structure reconstruction network, the proposed PartNet-Symh dataset is used, which adds a recursive hierarchical organization of fine-grained parts to each shape to further enhance the data. This dataset contains 22,369 3D shapes, covering 24 shape categories. Each model has seven folders; the obbs folder is mainly used in this invention, which stores the entire shape obb for each model, including the original part obb, adjacent part relationships, and symmetry parameters. The corresponding hierarchical structure feature vector is obtained after processing this data. The SDF prediction network dataset uses the ShapeNetCore3D model. Based on the human eye's viewing angle, 24 random camera viewpoints are used to capture 3D data using Blender software to obtain 2D rendered images. Each rendered image is 137×137 pixels. Camera parameters for each image are saved during rendering. While our method can regress SDF values from inputs at arbitrary locations, thus obtaining reconstruction results at arbitrary resolutions, during training, we are only interested in sampling points on the surface of the 3D model. Therefore, SDF data and labels are obtained by Gaussian sampling of 32,768 points near the 3D model surface. In each iteration of the training process, a weighted sampling strategy is used to randomly select 2,048 points to calculate the loss and update the weights. Experiments on the hierarchical reconstruction network primarily focus on subjectively evaluating the training effect. Specifically, we evaluate whether the hierarchical structure corresponding to a single view conforms to human visual expectations to obtain the optimal training results. The evaluation results are as follows: Figure 5 As shown.
[0036] In summary, this invention constructs a single-view system that integrates structure perception and geometric features. Figure 3 The network is reconstructed; its network is easy to train, such as... Figure 6The reconstruction method presented in this paper achieves higher accuracy. The 3D SDF representation is divided into global SDF and local SDF. Hierarchical structural features are decoupled from global features and fused with the original global features, resulting in a more accurate global SDF and better constraint on local features, reducing outliers. By fusing local blocks projected from feature maps of 3D points at different scales and predicting the local SDF of that point, the model can capture fine-grained features. Parts not described in detail in this invention are well-known techniques to those skilled in the art.
Claims
1. A single-view 3D reconstruction method integrating structure perception and geometric features, characterized in that, Includes the following steps: Step 1.1 Train a single-view 3D reconstruction network that integrates structure awareness and geometric features using 3D models and 2D rendered images from the ShapeNet Core dataset. The training of the single-view 3D reconstruction network is divided into training the structural feature extraction network and training the symbolic distance function (SDF) prediction network. The single-view image is processed by the ResNet-50 feature extraction network to extract hierarchical structural features, which are then encoded into root features by a multilayer perceptron (MLP). The root node features of the code are reconstructed through a recurrent neural network. Each root node is first classified to determine its type: leaf node, symmetric node, or connection node. Leaf node roots regress the directed bounding box values, symmetric nodes regress the symmetry parameters and generate child nodes, and connection nodes generate two child nodes. During decoding, the node classification network is recursively called to determine which decoder to send the root node to, until the corresponding directed bounding box leaf node is decoded, thus reconstructing the complete hierarchical structure. During the continuous optimization of the structure reconstruction, the root node features are gradually decoupled from the single-view image features to contain only hierarchical structure features for subsequent 3D model reconstruction. The structure reconstruction network uses stochastic gradient descent to optimize the structure reconstruction network and trains the recurrent neural network decoder through backpropagation of the structure. Step 2.1 Input the single view to be reconstructed into the single view 3D reconstruction network, extract structural features and encode the structural features into structural feature vectors; extract global geometric features, and extract local geometric features from the global geometric features based on spatial sampling point information and camera parameters; the local geometric feature vector extraction method is as follows: global feature encoding is performed through the ResNet50 network. For spatial sampling point p, it is projected onto pixel q of the single view to be reconstructed according to the camera parameters. The position of pixel q in the feature extraction network of different size feature layers is indexed, and then the values of these corresponding positions are concatenated to obtain the local feature vector; Step 2.2 After fusing the global geometric feature vector, structural feature vector, and spatial sampling point information, a coarse representation of the 3D model of the single view to be reconstructed is obtained by decoding. Step 2.3 After fusing local features and spatial sampling point information, a refined representation of the 3D model of the single view to be reconstructed is obtained; Step 2.4 The coarse and fine representations of the 3D model of the single view to be reconstructed are fused to obtain the symbolic distance function representation (SDF) of the 3D model of the single view to be reconstructed. Finally, the surface with a symbolic distance function representation (SDF) isosurface of 0 is obtained as the surface of the 3D model of the single view to be reconstructed.
2. The method according to claim 1, characterized in that: In step 1.1, the structural feature extraction network is first trained separately, and then the parameters of the structural feature extraction network are frozen and the SDF prediction network is trained using the dataset to represent the symbolic distance function.
3. The method according to claim 1, characterized in that: The structural features in step 2.1 are obtained as follows: First, the ResNet50 network is used to extract structural features from the single view to be reconstructed. Then, the extracted structural features are encoded by an encoder into an implicit vector representing the structural features, namely the structural feature vector.
4. The method according to claim 1, characterized in that: In step 2.2, a multilayer perceptron is used to map the given spatial sampling point p to a higher-dimensional feature space to obtain high-dimensional features. Then, the high-dimensional features are fused with global features and structural features respectively, and then decoded to obtain the symbolic distance function (SDF) value of the spatial sampling point, which serves as a coarse representation of the 3D model of the single view to be reconstructed.
5. The method according to claim 1, characterized in that: Step 1.1 The loss function Loss for training the structural feature extraction network is as follows: Where, n b n s n c These represent the number of bounding boxes, the number of symmetric nodes, and the number of connected nodes in the structural features of each 3D model, respectively. , , These represent the bounding box location, symmetry parameters, and node classifications predicted by the network. , , These represent the actual bounding box location, symmetry parameters, and node classification, respectively.
6. The method according to claim 1, characterized in that: Step 1.1 Training the symbolic distance function to represent the loss of the SDF prediction network. The LSDF is defined as follows: Where f(I, p) represents the predicted SDF value using a single-view image I and sampling point p as input, and SDF... I (p) represents the true SDF value, and m1 and m2 represent different weights. It is a threshold.