A hierarchical 3D reconstruction method based on an auto-decoder
By using improved convolutional neural networks and neural signed distance fields in the automatic decoder framework, the problem of difficult to reconstruct hierarchical structures and internal structures in the prior art is solved, and a high-precision three-dimensional shape reconstruction is achieved.
Patent Information
- Application Number
- CN202410931885.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-12
AI Technical Summary
Existing three-dimensional reconstruction technologies based on deep learning are difficult to effectively reconstruct hierarchical structures or internal structures, and the reconstruction accuracy is low.
Using a hierarchical three-dimensional reconstruction method based on an automatic decoder, data features are extracted through an improved convolutional neural network, and signed distance values of the nearest surface of any spatial point are predicted in the neural signed distance field to reconstruct the three-dimensional shape with hierarchical structure.
The three-dimensional shapes with hierarchical structures and internal structures are realized with high accuracy, including non-watertight shapes and multi-layer surface shapes, improving the accuracy and effect of three-dimensional reconstruction.
Smart Images

Figure CN118691761B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer three-dimensional reconstruction, computer vision, big data cloud computing, etc., and specifically relates to a hierarchical three-dimensional reconstruction method based on an auto-decoder. Background Art
[0002] For the method of explicit mesh representation, a mesh is a structure composed of a set of vertices and edges. Mesh representation is a method of decomposing a three-dimensional object into a series of planes or curved surfaces. These planes or curved surfaces intersect to form a mesh structure composed of line segments and vertices. Each line segment is usually represented as a connection between two vertices, and each vertex has coordinates in three-dimensional space. Graph convolution can be directly applied on the mesh for geometric learning. The mesh can also be used as the output representation of shape reconstruction. However, most mesh-based methods deform a template and are restricted by a fixed topology. Recent methods directly predict the vertices and faces of the shape, but are prone to generating self-intersecting faces, thus losing the surface continuity. 3D point cloud is a common method for representing three-dimensional shapes. A 3D point cloud is a dataset composed of a series of three-dimensional points, used to represent the surface or spatial distribution of an object. The point cloud can also be used as the output representation of 3D reconstruction. Different from other representation forms, the point cloud does not contain topological information, so usually a series of complex post-processing steps are required to generate a renderable surface. Voxel is a representation method widely used in shape learning and is similar to pixels in two-dimensional space in three-dimensional space. It is a cubic unit in a three-dimensional mesh and is usually used to represent the volume data of a three-dimensional object. Each voxel has a position coordinate and possible attribute values, such as density, color, or other attributes. Since the storage space grows in a cubic form, the resolution of the mesh is restricted in practice. Therefore, it is difficult for voxel-based methods to reconstruct shapes with high-fidelity details.
[0003] "Watertight shapes" refer to those completely enclosed three-dimensional models without any gaps or holes. The surface of this kind of shape is completely continuous without any open boundaries. "Non-watertight shapes" refer to those shapes in three-dimensional modeling and computer graphics that do not have a completely enclosed surface. These shapes may contain voids, cracks, holes, or unconnected edges.
[0004] In recent years, many progresses have been made in three-dimensional shape learning using neural implicit functions. Although the output methods are different, these methods are all committed to learning a continuous function to predict the relationship between query points in three-dimensional space and the surface. Traditional implicit expression methods will result in the loss of the internal structure of the shape. Therefore, how to improve the ability of the network architecture to learn hierarchical structures and enhance the reconstruction fineness of shapes with internal structures is an important problem to be solved in this field. Summary of the Invention
[0005] Aiming at the problems that the current mainstream 3D reconstruction technology based on deep learning cannot reconstruct hierarchical structures or has low accuracy in reconstructing internal structures, the present invention provides a hierarchical 3D reconstruction method based on an auto-decoder. The method can learn the relationship between two points in space in the auto-decoder framework, extract data features using an improved convolutional neural network, predict the signed distance value of the nearest surface of any spatial point in the neural signed distance field, and reconstruct a 3D shape with a hierarchical structure.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A hierarchical 3D reconstruction method based on an auto-decoder, the method comprising the following steps:
[0008] Step 1, divide the ShapeNet dataset, including the following sub-steps:
[0009] Step 1.1, first import the dependencies of collections, glob, and numpy, and load the ShapeNet dataset path and configuration file;
[0010] Step 1.2, use the glob function to obtain all file paths under the specified path, and at the same time read the directories where all files are located;
[0011] Step 1.3, sort all samples in a non-random manner, and divide the ShapeNet dataset according to the set ratio, and calculate the corresponding number of sample quantities;
[0012] Step 1.4, save the division result to a.npz file in a unified format;
[0013] Step 2, perform data preprocessing on the divided ShapeNet dataset, including the following sub-steps:
[0014] Step 2.1, call the data scaling function to convert the original data into a 3D model file in OFF format. The data scaling function receives a file path as a parameter, indicating the 3D model file to be processed, determines the output file path according to the input file path. If the output file already exists, it is directly skipped. Use trimesh.load to load the 3D model file, obtain the model object, call the mesh function to convert the model object into a mesh object, and calculate the boundary size and center point coordinates of the model. Move the center point of the mesh object to the origin and perform a scaling operation according to the boundary size to ensure that the model adapts to the given size. Export the scaled model as a file in OFF format;
[0015] Step 2.2: Set the number of sampling point pairs to 15000, and set three sampling ranges: surface sampling of the three-dimensional shape, sampling within the bounding box, and spatial sampling of the three-dimensional shape. Define a function for generating labels that takes a file path as input, where the file path stores the scaled three-dimensional model to be processed. Use the label generation function to call the compiled file to generate corresponding binary flags for each pair of sampling points, and define three functions inside the function to downsample the data respectively to ensure that the number of sampling point pairs does not exceed the specified value. Balance the positive and negative samples and downsample them to the specified number. Convert the point cloud coordinates to grid coordinates and obtain the data of the signed distance value for each sampling point.
[0016] Step 2.3: Define a voxelized point cloud function, parse the input file path, and construct the file path of the sampled point cloud data. Initialize an all-zero array to store the voxelized point cloud data, accept the minimum value, maximum value, and resolution as inputs, generate an array of coordinates of three-dimensional grid points, use a KD tree for nearest neighbor search, set the voxel corresponding to the nearest neighbor point to 1, compress the voxelized data, and save the point cloud data, the compressed voxel data, and other relevant information as a unified format of.npz file.
[0017] Step 2.4: Save the datasets processed in Step 2.2 and Step 2.3 to the determined file directory to provide a data source for the training and reconstruction of the model.
[0018] Step 3: Build a network architecture, which includes an encoder (i.e., a feature extraction network), a point embedding network, an auto-decoder, and an SDF network. Input the dataset preprocessed in Step 2 into the network architecture built in Step 3 for training, adjust the hyperparameters and weights of the network architecture to obtain a trained neural network model, including the following sub-steps:
[0019] Step 3.1: Use the deep 3D convolutional network of the encoder to extract the spatial features of the input data layer by layer. Apply a batch normalization layer and a ReLU activation function after each 3D convolutional layer to enhance the non-linear ability and stability of the encoder.
[0020] Step 3.2: Construct an encoder, i.e., a feature extraction network. The encoder extracts deep features from the processed data using the constructed convolutional layers. After the input of the first layer, the first initial feature map is obtained. Then, one feature map is output for every two convolutional layers, and finally, the obtained feature maps are used as one of the input data for the point embedding layer. Construct a point embedding network that expands the point coordinate dimension, displaces the point coordinates according to the displacement vector, then performs grid sampling on the feature maps generated by the encoder, extracts the features of the displaced point coordinates, combines the features of all layers and the features corresponding to each displaced point coordinate along the feature dimension to form a comprehensive feature representation, then adjusts the shape of the features to unify the features from different displacements and levels into a continuous vector, and finally combines the position information of the original points with the extracted features. Construct an auto-decoder, and use a fully connected network implemented in the form of a seven-layer 1D convolutional layer as the auto-decoder part.
[0021] Step 3.3: Construct an SDF network. Input the dataset preprocessed in Step 2 into the encoder, i.e., the feature extraction network, constructed in Step 3.2. Use a constructed ten-layer fully connected network as the SDF network, i.e., the SDF Network part, to learn the signed distance values of the sampling points, and use the auto-decoder to learn the spatial relationship between a pair of sampling points.
[0022] Step 4: Input the data used for reconstruction in Step 2 into the neural network model trained in Step 3, and use the marching cubes algorithm to reconstruct a shape with internal structure, including the following sub-steps:
[0023] Step 4.1: Load the parameters of the neural network model trained in Step 3, the threshold, and the preprocessed data obtained in Step 2.3.
[0024] Step 4.2: Define the initial resolution of the grid, and then use a step-by-step refinement operation to increase the resolution of the grid. In each level of refinement, calculate the grid center.
[0025] Step 4.3: Use the signs of the signed distance values of each point to take points with opposite signs as a pair of points, and use the auto-decoder to predict the binary flags of this pair of points.
[0026] Step 4.4: Use the SDF network of the model to predict the signed distance values corresponding to the points in each grid, determine the points participating in the reconstruction algorithm according to the set threshold, and reconstruct a three-dimensional shape with internal structure.
[0027] Furthermore, the sub-step of using a deep 3D convolutional network of the encoder in Step 3.1 to gradually extract the spatial features of the input data, and applying a batch normalization layer and a ReLU activation function after each 3D convolutional layer is as follows:
[0028] For the first convolutional layer, the number of input channels is 1, the number of output channels is 16, the kernel size is 3, and padding of 1 is used to maintain the spatial dimensions. This is followed by two consecutive convolutional layers. The first one increases the number of channels to 32, and the second one maintains 32 channels. Then, the next two convolutional layers increase the number of channels to 64, and the two convolutional layers after that increase the number of channels to 128. Finally, the last six convolutional layers maintain the number of channels at 128. Max pooling layers are applied respectively after the first, third, fifth, seventh, and ninth convolutional layers. The max pooling layer is used to halve the size to 128 after the first convolutional layer, to 64 after the third layer, to 32 after the fifth layer, to 16 after the seventh convolutional layer, and to 8 after the ninth convolutional layer;
[0029] After the convolutional layers, fully connected layers are built. First, a ten-layer fully connected network is built. Among them, the first fully connected layer expands the features to twice the hidden dimension, the second fully connected layer reduces the feature dimension to the hidden dimension, the third layer and subsequent layers maintain the hidden dimension, and the data of the input layer is input again at the fifth layer. Finally, the output layer outputs the signed distance value predicted by the model. Then, a seven-layer autoencoder of fully connected layers implemented in the form of 1D convolutional layers is built. The first layer processes the input features and maps them to twice the hidden dimension. The middle three layers maintain the hidden dimension, and the output layer outputs the label value predicted by the model.
[0030] Furthermore, the sub-steps of reconstructing the three-dimensional shape with internal structure in step 4.4 are as follows:
[0031] Use the Marching Cubes algorithm to determine the boundaries of the grid based on the binary flags, determine the surface of the grid according to the signed distance value, and finally save the generated three-dimensional shape to a.obj file in a unified format.
[0032] The technical solution provided by the present invention has the following beneficial effects compared with the prior art:
[0033] The reconstruction method of the present invention can effectively represent various three-dimensional shapes with complex multi-layer surfaces including non-watertight shapes. The present invention uses a technology different from the shape representation of traditional implicit functions to process shapes with internal structures. Starting from the framework of an auto-decoder, the feature extraction network is improved to reconstruct three-dimensional shapes with hierarchical structures with high precision, including shapes with non-seamless connections and shapes with multi-layer surfaces. The present invention adopts a hierarchical shape representation and an auto-decoder framework, which can more accurately predict the signed distance value between a three-dimensional space query point and the surface, and shows good performance in accurately reconstructing three-dimensional shapes with internal structures. The method of the present invention designs a new mapping relationship, improves the ability of the neural network model to learn hierarchical structures, and enables the neural network model to more effectively learn from data and reconstruct three-dimensional shapes with internal structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0035] Figure 1 is the network architecture diagram of a hierarchical three-dimensional reconstruction method based on an auto-decoder described in an embodiment of the present invention;
[0036] Figure 2 is the flowchart of steps 1 and 2 of a hierarchical three-dimensional reconstruction method based on an auto-decoder described in an embodiment of the present invention;
[0037] Figure 3 is the flowchart of steps 3 and 4 of a hierarchical three-dimensional reconstruction method based on an auto-decoder described in an embodiment of the present invention;
[0038] Figure 4 is the reconstruction effect diagram of the three-dimensional shapes with internal structures of vehicle a, vehicle b, and vehicle c in an embodiment of the present invention;
[0039] Figure 5 is the reconstruction effect diagram of the three-dimensional shape with internal structure of truck a in an embodiment of the present invention;
[0040] Figure 6 is the reconstruction effect diagram of the three-dimensional shape with internal structure of truck b in an embodiment of the present invention;
[0041] Figure 7 is the reconstruction effect diagram of the three-dimensional shape with internal structure of truck c in an embodiment of the present invention;
[0042] Figure 8 is the multi-angle hierarchical structure reconstruction effect diagram of vehicle d in an embodiment of the present invention;
[0043] Figure 9 It is the multi - angle hierarchical structure reconstruction effect diagram of truck d and car e in a certain embodiment of the present invention. Detailed implementation manners
[0044] In order to more clearly understand the above - mentioned objects, features, and advantages of the present invention, the solutions of the present invention will be further described below. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] In one embodiment, as Figure 1 shown, it is the network architecture diagram of a three - dimensional reconstruction method of a hierarchical structure based on an auto - decoder in a certain embodiment of the present invention; as Figure 2 shown, it is the flowchart of step 1 and step 2 of a three - dimensional reconstruction method of a hierarchical structure based on an auto - decoder in a certain embodiment of the present invention; as Figure 3 shown, it is the flowchart of step 3 and step 4 of a three - dimensional reconstruction method of a hierarchical structure based on an auto - decoder in a certain embodiment of the present invention, including the following steps:
[0046] Step 1: Divide the ShapeNet dataset, including the following sub - steps:
[0047] Step 1.1: First, import the dependencies of collections, glob, and numpy, and load the ShapeNet dataset path and configuration file;
[0048] Step 1.2: Use the glob function to obtain all file paths under the specified path, and at the same time read the directories where all files are located;
[0049] Step 1.3: Sort all samples in a non - random manner, and divide the ShapeNet dataset according to the set ratio, and calculate the corresponding number of sample quantities;
[0050] Step 1.4: Save the division result into a.npz file in a unified format;
[0051] Step 2: Pre - process the divided ShapeNet dataset, including the following sub - steps:
[0052] Step 2.1, call the data scaling function to convert the original data into a 3D model file in OFF format. The data scaling function takes a file path as a parameter, which represents the 3D model file to be processed. Determine the output file path according to the input file path. If the output file already exists, skip it directly. Use trimesh.load to load the 3D model file, obtain the model object, call the mesh function to convert the model object into a mesh object, and calculate the boundary size and center point coordinates of the model. Move the center point of the mesh object to the origin and perform a scaling operation according to the boundary size to ensure that the model fits the given size. Export the scaled model as a file in OFF format;
[0053] Step 2.2, set the number of logarithm of sampling points to 15000, set three sampling ranges: surface sampling of the 3D shape, sampling within the bounding box, and spatial sampling of the 3D shape. Define a function to generate labels, which takes a file path as input, and the file path stores the scaled 3D model to be processed; Use the function to generate labels, call the compiled file to generate corresponding binary flags for each pair of sampling points, and the binary flag B used to reveal whether the line segment intersects the surface is defined as the following binary flag: ; while the binary flags in other cases are represented as where, represents two points in space, represents a certain surface in space, represents the surface the point on, represents the line segment where it is located, and three functions are defined again inside the function to downsample the data respectively to ensure that the number of sampling point pairs does not exceed the specified value; balance positive and negative samples and downsample to the specified number; convert the point cloud coordinates to mesh coordinates and obtain the data of signed distance values for each sampling point.
[0054] Step 2.3, define the voxelized point cloud function, parse the input file path, and construct the file path of the sampled point cloud data file; Initialize a zero-filled array to store the voxelized point cloud data, accept the minimum value, maximum value, and resolution as inputs, generate an array of coordinates of 3D grid points, use the KD tree for nearest neighbor search, set the voxel corresponding to the nearest neighbor point to 1, compress the voxelized data, and save the point cloud data, the compressed voxel data, and other relevant information as a unified format of.npz file.
[0055] Step 2.4, save the datasets processed in Step 2.2 and Step 2.3 to the determined file directory to provide data sources for the training and reconstruction of the model;
[0056] Step 3, build a network architecture, including an encoder (i.e., a feature extraction network), a point embedding network, an auto-decoder, and an SDF network. Input the dataset preprocessed in Step 2 into the network architecture built in Step 3 for training, and adjust the hyperparameters and weights of the network architecture to obtain a trained neural network model, including the following sub-steps:
[0057] Step 3.1, use the deep 3D convolutional network of the encoder to extract the spatial features of the input data layer by layer, convert the input point cloud into a discrete voxel grid, and apply a 3D convolutional neural network to obtain multi-scale grid features , where represents the grid size, which varies with the scale, and is the depth feature with M channels. After each 3D convolutional layer, a batch normalization layer and a ReLU activation function are applied to enhance the non-linearity and stability of the encoder.
[0058] Step 3.2, construct the encoder (i.e., the feature extraction network). The encoder uses the constructed convolutional layers to extract deep features from the processed data, obtains the first initial feature map after the input of the first layer, and then outputs a feature map every two convolutional layers. Finally, the obtained feature map is used as one of the input data of the point embedding layer. Construct the point embedding network. The point embedding layer expands the point coordinate dimension, displaces the point coordinates according to the displacement vector, then performs grid sampling on the feature map generated by the encoder, extracts the features of the displaced point coordinates, merges the features of all layers and the features corresponding to each displaced point coordinate along the feature dimension to form a comprehensive feature representation, then adjusts the shape of the features, unifies the features from different displacements and levels into a continuous vector, and finally combines the position information of the original points with the extracted features. Construct the auto-decoder, and use the constructed seven-layer fully connected network implemented in the form of 1D convolutional layers as the auto-decoder part;
[0059] Step 3.3, construct the SDF network. Input the dataset preprocessed in Step 2 into the encoder (i.e., the feature extraction network) constructed in Step 3.2, use the constructed ten-layer fully connected network as the SDF network (i.e., the SDF Network part) to learn the signed distance values of the sampling points, and use the auto-decoder to learn the spatial relationship between a pair of sampling points;
[0060] Step 4, load the data used for reconstruction in Step 2 into the neural network model trained in Step 3, and use the marching cubes algorithm to reconstruct the shape with internal structure, including the following sub-steps:
[0061] Step 4.1, load the parameters of the neural network model trained in Step 3, the threshold, and the preprocessed data obtained in Step 2.3;
[0062] Step 4.2, define the initial resolution of the grid, and then use the step-by-step refinement operation to increase the resolution of the grid. In each level of refinement, calculate the grid center;
[0063] Step 4.3, using the signs of the signed distance values of each point, take the points with opposite signs as a pair of points, and use the auto-decoder to predict the binary flags of this pair of points. , , where is used for binary flag prediction, represents the predicted shape, represents the auto-decoder, represents the point embedding layer, represents the predicted shape at represents the corresponding true shape, represents the index of the observation, represents the number of training samples, represents the observed value and the corresponding true shape match;
[0064] Step 4.4, use the SDF network of the model to predict the signed distance value corresponding to each point in the grid. , where represents regression, represents the set threshold, represents the network, determine the points participating in the reconstruction algorithm according to the set threshold, and reconstruct the three-dimensional shape with internal structure.
[0065] The present invention uses a technology different from the shape representation of traditional implicit functions to process shapes with internal structures. Starting from the framework of the auto-decoder, the feature extraction network is improved to reconstruct three-dimensional shapes with hierarchical structures with high accuracy, including shapes with non-seamless connections and shapes with multi-layer surfaces. The specific three-dimensional shape structure diagram for reconstruction can be seen in Figures 4 to 9 .
[0066] Furthermore, the sub-step of using the deep 3D convolutional network of the encoder to extract the spatial features of the input data layer by layer in step 3.1, and applying a batch normalization layer and a ReLU activation function after each 3D convolutional layer is:
[0067] For the first convolutional layer, the input channels are 1, the output channels are 16, the kernel size is 3, and padding of 1 is used to maintain the spatial dimensions. This is followed by two consecutive convolutional layers. The first one increases the number of channels to 32, and the second one maintains 32 channels. The next two convolutional layers increase the number of channels to 64, and the two convolutional layers after that increase the number of channels to 128. Finally, the last six convolutional layers maintain the number of channels at 128. Max pooling layers are applied after the first, third, fifth, seventh, and ninth convolutional layers respectively. The max pooling layer halves the size to 128 after the first convolutional layer, to 64 after the third layer, to 32 after the fifth layer, to 16 after the seventh convolutional layer, and to 8 after the ninth convolutional layer;
[0068] Fully connected layers are built after the convolutional layers. First, a ten-layer fully connected network is built. Among them, the first fully connected layer expands the features to twice the hidden dimension, the second fully connected layer reduces the feature dimension to the hidden dimension, and the third layer and subsequent layers maintain the hidden dimension. The data of the input layer is input again at the fifth layer. The final output layer outputs the signed distance value predicted by the model. Then, a seven-layer fully connected layer autoencoder implemented in the form of 1D convolutional layers is built. The first layer processes the input features and maps them to twice the hidden dimension. The middle three layers maintain the hidden dimension, and the output layer outputs the label value predicted by the model.
[0069] Furthermore, the sub-steps of reconstructing the three-dimensional shape with internal structure in step 4.4 are as follows:
[0070] The Marching Cubes algorithm is used to determine the boundaries of the grid according to the binary flags and the surface of the grid according to the signed distance values. Finally, the generated three-dimensional shape is saved to an.obj file in a unified format, which mainly includes three stages: (1) Loop traversal: Each grid cell is traversed through nested loops, and the position inside the cube is determined according to the distance values of each cube vertex using the combined vertex function. (2) Interpolation and triangulation: The vertex interpolation function is used to perform linear interpolation on the required edges to determine the exact positions of the new vertices, thereby generating triangles. (3) Constructing the final grid: Finally, the entire hierarchical grid model is constructed by connecting all the calculated triangle vertices.
[0071] As Figure 4 shown, it is the rendering of the three-dimensional shape reconstruction with internal structure of car a, car b, and car c in a certain embodiment of the present invention;
[0072] As Figure 5 shown, it is the rendering of the three-dimensional shape reconstruction with internal structure of truck a in a certain embodiment of the present invention;
[0073] As Figure 6As shown, it is the three-dimensional shape reconstruction effect diagram of the internal structure of truck b in a certain embodiment of the present invention;
[0074] As Figure 7 shown, it is the three-dimensional shape reconstruction effect diagram of the internal structure of truck c in a certain embodiment of the present invention;
[0075] As Figure 8 shown, it is the multi-angle hierarchical structure reconstruction effect diagram of car d in a certain embodiment of the present invention;
[0076] As Figure 9 shown, it is the multi-angle hierarchical structure reconstruction effect diagram of truck d and car e in a certain embodiment of the present invention.
[0077] The above are only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the foregoing embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A hierarchical 3D reconstruction method based on automatic decoder, characterized in that: The following steps are involved: Step 1: Divide the ShapeNet dataset into data sets, including the following sub-steps: Step 1.1, first import collections, glob, numpy dependencies, load the ShapeNet dataset path and configuration file; Step 1.2, use the glob function to obtain all file paths under the specified path, and read the directories where all files are located; Step 1.3, sort all samples in a non-random way, divide the ShapeNet dataset according to the set ratio, and calculate the corresponding number of samples; Step 1.4, save the partitioning results into a .npz file in a unified format; Step 2: Preprocess the divided ShapeNet dataset, including the following sub-steps: Step 2.1, call the data scaling function to convert the original data into a 3D model file in OFF format. The data scaling function receives a file path as a parameter, which indicates the 3D model file to be processed. The output file path is determined according to the input file path. If the output file already exists, it is skipped directly. Use trimesh.load to load the 3D model file, obtain the model object, call the mesh function to convert the model object into a mesh object, and calculate the boundary size and center point coordinates of the model. Move the center point of the mesh object to the origin, and scale it according to the boundary size to ensure that the model fits the given size. Export the scaled model to a file in OFF format. Step 2.2, set the number of sampling point pairs to 15000, set three sampling ranges: surface sampling of three-dimensional shapes, sampling within bounding boxes, and spatial sampling of three-dimensional shapes; define a function to generate labels, receive a file path as input, and the file path stores the scaled three-dimensional model to be processed; use the label generation function to call the compilation file to generate a corresponding binary flag for each pair of sampling points, and define three functions again inside the function to downsample the data respectively, ensuring that the number of sampling point pairs does not exceed the specified value; balance positive and negative samples, and downsample to the specified number; convert the point cloud coordinates to grid coordinates, and obtain the signed distance value data for each sampling point; Step 2.3, define the voxelization point cloud function, parse the input file path, and construct the sampled point cloud data file path; initialize an all-zero array to store the voxelized point cloud data, accept the minimum value, maximum value and resolution as input, generate the coordinate array of the three-dimensional grid point, use the KD tree to perform the nearest neighbor search, set the voxel corresponding to the nearest neighbor point to 1, compress the voxelized data, and save the point cloud data, compressed voxel data and other related information as a unified format .npz file; Step 2.4, save the data set processed in step 2.2 and step 2.3 to a determined file directory to provide a data source for model training and reconstruction; Step 3: Build a network architecture, which includes an encoder, i.e., a feature extraction network, a point embedding network, an automatic decoder, and an SDF network. Input the data set preprocessed in step 2 into the network architecture built in step 3 for training, adjust the hyperparameters and weights of the network architecture, and obtain a trained neural network model, including the following sub-steps: Step 3.1, the deep 3D convolutional network of the encoder is used to extract the spatial features of the input data layer by layer. A batch normalization layer and ReLU activation function are applied after each 3D convolutional layer to enhance the nonlinear ability and stability of the encoder; Step 3.2, construct an encoder, i.e., a feature extraction network. The encoder uses the constructed convolutional layer to extract deep features from the processed data, and finally uses the obtained feature map as one of the input data of the point embedding layer; Construct a point embedding network, which expands the point coordinate dimension and shifts the point coordinate according to the displacement vector. Then, grid sampling is performed on the feature map generated by the encoder to extract the features of the shifted point coordinates. The features of all layers and the features corresponding to each shifted point coordinate are merged along the feature dimension to form a comprehensive feature representation. Then, the shape of the features is adjusted to unify the features from different displacements and levels into a continuous vector. Finally, the position information of the original point is combined with the extracted features. Construct an automatic decoder, using the constructed seven-layer fully connected network implemented in the form of 1D convolutional layers as the automatic decoder part. Step 3.3, construct an SDF network, input the data set preprocessed in step 2 into the encoder constructed in step 3.2, use the constructed ten-layer fully connected network as the SDF network, i.e., the SDF Network part to learn the signed distance value of the sampling points, and use the automatic decoder to learn the spatial relationship between a pair of sampling points; Step 4, inputting the data used for reconstruction in step 2 into the neural network model trained in step 3, and reconstructing the shape with internal structure using the marching cube algorithm, including the following sub-steps: Step 4.1, load the parameters and thresholds of the neural network model trained in step 3 and the preprocessed data obtained in step 2.3; Step 4.2, define the initial resolution of the grid, and then use the step-by-step refinement operation to increase the resolution of the grid. In each level of refinement, calculate the grid center; Step 4.3, using the sign of the signed distance value of each point, taking the points with opposite signs as a pair of points, and using the automatic decoder to predict the binary signs of the pair of points; In step 4.4, the SDF network of the model is used to predict the signed distance value corresponding to each point in the grid, and the points involved in the reconstruction algorithm are determined according to the set threshold to reconstruct the three-dimensional shape with internal structure.
2. The hierarchical structure 3D reconstruction method based on automatic decoder according to claim 1, characterized in that: In step 3.1, the deep 3D convolutional network of the encoder is used to extract the spatial features of the input data layer by layer. The sub-steps of applying a batch normalization layer and a ReLU activation function after each 3D convolutional layer are: For the first convolution layer, the input channel is 1, the output channel is 16, the kernel size is 3, and the edge padding is 1 to maintain the spatial size. Then there are two consecutive convolution layers, the first one is increased to 32 channels, the second one is maintained at 32 channels, the next two convolution layers increase the channel dimension to 64, the next two convolution layers increase the channel dimension to 128, and the last six convolution layers maintain the channel dimension at 128, and the maximum pooling layer is applied to the first, third, fifth, seventh, and ninth convolution layers respectively. The maximum pooling layer is used after the first convolution layer to halve the size to 128; The size is halved to 64 after the third layer, 32 after the fifth layer, 16 after the seventh convolutional layer, and 8 after the ninth convolutional layer; A fully connected layer is built after the convolutional layer. First, a ten-layer fully connected network is built. In the first fully connected layer, the features are expanded to twice the hidden dimension. The second fully connected layer reduces the feature dimension to the hidden dimension. The third and subsequent layers maintain the hidden dimension, and the input layer data is input for the second time in the fifth layer. Finally, the output layer outputs the signed distance value predicted by the model. Then, a seven-layer automatic decoder with fully connected layers implemented in the form of 1D convolutional layers is built. The first layer processes the input features and maps them to twice the hidden dimension. The middle three layers maintain the hidden dimension, and the output layer outputs the label value predicted by the model.
3. The hierarchical structure 3D reconstruction method based on automatic decoder according to claim 1, characterized in that: The sub-steps of reconstructing the three-dimensional shape with internal structure in step 4.4 are: The marching cubes algorithm is used to determine the boundaries of the grid based on binary flags, the surface of the grid based on signed distance values, and finally the generated 3D shape is saved to a unified .obj file format.
Citation Information
Patent Citations
Method for reconstructing three-dimensional structured model based on any visual angle pictures
CN113077554A
Three-dimensional surface reconstruction method based on multi-scale space fast Fourier coding
CN118154785A