A Method and Computer Device for a 3D Reconstruction Network Model Based on Graph Convolution
By constructing the network model based on graph convolution, generating voxel models and deformation of graph convolution neural networks, the problem of generating non-fixed structured three-dimensional grids in the existing technology is solved, and high-precision three-dimensional grid reconstruction is realized, which can restore object texture and line details.
Patent Information
- Application Number
- CN202110876136.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-07-30
AI Technical Summary
In the existing three-dimensional reconstruction technology, ordinary convolutional layers perform well when processing regular image data, but they are powerless to use data with non-fixed structures. The process of generating three-dimensional grids in two-dimensional images is difficult, especially when the object structure is large and the initial structure is initialized, the reconstruction effect is not ideal.
The three-dimensional reconstruction network model based on graph convolution is used. By generating voxel models, converting them into preliminary three-dimensional grid models, and using graph convolution neural networks for deformation, combining initial residual connections and identity mapping graph convolution networks, the excessive smoothing problem during deformation is solved, and a three-dimensional grid with object texture and line details are generated.
It realizes the generation of a high-precision three-dimensional grid model for two-dimensional pictures with non-fixed structures, simplifies the image vertex feature acquisition process, solves the problem of excessive smoothing during deformation, and can restore the detailed characteristics of the object.
Smart Images

Figure CN113658323B_ABST
Abstract
Description
Technical Field
[0001] It relates to the field of 3D reconstruction, and specifically relates to a method and a computer device for a 3D reconstruction network model based on graph convolution. Background Art
[0002] 3D reconstruction has important practical significance. For example, in medical treatment, medical images can be reconstructed into corresponding 3D structures to facilitate doctors' diagnosis of information; in the research of ancient cultural relics, 3D digital restoration technology can help researchers restore the 3D shapes of ancient buildings and ancient porcelain. These applications all illustrate the practical significance of 3D reconstruction. There are various representation forms of 3D in computers, among which meshes can better display the details of objects. A mesh contains the coordinate information, edge information, and face information of a 3D object. This representation form makes it easier to see the structure of the object, but it is more difficult for computers to process.
[0003] There is a large amount of irregular data in 3D models. For example, the data structures of vertices and edges in 3D meshes. Ordinary convolutional layers perform well in processing regular image data, but are powerless for data with non-fixed structures.
[0004] The process of generating a 3D mesh from a 2D image is very difficult. Therefore, in many methods, an initialized mesh structure is used, such as a cube mesh or an ellipsoid mesh, and these initialized structures are deformed into ideal structures. However, the reconstruction effect of such methods is not ideal when the object structure has a large gap from the initialized structure. Summary of the Invention
[0005] The present invention solves the problem that in the existing 3D reconstruction technology, ordinary convolutional layers perform well in processing regular image data, but are powerless for data with non-fixed structures, and the process of generating a 3D mesh from a 2D image is very difficult. The method provided in this application can use a single image to restore a 3D mesh model with details such as object texture and lines. The solution adopted in this application is as follows:
[0006] A method for a 3D reconstruction mesh model based on graph convolution, the method includes:
[0007] Steps of generating a voxel model according to image features;
[0008] Steps of converting the voxel model into a preliminary 3D mesh model;
[0009] Steps of collecting image vertex features;
[0010] Steps of deforming the preliminary 3D mesh model into a new 3D mesh.
[0011] Further, the step of generating a voxel model based on image features is specifically as follows:
[0012] Generate a voxel model by learning image features from convolutional layers.
[0013] Further, the step of generating a voxel model based on image features further includes:
[0014] The step of collecting two-dimensional features of each vertex of the image;
[0015] The step of collecting three-dimensional features of each vertex of the image.
[0016] Further, the step of converting the voxel model into a preliminary three-dimensional mesh model is specifically as follows:
[0017] Convert the voxel model into the preliminary three-dimensional mesh according to the cubification method.
[0018] Further, the step of collecting vertex features of the image is specifically as follows: including:
[0019] The step of splicing two-dimensional features and three-dimensional features of each vertex of the image;
[0020] The step of obtaining features of each vertex.
[0021] Further, the step of collecting two-dimensional features of each vertex of the image is specifically as follows:
[0022] Obtain two-dimensional features of each vertex of the image according to the bilinear interpolation algorithm.
[0023] Further, the step of collecting three-dimensional features of each vertex of the image is specifically as follows:
[0024] Obtain three-dimensional features of each vertex of the image according to the trilinear interpolation algorithm.
[0025] Further, the step of deforming the preliminary three-dimensional mesh model into a new three-dimensional mesh is specifically as follows:
[0026] Deform the preliminary three-dimensional mesh model according to the graph convolutional neural network.
[0027] Further, the way of deforming the preliminary three-dimensional mesh model according to the graph convolutional network with initial residual connection and identity mapping is specifically as follows: through the formula:
[0028]
[0029] where H (0) represents the initial feature information of nodes in the graph, l represents the number of layers, H is the feature of each layer, α l and βl represent two hyperparameters, σ represents the non-linear function ReLU, represents A + u, A represents an N×N-dimensional matrix formed based on the relationships between each node, N is a constant, and u is the identity matrix, represents the degree matrix of, I n represents the identity mapping, and W represents the weight matrix.
[0030] A computer device includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method of a three-dimensional reconstruction network model based on graph convolution described in this application.
[0031] The advantages of this application are as follows: The present invention realizes the effect of generating a three-dimensional mesh for a two-dimensional picture with a non-fixed structure. The specific effects include:
[0032] 1. First, learn the image features according to the convolutional layer and generate a rough voxel model. Convert the voxel model into a rough three-dimensional mesh model according to certain rules, and deform the mesh according to the graph convolutional neural network to output a more accurate three-dimensional mesh model, which can use a single image to recover a three-dimensional mesh model with details such as object texture and lines.
[0033] 2. Collect the two-dimensional and three-dimensional features of the image during the process of learning the image features according to the convolutional layer and generating a rough voxel model, which simplifies the process of collecting the vertex features of the image.
[0034] 3. Deform the initial three-dimensional mesh model through the graph convolutional network with initial residual connection and identity mapping, which not only utilizes the initial features of the nodes, but also adds the weight matrix and the identity matrix to solve the problem of "over-smoothing" during the deformation process.
[0035] The present invention is applicable to the work of three-dimensional reconstruction of images in fields such as medical treatment, virtual reality, and ancient cultural relic research. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a structural diagram of a method for a three-dimensional reconstruction grid based on graph convolution proposed in this application.
[0037] Figure 2 is a schematic diagram of projecting image features to corresponding grid vertices mentioned in Embodiment 4.
[0038] Figure 3 is a result diagram of the reconstructed grid instance mentioned in Embodiment 12. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] Embodiment 1. Refer toFigure 1 This embodiment provides a method for a three-dimensional reconstruction grid model based on graph convolution. The method includes:
[0040] The step of generating a voxel model according to image features;
[0041] The step of converting the voxel model into a preliminary three-dimensional grid model;
[0042] The step of collecting image vertex features;
[0043] The step of deforming the preliminary three-dimensional grid model into a new three-dimensional grid.
[0044] Refer to Figure 1 , Figure 1 shows the overall model of the reconstructed grid. The input is a single image. First, the ResNet-50 network is used to learn image features. Further, three-dimensional features are extracted through a three-dimensional convolution module, and a rough voxel model is generated. Then, the voxel model is converted into a three-dimensional grid model. During the process of deforming the three-dimensional grid, the two-dimensional and three-dimensional features learned by the convolutional layer are projected onto the grid vertices, and a graph convolutional neural network with an identity mapping is used to deform the three-dimensional grid, and finally the output result is obtained.
[0045] Embodiment 2: This embodiment further limits the method for a three-dimensional reconstruction grid model based on graph convolution provided in Embodiment 1. The step of generating a voxel model according to image features is specifically:
[0046] Generate a voxel model according to the image features learned by the convolutional layer.
[0047] This embodiment generates a rough voxel model according to the input image and generates an initialized grid based on the voxel model, which can roughly reflect the structure of the object.
[0048] Embodiment 3: This embodiment further limits the method for a three-dimensional reconstruction grid model based on graph convolution provided in Embodiment 2. The step of generating a voxel model according to image features further includes:
[0049] The step of collecting two-dimensional features of each vertex of the image;
[0050] The step of collecting three-dimensional features of each vertex of the image.
[0051] Embodiment 4: This embodiment further limits the method for a three-dimensional reconstruction grid model based on graph convolution provided in Embodiment 1. The step of converting the voxel model into a preliminary three-dimensional grid model is specifically:
[0052] Convert the voxel model into the preliminary three-dimensional grid according to the cubification method.
[0053] For example, when processing an image of 137×137, the specific process is as follows: First, the image of size 137×137 extracts features at four levels through the ResNet-50 network, and the feature sizes are: 256×56×56, 512×28×28, 1024×14×14, 2048×5×5. These two-dimensional features are saved for subsequent use. Further, the last feature is decoded, and three-dimensional features are extracted from it using three-dimensional convolution. The size of the extracted three-dimensional features after splicing is: 36×32×32×32, and a voxel model is generated using these three-dimensional features. The above process not only undertakes the task of reconstructing voxels, but also the image features extracted by it will be projected onto each vertex of the grid, facilitating the graph convolutional network to update the coordinates of the vertices using the vertex features.
[0054] Refer to Figure 2 , Figure 2 which shows the process of two-dimensional feature projection. The vertex coordinates of the three-dimensional grid are relative to the camera coordinate system. The first thing to do is to convert the points {v c} in the camera coordinate system to the pixel coordinate system:
[0055] v = Kv c
[0056] where K represents the internal parameter matrix of the camera, and v represents the coordinates in the pixel coordinate system.
[0057] Next, the two-dimensional feature of this point is obtained using the method of bilinear interpolation:
[0058]
[0059] where f 1,1 represents the feature at the upper left corner of v i '. Different f i,j are defined similarly. f 1,2 represents the feature at the lower left corner of v i ', f 2.1 represents the feature at the upper right corner of v i ', and f 2.2 represents the feature at the lower right corner of v i '. (x, y) represents the vertex coordinates, and the number of two-dimensional features corresponding to each point depends on the feature dimension learned by the previous image encoding layer.
[0060] Similarly, the three-dimensional feature of the vertex can also be obtained using the method of trilinear interpolation. The difference is that both the three-dimensional feature and the grid vertex are in the camera coordinate system, and direct interpolation can be performed:
[0061]
[0062] Among them, f i,j,k represents the feature at the corresponding position, k represents the coordinate in the third direction of the feature of v i ', and z represents the coordinate in the third direction of the vertex coordinates.
[0063] For points that exceed the pixel boundary or the grid boundary, their features are set to 0.
[0064] Then, the two-dimensional features and three-dimensional features of each vertex are concatenated to obtain the final feature of each vertex. In this way, the feature of each vertex contains two-dimensional features and three-dimensional features at different scales, and the deformable mesh module can utilize more information.
[0065] Furthermore, a mesh is generated from the voxels according to the cubification method. The positions where the voxel probability is greater than the threshold t are all changed into cubes composed of three-dimensional meshes. Each cube consists of 8 vertices, 18 edges, and 12 faces. In order to form the mesh surface, the faces existing inside the mesh also need to be eliminated. This algorithm meets the requirements for the speed of generating the mesh and is superior to another commonly used algorithm: the Marching Cubes algorithm. This method generates the mesh by calculating the isosurface, and the visualization result is good. However, due to its large computational amount, the speed of forming the mesh is relatively slow and it is not applicable to the network of this embodiment.
[0066] In this way, an initialized mesh is generated. The number of vertices of this mesh is closely related to the size of the voxel resolution. The rough voxel resolution generated by this embodiment is 32×32×32, and the number of generated vertices ranges from 600 to 1400. The subsequent network deforms the mesh by moving these vertices, and during this process, the relationship between the edges and faces of the vertices is maintained.
[0067] Embodiment 5: This embodiment further limits a method for a three-dimensional reconstruction mesh model based on graph convolution provided in Embodiment 3. The step of collecting the vertex features of the image specifically includes:
[0068] The step of concatenating the two-dimensional features and three-dimensional features of each vertex of the image;
[0069] The step of obtaining the features of each vertex.
[0070] Embodiment 6: This embodiment further limits a method for a three-dimensional reconstruction mesh model based on graph convolution provided in Embodiment 3. The step of collecting the two-dimensional features of each vertex of the image specifically is:
[0071] Obtaining the two-dimensional features of each vertex of the image according to the bilinear interpolation algorithm.
[0072] Embodiment 7. This embodiment further limits a method for a 3D reconstruction mesh model based on graph convolution provided in Embodiment 3. The step of collecting the 3D features of each vertex of the acquired image is specifically as follows:
[0073] Obtain the 3D features of each vertex of the image according to the trilinear interpolation algorithm.
[0074] Embodiment 8. This embodiment further limits a method for a 3D reconstruction mesh model based on graph convolution provided in Embodiment 5. The step of deforming the preliminary 3D mesh model into a new 3D mesh is specifically as follows:
[0075] Deform the preliminary 3D mesh model according to the graph convolutional network with initial residual connection and identity mapping.
[0076] Embodiment 9. This embodiment further limits a method for a 3D reconstruction mesh model based on graph convolution provided in Embodiment 9. The way of deforming the preliminary 3D mesh model according to the graph convolutional network with initial residual connection and identity mapping is specifically as follows: Through the formula:
[0077]
[0078] where, H (0) represents the initial feature information of the nodes in the graph, l represents the number of layers, H is the feature of each layer, α l and β l represent two hyperparameters, σ represents the non-linear function ReLU, represents A + u, A represents a matrix formed based on the relationships between each node, which is an N×N-dimensional matrix, N is a constant, and u is the identity matrix, represents the degree matrix of, I n represents the identity mapping, and W represents the weight matrix.
[0079] The deformation formula of the classical graph convolutional neural network is to deform the grid of the graph convolutional neural network (GCN) with eigen mapping and residual connection, and output the final grid result. The deformation formula of the classical graph convolutional neural network is:
[0080]
[0081] The above-mentioned graph convolutional network can make good use of the associations between nodes to learn features. However, GCN cannot be directly applied to deep structures. Relevant experiments have shown that GCN performs best when it has two layers. Stacking more layers above this actually makes the performance of the model worse. This problem is called over-smoothing. And shallow GCNs cannot extract high-order features of nodes. Using strategies similar to residual connections can only alleviate the problem of over-smoothing and cannot make deep structures perform better than two-layer structures. Based on the above considerations, this embodiment contemplates a graph convolutional network with initial residual connections and identity mapping, and its iterative formula can be written as:
[0082]
[0083] The initialized residual connection ensures that the final representation of each node retains part of the information from the input layer. This formula not only utilizes the initial features of the nodes but also adds the weight matrix to the identity matrix. As the number of network layers deepens, the performance of this model improves and it outperforms traditional shallow graph convolutional models.
[0084] Specifically, this embodiment uses graph convolutional layers in the case of a grid to learn features in combination with the above strategy. The number of graph convolutional layers is 12. Every four layers, the vertex positions are updated using the features, and the image features are re-projected based on the new positions. Finally, the vertex coordinates in the grid and the vertex indices corresponding to the edges in the grid are output.
[0085] Embodiment Ten. This embodiment provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes a method for a three-dimensional reconstruction grid model based on graph convolution according to any one of Embodiments One to Nine.
[0086] Embodiment Eleven. This embodiment provides a method for calculating the loss function of a grid during the process of generating a voxel model. The calculation method is as follows:
[0087] Through the formula:
[0088]
[0089] where N represents the total number of voxel blocks and p represents the predicted probability value.
[0090] The formula is: During the process of training voxels, the loss function uses the cross-entropy loss function in the case of voxels;
[0091] In the stage of training the grid, this embodiment uses the difference between the points of the target grid and the predicted grid to define the loss function. First, 5000 points are sampled from the surfaces of the target grid and the predicted grid. Let P represent the set of target point clouds and Q represent the set of predicted point clouds. Calculate the nearest vertices between them:
[0092] Λ P,Q ={(p, argmin q ||p - q||: p ∈ P)}
[0093] where p represents the midpoint of P and q represents the midpoint of Q.
[0094] The Chamfer loss between P and Q can be written as:
[0095]
[0096] where |P| and |Q| represent the number of midpoints of P and Q respectively. The Chamfer loss constrains the distance between the vertices of the two point clouds, but it does not consider the edges in the grid. Therefore, it is necessary to define the loss function of the edges:
[0097]
[0098] where V represents..., represents the set of edges in the predicted grid, and |E| represents the number of edges in the grid. Finally, the loss function of the three-dimensional grid reconstructed by the method provided in this application will be composed of the above three loss functions weighted.
[0099] Embodiment Twelve. Refer to Figure 3 To illustrate this embodiment, this embodiment is the result of an experiment and experiment on an image according to a three-dimensional reconstruction grid method based on graph convolution proposed in this application. Specifically:
[0100] This embodiment uses the ShapeNet dataset. The difference is that the finally generated is a triangular mesh model. At the same time, the camera internal parameters and external parameter matrices corresponding to different images are used during the training and testing processes.
[0101] This article adopts the F-Score evaluation criterion and gives the distance threshold d. It can be written formulaically as:
[0102]
[0103] where P(d) and R(d) represent the precision and recall rate with respect to the distance threshold d respectively.
[0104] can be calculated as:
[0105]
[0106] where n R and n G represent the numbers of the predicted point cloud and the target point cloud respectively. For the evaluation criterion of F-Score, the larger the value, the better the reconstruction effect.
[0107] This embodiment uses a single RGB image with a resolution of 137×137 as the input for training and testing. The network for extracting image features is ResNet-50 pre-trained on ImageNet. Next, 3D convolution and 3D transposed convolution are used to generate voxels with a resolution of 32×32×32, and they are converted into a triangular mesh model. 12 graph convolutional network layers are used to deform the mesh in the deformation module. It is trained for 20 epochs using the Adam optimizer, the learning rate is set to 0.0001, the batch size is set to 4, the voxel threshold is set to 0.15, and the weights of the loss function are set as: λ voxel =1, λ cham =1, λ edge =0.2. The model of this embodiment is implemented using the pytorch and pytorch3d libraries.
[0108] Table 2-1 shows the F-Score values of the reconstructed meshes of this model under different object types. Among them, the distance threshold d = 0.0001.
[0109] Table 2-1 F-Score values of the results reconstructed by the method provided in this application on the test set under different thresholds d
[0110]
[0111] It can be seen that the method provided in this application has a better reconstruction effect on the reconstruction of gun and aircraft types for the 3D mesh model. At the same time, Table 2-2 compares the results reconstructed by the method provided in this application with those of other reconstruction methods in recent years to illustrate the effect of the results reconstructed by the method provided in this application. Among them, N3MR is a method for generating a mesh model based on weakly supervised. It can be seen from Table 2-2 that the method provided in this application has obtained better reconstruction results. It can be determined through the index that the results reconstructed by the method provided in this application are improved compared with those of several algorithms in the table under different distance thresholds d. The index is the F-score evaluation index, and the higher the value, the better the reconstruction effect.
[0112] Table 2-2 Comparison of F-Score values of the results reconstructed by the method provided in this application with those of other models under different thresholds d
[0113]
[0114] Figure 3 Several examples of reconstruction by the method provided in this application are shown. It can be seen that the model reconstruction grid can obtain results similar to the target grid and can describe detailed features such as textures and lines.
[0115] Among them, from left to right are the input image, the target grid, the predicted grid, and the predicted grid of the model after removing the L edge loss function.
[0116] Related research shows that removing the L edge loss function can make the evaluation criteria perform better, and this embodiment also conducts experiments on this. As shown in Table 2-3, the results of setting the weight of L edge to 0 and keeping other settings unchanged are shown. Figure 3 The comparative experimental results after removing the L edge loss function are also shown.
[0117] It can be seen that the evaluation index does become better when the L edge loss function is removed. However, this loss constrains the length of the edges. If this loss function is removed, many irregular and repetitive faces appear, and these meshes will be displayed unclearly. Therefore, when reconstructing the 3D mesh model by the method provided in this application, the L edge loss function is still added to train the model in order to generate triangular meshes with better sensory effects.
[0118] Table 2-3 F-Score values on the test set of the results reconstructed by the method provided in this application under different thresholds d after removing the L edge loss
[0119]
[0120] This embodiment first analyzes the characteristics of the triangular mesh representation method. By the method provided in this application, the input 2D features and 3D features are first obtained through an encoding layer based on ResNet-50 and 3D convolution, and the reconstructed voxel model is efficiently converted into a triangular mesh by the cubing method. The features learned by the encoding layer are projected onto the corresponding vertices by linear interpolation. Finally, a graph convolutional network with residual connections and identity mapping is used to deform the mesh, and the network is trained by the Chamfer loss and an edge loss function, taking into account the relationship of the edges in the mesh while constraining the vertex distance. The data of the model on the test set shows that the triangular mesh reconstructed by the method provided in this application can show detailed features such as the lines and textures of the object, and the visualization effect is clearer than that of the voxel.
Claims
1. A method for a three-dimensional reconstruction grid model based on graph convolution, characterized in that The method comprises: The step of generating a voxel model according to image features; The step of converting the voxel model into a preliminary three-dimensional mesh model; The step of collecting image vertex features as initial feature information of nodes; The step of deforming the preliminary three-dimensional mesh model into a new three-dimensional mesh; The deformation of the preliminary 3D mesh model is achieved by the formula: Among them, H (0) represents the initial feature information of the nodes in the figure, l represents the number of layers, and H represents the features of each layer. α l and β l represent two hyperparameters, σ represents the non-linear function ReLU, represents A + u, where A represents an N×N-dimensional matrix formed based on the relationships between each node, N is a constant, and u is the identity matrix. represents the degree matrix of, and I n represents the identity mapping, and W represents the weight matrix.
2. The method for a three-dimensional reconstruction grid model based on graph convolution according to claim 1, wherein, The steps of generating a voxel model based on image features are specifically as follows: Generate voxel models based on convolutional layer learning of image features.
3. The method for a three-dimensional reconstruction grid model based on graph convolution according to claim 2, wherein The step of generating a voxel model according to image features further includes: The step of collecting two-dimensional features of each vertex of the image; The step of collecting three-dimensional features of each vertex of the image.
4. The method for three-dimensional reconstruction of a mesh model based on graph convolution according to claim 1, characterized in that: The steps of converting the voxel model into a preliminary three-dimensional mesh model are specifically as follows: The voxel model is converted into the preliminary three-dimensional mesh according to a cubosome method.
5. The method for a 3D reconstruction mesh model based on graph convolution according to claim 3, wherein The steps of collecting image vertex features specifically include: A step of splicing the two-dimensional features and the three-dimensional features of each vertex of the image; Steps to get the features of each vertex.
6. A method for a three-dimensional reconstruction grid model based on graph convolution according to claim 3, characterized in that The steps of collecting the two-dimensional features of each vertex of the image are specifically as follows: The two-dimensional features of each vertex of the image are obtained according to a bilinear interpolation algorithm.
7. A method for a three-dimensional reconstruction mesh model based on graph convolution according to claim 3, characterized in that The steps of collecting the three-dimensional features of each vertex of the image are specifically as follows: The three-dimensional features of each vertex of the image are obtained according to a trilinear interpolation algorithm.
8. A method for a three-dimensional reconstruction grid model based on graph convolution according to claim 5, characterized in that The steps of transforming the preliminary three-dimensional mesh model into a new three-dimensional mesh are specifically as follows: The preliminary 3D mesh model is deformed according to a graph convolutional network with initial residual connections and identity mapping.
9. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes a method for reconstructing a three-dimensional network model based on graph convolution according to any one of claims 1 to 8.
Citation Information
Patent Citations
Three-dimensional model reconstruction method and device, equipment and storage medium
CN111369681A