A 3D Reconstruction Method Based on Integrating Multi-Perspective Features and Deep Learning
Through the combined full convolution network and point convolution technology combined with multi-view features, a three-dimensional model is generated, which solves the problems of low resolution of three-dimensional model reconstruction and incomplete intermediate representation in the existing technology, and achieves a more efficient and robust three-dimensional model reconstruction effect.
Patent Information
- Application Number
- CN202210219791.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-08
AI Technical Summary
The prior art is difficult to effectively combine multi-view features and deep learning in three-dimensional model reconstruction, resulting in limited resolution of the generative model, and the intermediate representation cannot completely preserve the adjacency relationship between model elements, and there are noise and diversity robustness problems.
A full convolution network is used to convolutionize multi-view image data, extract global and local feature information, and generate the final feature information through point convolution and voting mechanisms. Finally, a three-dimensional model is generated using the Marching Cube algorithm.
It improves the resolution and integrity of three-dimensional model reconstruction, enhances the robustness to noise and diversity, simplifies the network structure, improves training accuracy, and improves the effect of multi-view image fusion.
Smart Images

Figure CN114708380B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a 3D reconstruction method, in particular to a 3D reconstruction method based on the fusion of multi-view features and deep learning. Background Art
[0002] In recent years, with the emergence of more and more 3D modeling software and the wide application of depth sensors such as Kinect on platforms for collecting depth data, 3D model data has shown an explosive growth on the Internet, and there have been a large number of forms of 3D models, such as point clouds, voxels, meshes, implicit fields, etc. This trend has made the analysis of 3D models a hot research field. At present, the research in the field of image analysis has achieved fruitful results, and the introduction of deep learning frameworks has further improved the results. In this context, a large amount of research has been invested in 3D reconstruction technologies based on monocular, binocular, and multi-view images. However, convolutional operations on 2D images cannot be directly applied to 3D models, making it extremely difficult to apply deep learning methods to the analysis of 3D models. Therefore, a large number of 3D model analysis methods rely on manually tuned descriptors to extract features. Although recently there have been data structures such as trees and graphs that organize 3D model data into intermediate representations, making convolutional operations feasible, it is very difficult for such structures to completely maintain the original adjacency relationships between meshes or points. At the same time, due to the large amount of point information and feature information involved in the research of the 3D field, it is also very difficult to effectively improve the resolution of the generated model.
[0003] It can be seen that the problem of 3D model reconstruction is very challenging for the following reasons:
[0004] 1. Prediction of invisible parts during reconstruction, such as the bottom of the model that cannot be clearly shown in the picture, etc.;
[0005] 2. Accurately detecting the detailed parts of the model components, such as small patterns and holes on the model, which require more subtle geometric information;
[0006] 3. Local and global features must be combined for analysis to achieve better generation results;
[0007] 4. The analysis method must be robust to noise, downsampling, and the diversity of similar models;
[0008] In recent years, the field of 3D model reconstruction has developed vigorously, emerging two major categories: unsupervised traditional methods and supervised data-driven methods.
[0009] Unsupervised methods, such as Literature 1: Furukawa Y, Ponce J. Accurate, Dense, and Robust Multiview Stereopsis[J]. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 2010, 32(8): p. 1362-1376., Literature 2: Vu H H, Labatut P, Pons J P, et al. High Accuracy and Visibility-Consistent Dense Multiview Stereo[J]. IEEE Transactions on Pattern Analysis & Machine Intelligence, 2011, 34(5): 889-901., Literature 3: Galliani S, Lasinger K, Schindler K. Massively Parallel Multiview Stereopsis by Surface Normal Diffusion[C] / / IEEE International Conference on Computer Vision. IEEE Computer Society, 2015., Literature 4: Schnberger J L, Zheng E, Pollefeys M, et al. Pixelwise View Selection for Unstructured Multi-View Stereo[C] / / European Conference on Computer Vision (ECCV). Springer, Cham, 2016., etc., rely on the prior knowledge of existing model sets to perform multi-view model reconstruction work. Most of these traditional unsupervised methods adopt operations such as manually increasing similarity weights, that is, they require a large amount of additional work and have many problems in integrity and do not have good generalization ability.
[0010] Supervised methods extract features from labeled training datasets and then use the trained model to perform reconstruction work on the test dataset. Depending on the different intermediate and output representations of the reconstruction model, it also has a great impact on the training of the final network and the effect of the generative model. For example, in the field of single-view image reconstruction, Literature 5: Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, and Honglak Lee. Perspective transformernets: Learning single-view 3d object reconstruction without 3d supervision. In NeurIPS, 2016. uses voxels as the representation of the model. Literature 6: Christian Shubham Tulsiani, and Jitendra Malik. Hierarchical surface prediction for 3d object reconstruction. In 3DV, 2017. Use octree as the representation of the model. Literature 7: Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3d object reconstruction from a single image. In CVPR, 2017. Use point cloud as the representation of the model. Literature 8: Chuhang Zou, Ersin Yumer, Jimei Yang, Duygu Ceylan, and Derek Hoiem. 3d-prnn: Generating shape primitives with recurrent neural networks. In ICCV, 2017. Use polyhedron as the representation of the model. Literature 9: Ayan Sinha, Asim Unmesh, Qixing Huang, and Karthik Ramani. Surfnet: Generating 3d shape surfaces using deep residual networks. In CVPR, 2018. Use 3d surface as the representation of the final model. Literature 10: Jiapeng Tang, Xiaoguang Han, Junyi Pan, Kui Jia, and Xin Tong. A skeleton-bridged deep learning approach for generating meshes of complex topologies from single rgb images. In CVPR, 2019. Use the skeleton of the object as the representation for generating the final model. Literature 11: Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan Russell, and Mathieu Aubry. AtlasNet: A Papier- Approach to Learning 3D Surface Generation. In CVPR, 2018. By customizing a set of methods for the model surface as the model representation, etc., these methods convert 3D models into intermediate structures such as voxels, trees, graphs, spectral spaces, etc., and then apply deep learning frameworks to predict the labels of model patches or point clouds, achieving good results. However, they still do not solve the problem that the intermediate form cannot completely preserve the adjacency relationship between model elements, and are limited by the above intermediate representation forms when generating the final results. The resolution of the generated results is limited, and the generation and storage of voxel and point cloud representation forms consume a large amount of memory and are not practical enough. Literature 12: Xu Q, Wang W, Ceylan D, et al. DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction[J]. 2019. proposed using the implicit field representation as the encoding method for 3D models, that is, it is agreed that the part on the model surface is 0, less than 0 inside the model, and greater than 0 outside the model. Process the 3D model in this way to generate the corresponding encoding method, then design a network to learn the global and local features of the model, and then add the two features to obtain the final feature vector. Use the encoding obtained in the previous agreed way as the ground truth to calculate the L1 Loss, and use this to carry out training. The implicit field representation solves the discretization problem of the point cloud representation, and also solves the problem of insufficient resolution of voxel, mesh and other representation forms. In addition, this representation form has continuity and is suitable for training in convolutional networks. At the same time, compared with other representation forms, CD and EMD are used as the calculation methods for loss. This literature simplifies the loss calculation method and increases the efficiency of the training process, and finally also achieves good results in the field of single-view image-based reconstruction. However, limited by the input of a single image, the information of the model cannot be fully obtained. For a single image, the model has different degrees of occlusion. Therefore, the integrity of the finally generated model depends largely on the angle of the selected single image. The result of inputting a single front-view image is often better than that of inputting a single side-view image, which also makes the finally generated result uncertain.
[0011] Currently, in the field of 3D model reconstruction based on images, there are methods such as monocular, binocular, and multiocular. The above-cited reference 12 uses a monocular perspective, and the generated model is also limited by the angle of the input image, resulting in significant differences in the generated results. Reference 13: Han, X., Leung, T., Jia, Y., Sukthankar, R., Berg, A.C.: Matchnet: Unifying feature and metric learning for patch-based matching. Computer Vision and Pattern Recognition (CVPR) (2015) uses a binocular view and generates a model using patch matching, but it is also limited by insufficient information, making the generated model not achieve very good results. However, 3D reconstruction based on a multiocular perspective is not simply an increase in images. Due to the randomness of the angles of the input images, how to combine the information of multiple images becomes a challenge. Reference 14: Kar A, Hne C, Malik J. Learning a Multi-View Stereo Machine[J]. 2017. uses the method of projecting 2D images back into 3D space using camera parameters, fuses the 3D feature vectors of multiple perspectives through an RNN network to form a 3D cost volume, and then performs optimization operations using 3D convolution. Finally, the 3D final result is obtained. However, since the intermediate representation forms are all in the form of voxels, it will incur a large cost during the training process, and the generated results are also limited by the voxel representation method. Reference 15: Yao Y, Luo Z, Li S, et al. MVSNet: Depth Inference for Unstructured Multi-view Stereo[J]. 2018. First, feature extraction is performed on 2D images to obtain feature maps, and then a 3D cost volume is constructed based on the camera frustum of the reference view through differentiable homography transformation. Then, 3D convolution is used to regularize the cost volume, and the initial depth map is regressed. The initial depth map is then optimized through the reference image to obtain the final depth map.Since Document 15 has the problem of too many parameters in the process of optimizing the cost volume, which seriously consumes memory and makes the method have poor scalability and unable to be applied to high-resolution scenarios. On this basis, Document 16: Yao Y, Luo Z, Li S, et al. Recurrent MVSNet for High-resolution Multi-view Stereo Depth Inference[J]. IEEE, 2019. In the process of cost volume optimization, uses multi-layer gated recurrent unit (GRU) instead of 3D convolution, reducing the memory consumption from cubic growth to quadratic growth, and can be applied to high-resolution scenarios. However, the method based on MVSNet is also limited by the voxel representation method, resulting in the inability to flexibly generate multi-resolution models as needed. Summary of the Invention
[0012] Object of the Invention: The technical problem to be solved by the present invention is to provide a three-dimensional reconstruction method based on fusing multi-view features and deep learning in view of the deficiencies of the prior art.
[0013] To solve the above technical problems, the present invention discloses a three-dimensional reconstruction method based on fusing multi-view features and deep learning, including the following steps:
[0014] Step 1, collect multi-view image data and corresponding signed distance function (SDF) data as the input image data.
[0015] Step 2, perform convolution operations on the input image data with the fully convolutional network vgg to obtain one-dimensional global feature information and multi-layer two-dimensional image feature information under multiple views respectively.
[0016] Step 3, perform max pooling operations on the global feature information under multiple views in Step 2 to obtain the overall global feature information.
[0017] Step 4, input the point information of the three-dimensional model, that is, the three-dimensional query point, project it to obtain the corresponding two-dimensional query point, query the multi-layer feature information of the image with the two-dimensional query point to obtain local point information, and splice the multi-layer local point information to obtain local feature information.
[0018] Step 5, perform point convolution operations (reference: A network for extracting 3D point features, reference: (Qi CR, Su H, Mo K, et al. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation[C] / / 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017.) to obtain corresponding point features;
[0019] Step 6, combine the point features obtained in Step 5 with the overall global feature information obtained in Step 3 and the local feature information obtained in Step 4 respectively to obtain global features and local features with point feature information;
[0020] Step 7, combine the global features with point feature information in Step 6 with the local features with point feature information respectively, and generate the final feature information by voting;
[0021] Step 8, for the final feature information obtained in Step 7, generate the final 3D model through the marching cubes algorithm (reference: Marching Cube, reference: William E Lorensen and Harvey E Cline. Marching cubes: A high resolution 3d surface construction algorithm. In ACM siggraph computer graphics, 1987.).
[0022] Step 1 of the present invention includes the following steps:
[0023] Step 1-1: Obtain the original 3D dataset D from the Shapenet database (an official 3D data test set for training and testing, containing data required for 10 patents and used for subsequent testing and training), and obtain the 2D image information P of 24 viewpoints and the corresponding camera parameter information C from the 3DR2N2 dataset (a dataset obtained by projecting the 3D dataset in ShapeNet. Reference: Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In ECCV, 2016. It contains 2D image information and corresponding camera parameter information).
[0024] Step 1-2: Generate the corresponding SDF data from the 3D dataset D and store it in the corresponding h5 file.
[0025] Step 1-3: Store the 2D image information P and the corresponding camera parameter information C in the h5 file.
[0026] Step 1-2 in the present invention includes the following steps:
[0027] Step 1-2-1: Normalize the 3D dataset D to generate the dataset N such that the numerical range of the dataset is between -1 and 1.
[0028] Step 1-2-2: Use the isosurface tool (a tool for processing 3D data into SDF data. Reference: FunShing Sin, Daniel Schroeder, and Jernej Barbiˇc. Vega: non-linear fem deformable object simulator. In Computer Graphics Forum, volume 32, pages 36–48. Wiley Online Library, 2013.) to process the dataset N into an SDF dataset.
[0029] Step 1-2-3, using the Marching Cube method (Reference: William E Lorensen and Harvey ECline. Marching cubes: A high resolution 3d surface construction algorithm. In ACM siggraph computer graphics, 1987.), process the SDF dataset to generate the corresponding three-dimensional dataset O;
[0030] Step 1-2-4, store the SDF data information in the dataset O into the h5 file.
[0031] In the present invention, Step 2 includes the following steps:
[0032] Step 2-1, randomly divide the input three-dimensional mesh model dataset D = {D Train , D Test} into a training set D Train = {d 1 , d 2 , … d i , …, d n} and a test set D Test = {d n+1 , d n+2 , …, d n+j , …, d n+m} in equal numbers, where d i represents the i-th model in the training set, and d n+j represents the j-th model in the test set, n represents the number of the training set, and m represents the number of the test set;
[0033] Step 2-2, for the training set D Train , collect the projection maps P Train = {P 1 , P 2 , … P i , …, p n} in each viewing direction and the camera parameter information C Train = {C 1 , C 2 , … C i , …, c n} and input them into the fully convolutional network for training to obtain the trained fully convolutional network that fuses multi-view features, and at the same time obtain the multi-view multi-layer two-dimensional picture feature information L and the one-dimensional global feature G, where P i refers to the projection map set of the 24 viewing directions of the i-th model d Train in the training set D i , and C iRefers to the training set D Train The i-th model d in i The camera parameter set of 24 perspectives.
[0034] Step 2-2 includes the following steps:
[0035] Step 2-2-1, input the projection map P under each perspective of the training set Train , where for each model P i Corresponding to the projection map information set and camera parameter information set C of 24 perspectives i , select the pictures p of n perspectives from it extract ={p i , p m , …, p k} and input them into the vgg convolutional network, where the value range of n is from 1 to 24;
[0036] Step 2-2-2, input the selected pictures p of n perspectives extract into the vgg convolutional network (a network that can effectively extract picture features, reference: Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. In this patent, the vgg16 network model is used). After convolutional operations and pooling operations, hierarchically extract the feature vectors of 224×224×64 dimensions, 112×112×128 dimensions, 56×56×256 dimensions, 28×28×512 dimensions, and 14×14×512 dimensions under n perspectives as the hierarchical picture features L={l 1 , l 2 , …, l n}, then expand the 14×14×512-dimensional feature vector of the last layer into 1 dimension, perform a fully convolutional operation, and finally obtain the 1×128-dimensional feature vectors under n perspectives as the global vectors G={g 1 , g 2 , …, g n}.
[0037] In the present invention, step 3 includes the following steps:
[0038] Perform a max pooling operation on the global feature vector G obtained in step 2, that is, select the maximum value in each dimension to form a 128-dimensional feature vector that fuses the features of each perspective, and obtain the overall global feature G.
[0039] In the present invention, step 4 includes the following steps:
[0040] Step 4-1: Project the input 3D query point information D using the camera parameters to obtain the 2D query point Q in the corresponding view.
[0041] Step 4-2: Use the 2D query point Q obtained in Step 4-1 to query the hierarchical image features L in Step 2 respectively, and obtain local point feature information of 1×64 dimension, 1×128 dimension, 1×256 dimension, 1×512 dimension, and 1×512 dimension respectively. Then splice them to obtain local feature information of 1×1472 dimension.
[0042] In the present invention, Step 5 includes: Input the 2D query point Q into the PointNet network (Reference: Qi CR, Su H, Mo K, et al. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation[C] / / 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017.), and extract the feature vector F of 1×512 of the point feature information.
[0043] In the present invention, Step 6 includes the following steps:
[0044] Step 6-1: Splice the global feature vector G obtained in Step 3 and the point feature vector F obtained in Step 5 to obtain a global feature with point feature information of 1×640 dimension, and then perform a convolution operation to obtain a global feature vector G of 1×256 dimension.
[0045] Step 6-2: Splice the local features L in n views obtained in Step 4 and the point feature vector F obtained in Step 5 to obtain local features with point features of 1×1984 dimension in n views, and then perform a convolution operation to obtain local features L of 1×256 dimension.
[0046] In the present invention, Step 7 includes the following steps:
[0047] Step 7-1: Add the global feature vector G obtained in Step 6 to the local features L in n views obtained in Step 6 respectively to obtain n pieces of feature information W of 1×256 dimension = {w 1 , w 2 , …, w n}
[0048] Step 7-2: Perform a voting operation on the feature W obtained in Step 7-1 to finally obtain a 1×256-dimensional feature information SDF. Among them, the voting method is to count the number of positive and negative values of the SDF at each layer at the same point for the finally obtained multi-view features, and take the average value of the larger value as the SDF prediction value of the current point;
[0049] Step 7-3: Calculate the loss value
[0050]
[0051]
[0052] where |·| represents the L1-norm, that is, the L1 norm; f(I, p) represents the network output result, that is, the network prediction value, where the input is the picture I and the 3D point p; SDF I (p) represents the true value of SDF corresponding to the point p; δ is a threshold; m represents the relative position of the penalty point in the model. When the point is inside the model, a penalty term is obtained. Take m1>m2, where m1 is the value of the penalty term when the true value is less than the threshold δ, and m2 is the value of the penalty term when the true value is greater than the threshold δ;
[0053] Step 7-4: Perform backpropagation to finally obtain a fully convolutional network trained by fusing multi-view features.
[0054] Beneficial effects: The method of the present invention is dedicated to solving the problem of generating a corresponding 3D model from multi-view pictures based on a model. First, use the vgg network to extract features from the pictures, extract features in different perceptual fields from multiple layers of the vgg as multi-layer local information, and take the features in the last perceptual field as global features. Then, combine the features of 3D information to obtain features under multiple views. Through the voting method, finally obtain the final optimal features. In the whole process, the method proposes a method of combining multi-view picture information. Different from the previous max-pooling operation, according to the characteristics of SDF data, first judge the number of positive SDF and negative SDF, select the larger category, and then take its average value as the SDF value of the current point. Doing so not only satisfies the continuity of the SDF value but also conforms to the meaning expressed by the SDF value. At the same time, placing the operation of selecting the "optimal view" after the network reduces the burden on the network and improves the training efficiency; and only performs one training on the point features, then combines them with the local features and global features respectively, and then performs feature extraction, changing the way of obtaining the point features twice and then combining them with the local features and global features respectively, simplifying the network structure and also improving the accuracy of network training. The above operations simplify the network as much as possible and also propose a reasonable way of fusing multi-view pictures, thereby further improving the effect of 3D reconstruction based on multi-view pictures. Brief Description of the Drawings
[0055] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0056] Figure 1 This is the flowchart of the present invention.
[0057] Figure 2 This is a schematic diagram of a sample input picture.
[0058] Figure 3 This is a schematic diagram of the result after three-dimensional reconstruction.
[0059] Figure 4 This is the system framework diagram of the method of the present invention.
[0060] Figure 5 This is a schematic diagram of pictures of 24 perspectives of a single model in the present invention.
[0061] Figure 6 This is a schematic diagram for comparing the reconstruction results of the method of the present invention with other methods. Specific Embodiments
[0062] The following further describes the present invention in conjunction with the accompanying drawings and embodiments.
[0063] As Figure 1 shown, the present invention discloses a method for three-dimensional model reconstruction of pictures with multiple perspectives based on a model. The present invention uses an implicit field as the representation method of the three-dimensional model, processes the three-dimensional model into a representation method of SDF values as the ground truth, and simultaneously obtains pictures of 24 perspectives of the model and corresponding camera data; inputs the multi-perspective pictures of the model training set into the vgg network to obtain the global features of the model and multi-layer local feature information under multiple perspectives; performs max pooling operations on the global features under multiple perspectives to obtain the overall global features; projects the query points onto the two-dimensional plane using the camera parameter information, and queries the corresponding features in the multi-layer local feature information to form multi-layer point feature information; inputs the three-dimensional point information of the model into the PointNet network to obtain point position feature information; splices the point position feature information with the local features and global features respectively, and performs convolution operations to obtain local information features and global information features with point position information features; adds the global information features to the local feature information under multiple perspectives respectively; compares the number of positive values and negative values of each query point in each perspective, and takes the average value of the larger number as the final SDF feature of the model; uses the method of L1 Loss to calculate the gap between the predicted SDF value and the ground truth; finally uses the Marching Cube method to generate the corresponding three-dimensional model. For a given set of image D = {D Train , D Test}, and randomly divided into the training set D Train ={d 1 , d 2 ,…d i ,…, d n} and the test set D Test ={d n+1 , d n+2 ,…, d n+j ,…, d n+m}, where d i represents the i-th model in the training set, d n+j represents the j-th model in the test set, n represents the number of training sets, and m represents the number of test sets. The present invention completes the three-dimensional reconstruction of multi-views in the test set D Test through the following steps, and the target task is as Figure 2 shown, and the flow chart is as Figure 1 and Figure 4 shown:
[0064] Specifically, it includes the following steps:
[0065] Step 1, collect multi-view image data and corresponding SDF data;
[0066] Step 2, perform convolution operations on the input image data using the vgg network to obtain one-dimensional global features and multi-layer two-dimensional image feature information under multiple views respectively;
[0067] Step 3, perform max pooling operations on the global information under multiple views to obtain the overall global information;
[0068] Step 4, project the input query point information, query the image feature information with two-dimensional query points to obtain local point information, and splice them to obtain local feature information;
[0069] Step 5, perform point convolution operations using the PointNet network to obtain the point features of the model;
[0070] Step 6, combine the point features with the global features and local features respectively to obtain global features and local features with point feature information;
[0071] Step 7, combine the global features with the local features respectively, and generate the final feature information by voting;
[0072] Step 8, generate a three-dimensional model by the Marching Cube method.
[0073] Step 1 includes the following steps:
[0074] Step 1-1, assume that a single 3D SDF model s is input, that is, the label information (in the format of.h5, recording the data of 32,768 points of the 3D model organized in the SDF manner. This 3D model s is taken from the ShapeNetCore standard 3D model dataset containing 16 types of 3D models and has been normalized and processed into SDF data) and the set of 24-view images p corresponding to the model s (in the format of h5, recording the pixel information of the images and the corresponding camera data information);
[0075] Step 1-2, randomly select n-view images from the 24 views for subsequent convolution operations;
[0076] Among them, Step 1-1 includes the following steps:
[0077] Step 1-1-1, perform normalization processing on the original 3D model o (in the format of.obj, recording the point and face information of the 3D model), so that the center point of the 3D model is centered and the value range of the 3D points is within [-1, 1]. Then perform SDF processing on the normalized 3D model, that is, make the points on the surface of the 3D model be 0, the points inside the 3D model be less than 0, and the points outside the 3D model be greater than 0. Finally, sample the information of 256×256×256 points, and write the coordinate information of each obtained point and the corresponding SDF value information into a single file fs (a file in the format of.h5);
[0078] Step 1-1-2, obtain the 24-view 2D image information P and the corresponding camera parameter information C from the 3DR2N2 dataset, and write them into the same file fp (a file in the format of.h5, recording the pixel information of the 2D image and the corresponding camera parameter information);
[0079] Step 2 includes the following steps:
[0080] Step 2-1, randomly divide the input 3D mesh model dataset D = {D Train , D Test} into a training set D Train = {d 1 , d 2 , … d i , …, d n} and a test set D Test = {d n+1 , d n+2 , …, d n+j , …, d n+m} in equal numbers, where d i represents the i-th model in the training set, d n+j represents the j-th model in the test set, n represents the number of training sets, and m represents the number of test sets;
[0081] Step 2-2, for the training set D Train , collect the projection images P Train ={P 1 , P 2 , … P i , …, p n} and the camera parameter information C Train ={C 1 , C 2 , … C i , …, c n} and input them into the fully convolutional network for training to obtain a trained fully convolutional network that fuses multi-view features, and obtain the local feature L and the global feature G under multiple views, where P i refers to the projection image set of the 24 views of the i-th model s Train in the training set D i , and C i refers to the camera parameter set of the 24 views of the i-th model s Train in the training set D i ;
[0082] Among them, step 2-2 includes the following steps:
[0083] Step 2-2-1, input the projection images P Train under each view of the training set, where each model p i corresponds to the projection image information set P i of 24 views and the camera parameter information set C i , and select n (n ranges from 1 to 24) views of pictures p extract ={p i , p m , …, p k} and input them into the vgg convolutional network;
[0084] Step 2-2-2, as Figure 4 shown, input the selected n views of pictures p extract into the vgg convolutional network, where the vgg network uses the template of vgg16 provided in pytorch. After convolutional operations and pooling operations, layer by layer, extract the feature vectors of 224×224×64 dimensions, 112×112×128 dimensions, 56×56×256 dimensions, 28×28×512 dimensions, and 14×14×512 dimensions under n views as the hierarchical picture features L={l 1 , l 2 , …, l n}, and then expand the 14×14×512 - dimensional feature vector of the last layer into 1 - dimensional, perform a fully - convolutional operation, and finally obtain 1×128 - dimensional feature vectors from n perspectives as the global vector G = {g 1 , g 2 , …, g n}.
[0085] Step 3 includes the following steps:
[0086] Step 3 - 1: Perform a max - pooling operation on the global feature vector G obtained in Step 2, that is, select the maximum values in each dimension and combine them into a 128 - dimensional feature vector that integrates the features of each perspective, and obtain the overall global feature G.
[0087] Step 4 includes the following steps:
[0088] Step 4 - 1: For the input three - dimensional query point information p, perform a projection operation using the camera parameters to obtain the two - dimensional query point Q in the corresponding perspective;
[0089] Step 4 - 2: Use the two - dimensional query point Q in Step 4 - 1 to query the hierarchical image features L in Step 2 respectively, obtain 1×64 - dimensional, 1×128 - dimensional, 1×256 - dimensional, 1×512 - dimensional, 1×512 - dimensional local point feature information respectively, and perform splicing to obtain 1×1472 - dimensional local feature information;
[0090] Step 5 includes the following steps:
[0091] Step 5 - 1: Input the two - dimensional query point Q into the PointNet network to extract the feature information of the point, a 1×512 - dimensional feature vector F.
[0092] Step 6 includes the following steps:
[0093] Step 6 - 1: Concatenate the global feature G obtained in Step 3 and the point feature F obtained in Step 5 to obtain a 1×640 - dimensional global feature with point feature information, and then perform a convolutional operation to obtain a 1×256 - dimensional global feature G;
[0094] Step 6 - 2: Concatenate the local features L from n perspectives obtained in Step 4 and the point feature F obtained in Step 5 to obtain local features with point features of 1×1984 - dimensional from n perspectives, and then perform a convolutional operation to obtain a 1×256 - dimensional local feature L;
[0095] Step 7 includes the following steps:
[0096] Step 7 - 1: Add the global feature G obtained in Step 6 to the local features L from n perspectives obtained in Step 6 respectively to obtain n 1×256 - dimensional feature information W = {w1 , w 2 , …, w n}
[0097] Step 7-2: Perform a voting operation on the feature W obtained in Step 7-1 to finally obtain a 1×256-dimensional feature information SDF. Among them, the voting method is to count the number of positive and negative values of the multi-view features obtained finally at the same point in each layer of SDF, and take the average value of the larger value as the SDF prediction value of the current point.
[0098] Step 7-3: Calculate the Loss value
[0099]
[0100]
[0101] In the present invention, according to experience, m = 4 and δ = 0.01 are set; m will penalize the relative position of the point in the model. When the point is inside the model, a larger penalty term will be obtained, so that the points inside the model play a greater role in the training of the network parameters, which also meets the requirements of the experiment. δ is a threshold, and the points within this threshold range will play a greater role in the change of the network parameters.
[0102] Step 7-4: Perform backpropagation to finally obtain a fully convolutional network that integrates multi-view features after training.
[0103] Step 8 includes the following steps:
[0104] Step 8-1: Use the Marching Cube method for the SDF information obtained in Step 7 to obtain the final three-dimensional model.
[0105] Example:
[0106] The target task of the present invention is as Figure 2 and Figure 3 shown. Figure 2 is the original model, Figure 3 is Figure 2 the result after three-dimensional reconstruction. The structural system of the whole method is as Figure 4 shown. The following will illustrate each step of the present invention according to the example.
[0107] Step (1): Collect data for the three-dimensional model dataset S. Taking the model s as an example, it is specifically divided into the following steps:
[0108] Step (1.1): Obtain 24-view images and the corresponding camera parameters.
[0109] Step (1.1.1): Randomly select 24 different views, such as Figure 5As shown, the selected viewing angle has no fixed angle. The pictures are obtained by placing the camera at a fixed position and rotating the model s. The size of the pictures is 137×137 pixels and the format is jpg.
[0110] Step (1.2) writes the pictures of 24 viewing angles and their corresponding camera parameters into an h5 file for unified management.
[0111] Step (1.3) uses the isosurface tool to process the model s in obj format into a file in SDF value representation format, which is in h5 format.
[0112] In step (2), the input picture data is convolved using the vgg network to obtain one-dimensional global features and multi-layer two-dimensional picture feature information under multiple viewing angles, as Figure 4 shown:
[0113] Step (2.1) randomly divides the input three-dimensional data set D = {D Train , D Test} into a training set D Train = {d 1 , d 2 , … d i , …, d n} and a test set D Test = {d n+1 , d n+2 , …, d n+j , …, d n+m}, where d i represents the i-th model in the training set, and d n+j represents the j-th model in the test set. n represents the number of the training set, and m represents the number of the test set;
[0114] Step (2.2), for the training set S Train , collect its projection maps P Train = {P 1 , P 2 , … P i , …, p n} and the camera parameter information C Train = {C 1 , C 2 , … C i , …, c n} and input them into the fully convolutional network for training to obtain a trained fully convolutional network that fuses multi-view features, and obtain local features L and global features G under multiple viewing angles, where P i refers to the i-th model s Train in the training set D iThe projection atlas under 24 perspectives, C i refers to the training set S Train the i-th model s in i the camera parameter set of 24 perspectives. This step can be specifically divided into the following steps:
[0115] Step (2.2.1), input the projection maps P under each perspective of the training set Train , where each model p i corresponds to the projection map information set P of 24 perspectives i and the camera parameter information set C i , select n (n ranges from 1 to 24) perspective pictures p extract ={p i , p m , …, p k} and input them into the vgg convolutional network;
[0116] Step (2.2.2), input the selected n-perspective pictures p extract into the vgg convolutional network. Among them, the vgg network adopts the vgg16 template provided in pytorch. After convolutional operations and pooling operations, extract the feature vectors of 224×224×64 dimensions, 112×112×128 dimensions, 56×56×256 dimensions, 28×28×512 dimensions, and 14×14×512 dimensions under n perspectives layer by layer as the hierarchical picture features L={l 1 , l 2 , …, l n}. Then expand the 14×14×512-dimensional feature vector of the last layer into 1 dimension, perform a fully convolutional operation, and finally obtain the 1×128-dimensional feature vectors under n perspectives as the global vectors G={g 1 , g 2 , …, g n}.
[0117] Step (3), perform a max pooling operation on the global feature vectors G obtained in Step 2, that is, select the maximum values in each dimension and combine them into a 128-dimensional feature vector that integrates the features of each perspective to obtain the overall global feature G.
[0118] Step (4), project the input query point information, use the two-dimensional query point to query the picture feature information to obtain the local point information, and perform splicing to obtain the local feature information. This step specifically includes the following steps:
[0119] Step (4.1), project the three-dimensional query point information p of the input model s using the camera parameter C to obtain the two-dimensional query point Q under the corresponding perspective.
[0120] Step (4.2): Use the two-dimensional query point Q in step (4.1) to query the hierarchical image features L in step (2) respectively, and obtain local point feature information of 1×64 dimension, 1×128 dimension, 1×256 dimension, 1×512 dimension, and 1×512 dimension respectively. Then splice them to obtain local feature information of 1×1472 dimension, as Figure 4 shown.
[0121] Step (5): Input the two-dimensional query point Q into the PointNet network to extract a feature vector F of point position feature information of 1×512.
[0122] Step (6): Combine the point position information features with the global features and local features respectively to obtain global features and local features with point position information features. Step (6) specifically includes the following steps:
[0123] Step (6.1): Splice the global feature G obtained in step (3) and the point position information feature F obtained in step (5) to obtain a global feature with point feature information of 1×640 dimension, and then perform a convolution operation to obtain a global feature G of 1×256 dimension.
[0124] Step (6.2): Splice the local features L from n perspectives obtained in step (4) and the point position information feature F obtained in step (5) to obtain local features with point features of 1×1984 dimension from n perspectives, and then perform a convolution operation to obtain local features L of 1×256 dimension.
[0125] Step (7): Combine the global features with the local features respectively, and generate the final feature information by voting. Step (7) specifically includes the following steps:
[0126] Step (7.1): Add the global feature G obtained in step (6) to the local features L from n perspectives obtained in step 6 respectively to obtain n pieces of feature information W = {w 1 , w 2 , …, w n} of 1×256 dimension.
[0127] Step (7.2): Perform a voting operation on the feature W obtained in step (7.1) to finally obtain a feature information SDF of 1×256 dimension. Among them, the voting method is to count the number of positive and negative values of each layer of SDF at the same point for the finally obtained multi-perspective features, and take the average value of the larger value as the SDF prediction value of the current point.
[0128] Step (7.3): Calculate the Loss value
[0129]
[0130]
[0131] In the present invention, according to experience, m = 4 and δ = 0.01 are set; m will penalize the relative position of points in the model. When a point is inside the model, a larger penalty term will be obtained, making the points inside the model play a greater role in the training of network parameters, which also meets the requirements of the experiment. δ is a threshold, and the points within this threshold range will play a greater role in the change of network parameters.
[0132] Step (7.4) performs backpropagation, and finally obtains a fully convolutional network that fuses multi-view features after training.
[0133] Step (8) includes the following steps:
[0134] Step (8-1), for the SDF information obtained in step (7), uses the Marching Cube method to obtain the final three-dimensional model.
[0135] Result analysis:
[0136] The experimental environment parameters of the method of the present invention are as follows:
[0137] 1) The experimental platform parameters for data collection of the model are Ubuntu16.04 64-bit operating system, Intel(R) Core(TM) i7-6850K CPU 3.60GHz, 32GB of memory. Python programming language is used and combined with the isosurface tool to implement, and the programming and development environment is Visual Code;
[0138] 2) The experimental platform parameters for the training and testing process of the FCN fully convolutional network that fuses multi-view features are Windows10 64-bit operating system, Intel(R) Core(TM) i7-6850K CPU 3.60GHz, 32GB of memory, and the graphics card is Titan RTX GPU 24GB. Python programming language is used and the third-party open-source library Pytorch is adopted to implement.
[0139] The comparative experimental results of the method of the present invention with the method in Document 12 (abbreviation DISN) and the method in Document 16 (abbreviation LsmNet) (as shown in Table 1 and Table 2) are analyzed as follows:
[0140] Experiments were conducted on the model sets of 16 categories in the well-known 3D model dataset ShapeNetCore. The category names of each category of the dataset are shown in the first column of Table 1. The meanings of the category names are Airplane (airplane), Bench (bench), Display (display screen), Car (car), Chair (chair), Speaker (intercom), Lamp (lamp), Sofa (sofa), Table (table), Phone (mobile phone), Watercraft (washbasin), Cabinet (cup); the division of the training set and the test set is shown in the second column of Table 1; the comparison of the semantic segmentation annotation effect rendering diagrams is as Figure 6 shown; the comparison based on the IoU parameter is shown in Table 1 and Table 2. Among them, when setting the IoU parameter for comparison, the dimension of the model participating in the comparison is designed to be 64×64×64. The ratio of the intersection of the points of the prediction model and the ground truth model to the union of the points of the prediction model and the ground truth model is used as the IoU, and the larger the value, the higher the accuracy.
[0141] As Figure 6 shown, the comparison between the method of the present invention and DISN. As shown in the comparison of the IoU values in Table 1 and Table 2 (Table 1 shows the IoU comparison between the method of the present invention and other methods on the ShapeNetCore dataset, and Table 2 shows the IoU statistical comparison between the method of the present invention and other methods on the ShapeNetCore dataset), the method of the present invention is ahead of the DISN method, and both exceed the DISN method in Category Avg (average of category accuracy) and Dataset Avg (average accuracy of the entire dataset), which confirms that this experiment has a good effect on multi-view image fusion.
[0142] Table 1 IoU comparison table between the method of the present invention and other methods on the ShapeNetCore dataset
[0143]
[0144] Table 2 IoU statistical comparison table between the method of the present invention and other methods on the ShapeNetCore dataset
[0145] DISN The method of the present invention Category Avg. 0.56 0.64
[0146] In the self-comparison experiment, the voting operation was selected to be removed, and the method of separately performing max pooling on local features and global features proposed in Document 12 (abbreviated as DISN) was adopted. The comparison of the IoU values with the final experimental results is shown in Table 3, indicating that this voting operation can effectively fuse the feature information of multi-view images.
[0147] In addition, compared with DISN, the present method proposes a feature fusion method from multiple perspectives, which solves the problem of insufficient inference of DISN for occluded parts and makes the model reconstruction more complete. More importantly, the method of the present invention is based on the DISN method. Without introducing other additional parameters, it realizes the feature fusion from multiple perspectives and achieves better results. Therefore, it has a certain universality. Table 3 shows the comparison between the final result of the method of the present invention and the DISN method using max pooling.
[0148] Table 3 Comparison between the final result of the method of the present invention and the DISN method using max pooling
[0149]
[0150] The present invention provides an idea and method for a three-dimensional reconstruction method based on fusing multi-perspective features and deep learning. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation mode of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A 3D reconstruction method based on fusing multi-view features and deep learning, characterized in that, it includes the following steps: Step 1, collect multi-view image data and corresponding signed distance equation data SDF as the input image data; Step 2, perform convolution operations on the input image data using the fully convolutional network vgg to obtain one-dimensional global feature information and multi-layer two-dimensional image feature information from multiple perspectives respectively; Step 3, perform max pooling operations on the global feature information from multiple perspectives in Step 2 to obtain the overall global feature information; Step 4, input the point information of the 3D model, that is, the 3D query points, project them to obtain the corresponding 2D query points, query the multi-layer feature information of the images with the 2D query points to obtain local point information, and splice the multi-layer local point information to obtain local feature information; Step 5, perform point convolution operations to obtain corresponding point features; Step 6, combine the point features obtained in Step 5 with the overall global feature information obtained in Step 3 and the local feature information obtained in Step 4 respectively to obtain global features and local features with point feature information; Step 7, combine the global features with point feature information in Step 6 with the local features with point feature information respectively, and generate the final feature information by voting; Step 8, generate the final 3D model for the final feature information obtained in Step 7 through the marching cubes algorithm; Among them, Step 6 includes the following steps: Step 6-1, splice the global feature vector G obtained in Step 3 and the point feature vector F obtained in Step 5 to obtain a 1×640-dimensional global feature with point feature information, and then perform convolution operations to obtain a 1×256-dimensional global feature vector G; Step 6-2, splice the local features L from n perspectives obtained in Step 4 and the point feature vector F obtained in Step 5 to obtain 1×1984-dimensional local features with point features from n perspectives, and then perform convolution operations to obtain a 1×256-dimensional local feature L; Step 7 includes the following steps: Step 7-1, add the global feature vector G obtained in Step 6 to the local features L under n perspectives obtained in Step 6 respectively to obtain n 1×256-dimensional feature information W = {w 1 , w 2 , …, w n} Step 7-2, perform voting operations on the feature W obtained in Step 7-1, and finally obtain a 1×256-dimensional feature information SDF. Among them, the voting method is to count the number of positive and negative values of each layer of SDF at the same point for the finally obtained multi-view features, and take the average value of the larger value as the SDF prediction value of the current point; Step 7-3, calculate the loss value where |·| represents the L1-norm, i.e., the L1 norm; f(I, p) represents the network output result, i.e., the network prediction value, where the inputs are the image I and the 3D point p; SDF I (p) represents the ground truth SDF value corresponding to the point p; δ is a threshold; m represents the relative position of the penalty point in the model. When the point is inside the model, a penalty term is obtained, and m1 > m2 is taken, where m1 is the value of the penalty term when the ground truth is less than the threshold δ, and m2 is the value of the penalty term when the ground truth is greater than the threshold δ; Step 7-4, perform backpropagation, and finally obtain a trained fully convolutional network that fuses multi-view features.
2. The 3D reconstruction method based on fusing multi-view features and deep learning according to claim 1, characterized in that, Step 1 includes the following steps: Step 1-1, obtain the original 3D dataset D from the Shapenet database, obtain 24-view two-dimensional image information P and corresponding camera parameter information C from the 3DR2N2 dataset; Step 1-2, generate corresponding SDF data from the 3D dataset D and store it in the corresponding h5 file; Step 1-3, store the two-dimensional image information P and the corresponding camera parameter information C into the h5 file.
3. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 2, characterized in that, Step 1-2 includes the following steps: Step 1-2-1, normalize the three-dimensional dataset D to generate a dataset N such that the numerical range of the dataset is between -1 and 1; Step 1-2-2, use the isosurface tool to process the dataset N into an SDF dataset; Step 1-2-3, use the Marching Cube method to process the SDF dataset to generate the corresponding three-dimensional dataset O; Step 1-2-4, store the SDF data information in the dataset O into the h5 file.
4. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 3, characterized in that, Step 2 includes the following steps: Step 2-1, randomly divide the input three-dimensional mesh model dataset D = {D Train , D Test} into a training set D Train = {d 1 , d 2 , … d i , …, d n} and a test set D Test = {d n+1 , d n+2 , …, d n+j , …, d n+m}, where d i represents the i-th model in the training set, d n+j represents the j-th model in the test set, n represents the number of the training set, and m represents the number of the test set; Step 2-2, for the training set D Train , collect the projection images P Train = {P 1 , P 2 , … P i , …, p n} and the camera parameter information C Train = {C 1 , C 2 , … C i , …, c n} and input them into the fully convolutional network for training to obtain a trained fully convolutional network that fuses multi-view features. At the same time, obtain the multi-view multi-layer two-dimensional picture feature information L and the one-dimensional global feature G. Among them, P i refers to the projection image set of the 24 views of the i-th model d Train in the training set D i , and C i refers to the camera parameter set of the 24 views of the i-th model d Train in the training set D i .
5. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 4, characterized in that, Step 2-2 includes the following steps: Step 2-2-1: Input the projection images P from each perspective of the training set Train , where for each model P i corresponds to the projection image information set and the camera parameter information set C for 24 perspectives i . Select n perspective images p extract = {p i , p m , …, p k} and input them into the vgg convolutional network, where the value range of n is from 1 to 24; Step 2-2-2: Input the pictures p of the selected n perspectives extract into the vgg convolutional network. After convolutional and pooling operations, extract the feature vectors of dimensions 224×224×64, 112×112×128, 56×56×256, 28×28×512, and 14×14×512 under the n perspectives layer by layer, as the layered picture features L = {l 1 , l 2 , …, l n}. Then expand the feature vector of dimension 14×14×512 in the last layer into 1 dimension and perform a fully convolutional operation. Finally, obtain the feature vectors of dimension 1×128 under the n perspectives as the global vectors G = {g 1 , g 2 , …, g n}.
6. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 5, characterized in that, Step 3 includes the following steps: Perform a max pooling operation on the global feature vector G obtained in Step 2, that is, select the maximum values in each dimension to form a 128-dimensional feature vector that fuses the features of each view to obtain the overall global feature G.
7. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 6, characterized in that, Step 4 includes the following steps: Step 4-1, project the input three-dimensional query point information D using the camera parameters to obtain the two-dimensional query point Q in the corresponding view; Step 4-2, use the two-dimensional query point Q obtained in Step 4-1 to query the hierarchical image features L in Step 2 respectively, and obtain local point feature information of 1×64 dimensions, 1×128 dimensions, 1×256 dimensions, 1×512 dimensions, and 1×512 dimensions respectively, and splice them to obtain local feature information of 1×1472 dimensions.
8. A three-dimensional reconstruction method based on fusing multi-view features and deep learning according to claim 7, characterized in that, Step 5 includes: input the two-dimensional query point Q into the PointNet network to extract the feature vector F of the point with a dimension of 1×512.
Citation Information
Patent Citations
Projection full-convolution network three-dimensional model segmentation method based on fusion of multi-view-angle features
CN108389251A
Method of synthesis of a two-dimensional image of a scene viewed from a required view point and electronic computing apparatus for implementation thereof
RU2749749C1