An Image-Based 3D Scene Reconstruction Method and Device
By applying the methods of neural implicit functions and Manhattan hypothesis in three-dimensional scene reconstruction, semantic segmentation and global geometric constraints are performed, the problem of poor perspective consistency in traditional methods is solved, and high-quality three-dimensional scene reconstruction is achieved.
Patent Information
- Application Number
- CN202210435947.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-24
AI Technical Summary
Traditional image-based three-dimensional scene reconstruction methods are difficult to maintain consistency between different perspectives, resulting in the accuracy and completeness of reconstruction results.
Using a method based on neural implicit functions, the Manhattan hypothesis is used to perform semantic segmentation of areas such as ground and wall, perform global geometric constraints, and extract three-dimensional grid models to achieve high-quality reconstruction by jointly optimizing semantics and geometry.
Through global geometric constraints, the reconstruction accuracy and completeness of walls, ground and other areas are significantly improved, and high-quality three-dimensional scene reconstruction is achieved.
Smart Images

Figure CN114742966B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of scene reconstruction, and in particular, to a method and device for three-dimensional scene reconstruction based on images. Background Art
[0002] For the problem of three-dimensional scene reconstruction based on images, traditional methods use multi-view stereo matching methods to estimate the depth map of each view according to the principle of photometric consistency, and then use depth map fusion technology to obtain the final reconstruction result. Most artificial indoor scenes conform to the Manhattan assumption, that is, the ground, walls, and ceiling should be aligned with three mutually perpendicular principal directions. Traditional methods use the Manhattan assumption as a constraint for each view depth map to improve the effect, but it is difficult to ensure the consistency between different views, so there is still a large room for improvement in the reconstruction result. The present invention applies the Manhattan assumption to the reconstruction method based on neural implicit functions, first obtains the semantics of areas such as the ground and walls, so as to globally constrain the geometry, and jointly optimizes the semantics and geometry to achieve high-quality reconstruction. Summary of the Invention
[0003] The object of the present invention is to propose a method and device for three-dimensional scene reconstruction based on images in view of the deficiencies of the prior art.
[0004] The object of the present invention is achieved by the following technical solutions: In the first aspect, the present invention provides a method for three-dimensional scene reconstruction based on images, and the method includes:
[0005] (1) Using a neural network implicit function to learn the signed distance field and color field to represent the geometry and appearance of the scene, and rendering the neural network implicit function into a two-dimensional image through volume rendering technology.
[0006] (2) Using semantic segmentation technology to obtain masks for the wall and ground areas, and adding geometric constraints to the corresponding areas based on the Manhattan assumption;
[0007] (3) Learning the semantic field in three-dimensional space, jointly optimizing the semantics and geometry, and extracting a three-dimensional mesh model from the optimized neural network implicit function to obtain the reconstruction result.
[0008] Further, in step (1), a set of three-dimensional points is sampled along the ray from the camera to the pixel, the signed distance and color of the three-dimensional points are calculated using the neural network implicit function, and the image pixel color value is obtained by numerical integration on the ray.
[0009] Further, the signed distance field and color field are implemented by a multi-layer perceptron.
[0010] Further, the neural network implicit function representation is optimized by minimizing the error of the pixel value and depth value between the rendered two-dimensional image and the input image and the norm constraint of the signed distance field.
[0011] Further, in step (2), for the three-dimensional surface points corresponding to the pixels determined to be in the ground area, the following loss function is adopted:
[0012] L f (r) = |1 - n(x r )·n f |
[0013] where x r is the three-dimensional surface point coordinate corresponding to the camera ray r, and n(x r ) is the normal vector obtained by taking the gradient of the signed distance field at x r , and n f = <0, 0, 1> is the unit vector with the vertical upward direction, which is used to represent the normal direction of the assumed ground area;
[0014] For the three-dimensional surface points corresponding to the pixels determined to be in the wall area, the following loss function is adopted:
[0015] L w (r) = min k∈{-1,0,1} |k - n(x r )·n w |
[0016] where n w is a learnable unit vector, which is initialized as <1, 0, 0> and is used to represent the direction of one of the walls, and n w can be jointly optimized with the network parameters during the training process.
[0017] Further, in step (3), a multi-layer perceptron network is used to learn the semantics in the three-dimensional space, and the volume rendering method is used to obtain the semantics of each pixel in the image space. The rendered semantics are normalized by softmax to obtain the probabilities belonging to the ground, wall, and other areas, which are and
[0018] Further, the following loss function is used for the joint optimization of semantics and geometry:
[0019]
[0020] At the same time, the following loss function is used to achieve the supervision of semantics:
[0021]
[0022] where, and respectively represent the sets of camera rays corresponding to the pixels in the ground and wall areas; is the probability obtained by rendering, p k (r) is the prediction result of the two-dimensional semantic segmentation network.
[0023] In a second aspect, the present invention provides an image-based three-dimensional reconstruction device, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, the image-based three-dimensional reconstruction method described above is implemented.
[0024] In a third aspect, the present invention provides a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, the image-based three-dimensional reconstruction method described above is implemented.
[0025] Advantages of the present invention: The present invention can globally constrain geometry based on the Manhattan assumption, thereby improving the accuracy and integrity of the reconstruction of areas such as walls and floors to achieve high-quality reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of the input and output of the present invention.
[0027] Figure 2 is a schematic diagram of the present invention reconstructing geometry from a monocular image sequence of a scene based on an implicit function.
[0028] Figure 3 is a structural diagram of an image-based three-dimensional reconstruction device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The technical details and principles of the present invention are further described below with reference to the accompanying drawings:
[0030] The present invention proposes an indoor scene reconstruction method based on the Manhattan assumption. As Figure 1 shown, the present invention reconstructs three-dimensional geometry from an input monocular image sequence of an indoor scene. The specific process is as follows:
[0031] (1) Use a neural network implicit function to learn the signed distance field and color field to represent the geometry and appearance of the scene. Through volume rendering technology, the neural network implicit function is rendered into a two-dimensional image. This includes: sampling a set of three-dimensional points along the ray projected from the camera to the pixel, using the neural network implicit function to calculate the signed distance and color of the three-dimensional points, and calculating the image pixel color value through numerical integration on the ray. By minimizing the error between the rendered two-dimensional image and the input image, the neural network implicit function representation is optimized.
[0032] (2) Use semantic segmentation technology to obtain the masks of the wall and floor areas, and add geometric constraints to the corresponding areas based on the Manhattan assumption. Then learn the semantic field in 3D space, jointly optimize semantics and geometry to improve the effect. Finally, use the Marching cubes algorithm to extract the 3D mesh model from the optimized neural network implicit function.
[0033] As Figure 2 shown, in the indoor scene reconstruction method proposed by the present invention, the specific steps for constructing the implicit function representing the scene geometry and appearance are as follows:
[0034] 1. The present invention represents the indoor scene based on the model in the standard coordinate system. The model in the standard coordinate system is specifically represented by continuous signed distance and color, where the signed distance field and color field are implemented by a multi-layer perceptron. The present invention represents the signed distance prediction of the three-dimensional point x in the standard coordinate system as the following function:
[0035] d(x), z(x) = F d (x)
[0036] where F d is a multi-layer perceptron network with 8 fully connected layers, d(x) is the signed distance of the three-dimensional point x, and z(x) is the feature vector output by the network, containing the shape information of the scene at the three-dimensional point x.
[0037] Regarding the color function, the present invention takes the three-dimensional point x, the viewing direction v, the normal direction n(x), and the shape feature z(x) as the inputs of the function. The color function is defined as follows:
[0038] c(x) = F c (x, v, n(x), z(x))
[0039] where F c is a multi-layer perceptron network with 4 fully connected layers, and the normal direction n(x) is obtained by taking the gradient of d(x).
[0040] In the indoor scene representation method proposed by the present invention, the neural network implicit function representation is optimized by differentiable rendering. The specific steps are as follows:
[0041] 1. Differentiable volume rendering: Given a viewing angle, use a differentiable volume renderer to convert the neural network implicit function representation into a 2D RGB image. For each pixel of the image, the differentiable volume renderer accumulates the volume density and color on the camera ray through an integral equation to obtain the pixel color. In actual implementation, the present invention uses numerical integration for approximation. The present invention first calculates the corresponding camera ray r using the camera parameters, and then samples K three-dimensional points between the nearest point and the farthest point For each 3D point x, the present invention converts it into a voxel density σ(x) according to its signed distance d(x), and the conversion function is as follows:
[0042]
[0043] where β is a learnable parameter. Based on this, the rendering color of the pixel can be obtained as follows:
[0044]
[0045] where δ i =||x i+1 -x i ||2 is the distance between adjacent sampling points, represents the accumulated transparency along the ray. Using differentiable volume rendering, the present invention optimizes the neural network implicit function representation by minimizing the error between the rendered image of each frame and the input image.
[0046] 2. Optimize the neural network implicit function representation, specifically: for the input monocular image sequence, the camera parameters are known. For each image, the present invention optimizes the parameters F d ,F c ,β, to minimize the following objective function:
[0047]
[0048] where is the set of camera rays passing through the pixels of the image, is the pixel value obtained by volume rendering, and C(r) is the true pixel value. The present invention uses the following constraint on the magnitude of the signed distance field:
[0049]
[0050] where is the Hamiltonian operator, is the set of 3D points, which consists of two parts: the first part is the set of points uniformly sampled in the artificially divided 3D space containing the entire scene, and the second part contains the intersection points of the camera rays corresponding to each pixel and the scene surface.
[0051] In addition, the present invention uses the depth map obtained by the multi-view stereo method as an aid, and the corresponding loss function is:
[0052]
[0053] where is the set of camera rays corresponding to the pixels with valid values in all depth maps, is the depth value obtained by volume rendering (calculated by integrating the transparency of the sampled points on the camera ray r), and D(r) is the depth value obtained by the multi-view stereo matching method based on the PatchMatch algorithm.
[0054] Based on the Manhattan assumption, the present invention proposes plane constraints for the ground and wall regions. The specific steps are as follows:
[0055] 1. Ground region: For the three-dimensional surface points corresponding to the pixels determined to be in the ground region, the present invention uses the following loss function:
[0056] L f (r) = |1 - n(x r ) · n f |
[0057] where x r is the three-dimensional surface point coordinate corresponding to the camera ray r, and n(x r ) is the normal vector obtained by taking the gradient of the signed distance field at x r , and n f = <0, 0, 1> is the unit vector in the vertically upward direction, used to represent the normal direction of the assumed ground region.
[0058] 2. Wall region: For the three-dimensional surface points corresponding to the pixels determined to be in the wall region, the present invention uses the following loss function:
[0059] L w (r) = min k∈{-1,0,1} |k - n(x r ) · n w |
[0060] where n w is a learnable unit vector, initialized as <1, 0, 0>, used to represent the direction of one of the walls, and n w can be jointly optimized with the network parameters during the training process.
[0061] 3. The present invention uses a two-dimensional semantic segmentation network to predict the masks of the ground and wall regions in the image space and defines the loss function as:
[0062]
[0063] where and respectively represent the sets of camera rays corresponding to the pixels in the ground and wall regions.
[0064] The semantic and geometric joint optimization proposed by the present invention has the following specific steps:
[0065] 1. The present invention uses a multi - layer perceptron to learn semantic logits in three - dimensional space. The implicit function of the semantic logits is defined as:
[0066] s(x)=F s (x)
[0067] where F s is a multi - layer perceptron network with 4 fully - connected layers.
[0068] The present invention uses the method of volume rendering to obtain the semantic logits of each pixel in the image space as follows:
[0069]
[0070] The present invention renders the obtained semantic logits and obtains the probabilities of multiple categories through softmax normalization and is used to represent the probabilities that each pixel belongs to the ground, wall, and other regions.
[0071] 2. The present invention proposes the following loss function for joint optimization of semantics and geometry:
[0072]
[0073] At the same time, the following loss function is used to achieve the supervision of semantics:
[0074]
[0075] where is the rendered probability, and p k (r) is the prediction result of the two - dimensional semantic segmentation network DeepLabV3 +.
[0076] Finally, the present invention uses the weighted sum of L img 、L eik 、L depth 、L joint and L s as the total loss function, and optimizes the neural implicit function based on the Adam optimizer. For the optimized result, the Marching cubes algorithm is used to extract a three - dimensional mesh model from the optimized neural network implicit function to obtain the reconstruction result.
[0077] Corresponding to the foregoing embodiments of the three - dimensional scene reconstruction method based on images, the present invention also provides embodiments of a three - dimensional scene reconstruction device based on images.
[0078] See Figure 3, An apparatus for three-dimensional scene reconstruction based on images provided by an embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the method for three-dimensional scene reconstruction based on images in the above embodiment.
[0079] Embodiments of the apparatus for three-dimensional scene reconstruction based on images of the present invention can be applied to any device with data processing capabilities. Such a device with data processing capabilities can be a device or apparatus such as a computer. The apparatus embodiments can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the apparatus for three-dimensional scene reconstruction based on images of the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, generally according to the actual functions of the device with data processing capabilities where the apparatus in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0080] For the implementation processes of the functions and roles of each unit in the above apparatus, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0081] For the apparatus embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The apparatus embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0082] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for three-dimensional scene reconstruction based on images in the above embodiment.
[0083] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or is to be output.
[0084] The above embodiments are used to explain the present invention rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention. Through the method of the present invention, high-quality reconstruction of indoor scenes can be achieved, especially in areas such as the ground and walls, where more accurate and complete reconstruction results can be obtained compared to previous methods.
Claims
1. A method for three-dimensional scene reconstruction based on images, characterized in that, The method includes: (1) Using a neural network implicit function to learn the signed distance field and color field to represent the geometry and appearance of the scene, and rendering the neural network implicit function into a two-dimensional image through volume rendering technology; (2) Using semantic segmentation technology to obtain masks for the wall and floor regions, and adding geometric constraints to the corresponding regions based on the Manhattan assumption; for the three-dimensional surface points corresponding to the pixels determined to be in the floor region, the following loss function is used: L f (r) = |1 - n(x r )·n f | where x r are the three-dimensional surface point coordinates corresponding to the camera ray r, and n(x r ) is the normal vector obtained by taking the gradient of the signed distance field at x r , and n f = <0, 0, 1> is the unit vector with the vertical upward direction, which is used to represent the normal direction of the assumed ground area; For the three-dimensional surface points corresponding to the pixels determined to be in the wall region, the following loss function is used: L w (r) = min k∈{-1,0,1} |k - n(x r )·n w | where n w is a learnable unit vector, initialized as <1, 0, 0>, for representing the direction of one of the walls, and n w can be jointly optimized with the network parameters during training; Predict the masks of the ground and wall regions in the image space using a two-dimensional semantic segmentation network and Define the loss function as: wherein and respectively represent the sets of camera rays corresponding to the pixels in the ground and wall regions; (3) Learn the semantic field in three-dimensional space, jointly optimize semantics and geometry, and use the weighted sum of L img , L eik , L depth , L joint and L s as the total loss function. L img is to optimize the neural network implicit function representation by minimizing the error between the rendered image of each frame and the input image. L eik is the magnitude constraint of the signed distance field. L depth is the depth map loss obtained by using the multi-view stereo method. L joint is the joint optimization loss of semantics and geometry. L s is the semantic supervision loss, and the neural implicit function is optimized based on the Adam optimizer. Extract the three-dimensional mesh model from the optimized neural network implicit function to obtain the reconstruction result; Joint optimization of semantics and geometry is performed through the following loss function: respectively represent the probabilities that the semantics belong to the ground and the wall; at the same time, the following loss function is used to implement the supervision of the semantics: Among them, and respectively represent the sets of camera rays corresponding to the pixels in the ground and wall regions; is the probability obtained by rendering, and p k (r) is the prediction result of the two-dimensional semantic segmentation network.
2. The 3D reconstruction method based on images according to claim 1, wherein In step (1), a set of three-dimensional points is sampled along the ray projected from the camera to the pixel, the signed distance and color of the three-dimensional points are calculated using the neural network implicit function, and the image pixel color value is obtained through numerical integration on the ray.
3. The 3D reconstruction method based on images according to claim 1, wherein The signed distance field and color field are implemented by a multi-layer perceptron.
4. The 3D reconstruction method based on images according to claim 2, wherein The neural network implicit function representation is optimized by minimizing the error of the pixel values and depth values between the rendered two-dimensional image and the input image, as well as the norm constraint of the signed distance field.
5. A three-dimensional reconstruction method based on images according to claim 1, characterized in that, In step (3), a multi-layer perceptron network is used to learn the semantics in three-dimensional space, and the volume rendering method is used to obtain the semantics of each pixel in the image space. The rendered semantics are normalized through softmax to obtain the probabilities of belonging to the floor, wall, and other regions.
6. An image-based three-dimensional reconstruction device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, the image-based 3D reconstruction method described in any one of claims 1-5 is implemented.
7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, the image-based 3D reconstruction method described in any one of claims 1-5 is implemented.