Sparse view outdoor scene three-dimensional reconstruction method based on multi-view stereo and grid
Through the three-dimensional reconstruction method of sparse view outdoor scenes with multi-view stereoscopic and mesh, the bounding box and global-local feature fusion are optimized by K-means clustering, the efficient reconstruction problem of bounded-free scenes under sparse view is solved, and the three-dimensional reconstruction effect with high precision and low memory is achieved.
Patent Information
- Application Number
- CN202510732701.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the prior art, under the sparse view input conditions, it is difficult to achieve efficient and high-quality three-dimensional reconstruction of boundary-free outdoor scenes, and there are problems such as blurred boundaries, artifacts, high memory consumption and loss of details in rendered images.
A sparse view outdoor scene three-dimensional reconstruction method based on multi-view stereoscopic and mesh is adopted. By acquiring image data and sparse three-dimensional point clouds, the cost body and depth probability body are constructed, combined with light sampling and spatial constraints, feature fusion and end-to-end optimization are performed, and the bounding box is optimized by K-means clustering, and global-local feature adaptive fusion is designed to reduce memory consumption and improve reconstruction accuracy.
It realizes comprehensive coverage and geometric texture integrity of boundless scenes under sparse views, significantly reduces memory consumption, improves reconstruction accuracy and visual realism, and solves the problems of details loss and structural distortion in traditional methods.
Smart Images

Figure CN120259559A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional reconstruction in computer graphics, and specifically relates to a method for three-dimensional reconstruction of outdoor scenes with sparse views based on multi-view stereo and meshes. Background Art
[0002] With the development of computer vision, three-dimensional reconstruction and novel view synthesis have been widely applied in fields such as virtual reality, robot navigation, autonomous driving, and medical image analysis. As a novel three-dimensional scene reconstruction technology, Neural Radiance Field (NeRF) has demonstrated excellent capabilities in novel view synthesis and photo-realistic rendering. However, the network structure of traditional NeRF adopts a per-scene independent optimization strategy, and prior knowledge cannot be shared between different scenes. Therefore, a large number of views need to be input for training and optimization during the reconstruction process. At the same time, the scene representation method of NeRF based on a multi-layer perceptron (MLP) has problems of slow training and rendering speeds. These problems limit the practical application of NeRF in large-scale scene reconstruction.
[0003] For three-dimensional reconstruction with sparse view inputs, existing multi-view stereo (MVS)-based methods are limited by the viewport coverage range and the number of input views, resulting in blurring and artifacts in the boundary regions of the rendered images. In recent years, research has improved the training and rendering efficiency by introducing explicit mesh-based representation methods, but still faces challenges such as high memory consumption, lack of fine structures due to insufficient mesh resolution, and loss of background details in large-scale unbounded scenes. Therefore, how to achieve fast, efficient, and high-quality reconstruction of unbounded outdoor scenes under sparse view inputs is the key to large-scale three-dimensional scene reconstruction based on neural radiance fields.
[0004] The existing publicly disclosed technology has a publication number of "CN119359934A" and a name of "A Generalizable Neural Radiance Field Reconstruction Method Based on Multi-Modal Information Fusion". This method generates a multi-modal neural encoding volume through progressive complementary fusion, converts the encoding volume and the original RGB pixel volume into volume density and radiance; subsequently, a transformer network is used to perform cross-view fusion on the context features of the sampled rays and decode to generate density and radiance values. While solving the shape-radiance ambiguity problem, this method improves the surface reconstruction accuracy of the generalizable neural radiance field. However, the large-scale parameter calculations of the multi-modal feature fusion network and the transformer network constructed by this method significantly increase the memory and computing power overhead, especially when processing high-resolution large-scale scenes. In addition, sparse geometric supervision may be unevenly distributed in textureless areas, resulting in local over-smoothing or missing details in the reconstruction results. Summary of the Invention
[0005] (I) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the existing technology, the present invention provides a three-dimensional reconstruction method for sparse-view outdoor scenes based on multi-view stereo and grids, to overcome the dependence of traditional neural radiance fields on dense-view inputs and improve the surface reconstruction accuracy of unbounded outdoor scenes under sparse-view input conditions.
[0007] (2) Technical solutions
[0008] The present invention specifically adopts the following technical solutions to achieve the above objectives:
[0009] A three-dimensional reconstruction method for sparse-view outdoor scenes based on multi-view stereo and grids, comprising the following steps:
[0010] Step 1, obtain an image dataset of the scene to be reconstructed, including image data, camera pose parameters, and sparse three-dimensional point cloud information of the scene, and enclose the scene using a bounding box;
[0011] Step 2, extract multi-scale image features of the input views, construct a cost volume based on the multi-view stereo (MVS) framework, perform regularization processing on the cost volume to generate a neural encoding volume and a depth probability volume, and obtain a prior of the depth distribution of the scene through a probability depth estimation method;
[0012] Step 3, ray sampling and spatial constraint: perform adaptive sampling on the rays, and remap the rays of the unbounded scene using an equal-proportion spatial contraction function to limit the space within a finite grid range;
[0013] Step 4, extract local features of the target image and global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through an attention network;
[0014] Step 5, decode the fused features and the neural encoding volume features, output the volume density and color values of the sampling points, calculate the pixel RGB values based on the volume rendering equation, and finally perform end-to-end optimization by combining photometric consistency supervision and sparse geometric supervision of the grid.
[0015] Further, in the above Step 1, the bounding box is used to define the geometric boundary of the scene. In this process, it is necessary to determine the size range of the bounding box. First, divide the sparse point cloud in the dataset into several clusters based on the K-means clustering method, select the main cluster with the highest point density, calculate its centroid as the center point of the scene, extract the minimum and maximum values of the main cluster on the XYZ axes to generate the vertex coordinates of the compact bounding box, and then obtain the bounding box range.
[0016] Further, in the above Step 2, the process of constructing the cost volume and regularization is as follows:
[0017] Select M pictures as input images, and use 2D CNN for the input views Extract 2D neural features to obtain a feature map , where represents the real number space, and represent the dimensions of the matrix; obtain camera parameters , where is the camera intrinsic matrix, and represent the rotation and translation matrices of the camera respectively. Then, use the homography transformation matrix to transform all feature maps into the conical volume space of the reference view. The homography transformation matrix is as follows:
[0018]
[0019] where represents the normal direction of the reference view, represents the camera parameters of the reference view, represents the camera parameters of the input view. Obtain the depth at the input view to the warped feature map of the reference view:
[0020]
[0021] where represents the pixel coordinates of the reference view. Generate the cost volume at the reference view by calculating the variance of the deformed feature map. The formula is as follows:
[0022]
[0023] where represents the variance calculation. Use 3D CNN to regularize the reference view cost volume to generate the depth probability volume and the neural encoding volume. Among them, 3D CNN is composed of 3D U-Net to process the spatial information of the data; as shown in the formula, use the obtained probability volume to calculate the expectation along the depth direction to obtain the initial depth value of the corresponding pixel:
[0024]
[0025] where represents the number of depth channels, represents the probability of this pixel on the depth plane . Calculate the expectation for each pixel point to obtain the depth distribution of the image.
[0026] Furthermore, in step 3, the adaptive sampling guides the light rays through the image depth distribution information for more accurate scene surface sampling, reduces the sampling points in the invalid regions in the three-dimensional space, and performs non-uniform sampling through the designed metric distance mapping function. Therefore, the light ray sampling points are expressed as follows:
[0027]
[0028] where represents the light ray origin, represents the light ray direction, represents the metric distance. The non-uniform sampling strategy is realized by parameterizing the metric distance, so that the distribution of sampling points matches the scene characteristics. The metric distance and the normalized distance are mapped as follows:
[0029]
[0030] where are the near-plane distance and the far-plane distance defined according to the depth value respectively;
[0031] In addition, the space contraction function is designed based on the ray-voxel intersection detection theorem, and the sampling points on any light ray path are isometrically compressed into a finite grid space, so as to realize the global linear scaling of the unbounded scene space. The specific definition is as follows:
[0032]
[0033]
[0034] where represent the incident point coordinates and the exit point coordinates of the light ray entering the bounding box respectively, represents the light ray index, represents the coordinates of the first sampling point of the light ray, represents the scaling ratio of the light ray. Through the calculation, the light ray is compressed into a compact finite space.
[0035] Furthermore, the global-local feature dynamic fusion process in step 4 is as follows:
[0036] First, extract the global features. Based on tensor decomposition, construct a global spatial radiation field with a continuous grid, obtain the global consistency confidence feature vector and the appearance feature vector by interpolation for the discrete voxels adjacent to the query point, multiply the feature vectors by the color matrix and combine with the light ray direction Through the decoding function Obtain the global features And the global consistency confidence , and the calculation formula is as follows:
[0037]
[0038] Where the decoding function Adopts a multi-layer perceptron network, Represents the appearance feature vector obtained by interpolation, Represents the confidence feature vector obtained by interpolation, Represents the dimension of the spatial tensor decomposition, Represents the tensor decomposition channel index, Represents the concatenation operation, and the global consistency confidence Is used to evaluate the geometric consistency, texture richness and occlusion in the multi-view of the grid position, and dynamically adjust the weights of the global features and local features;
[0039] Then extract the local features. For any queried 3D point, project it onto the multi-resolution target image, extract the two-dimensional plane features at each scale through bilinear interpolation, and concatenate the features to generate a feature vector with a length of As the geometric local feature , where Represents the single-scale feature dimension;
[0040] Then adopt a lightweight attention network , using the global feature And the local feature And the confidence As the input to generate the fusion weight . Based on this weight, realize the adaptive fusion of the global-local features. The weight And the fusion feature calculation formulas are as follows:
[0041]
[0042]
[0043] Where the Sigmoid function is used to limit the weight to [0,1] to achieve a smooth transition, Represents the finally obtained fusion feature.
[0044] Furthermore, in step 5, the decoding process is based on a multi-layer perceptron MLP to predict the volume density And the color value ; where the th sampling point is represented as:
[0045]
[0046] Among them represents the neural coding body feature, which is used to infer more appearance perception information from the cost volume;
[0047] The volume rendering equation is expressed as the cumulative summation of the voxel density and color values along each ray, and the RGB value of the pixel corresponding to the ray direction is obtained. The volume rendering equation is as follows: where
[0048]
[0049] where represents the cumulative transmittance, represents the distance between adjacent sampling points, represents the number of ray sampling points;
[0050] The photometric consistency supervision and the sparse geometric supervision process of the grid are used to calculate the color loss and the voxel grid regularization, and finally a weighted sum is performed to obtain the final loss function:
[0051]
[0052] where is the weight coefficient.
[0053] (III) Beneficial effects
[0054] Compared with the prior art, the present invention provides a three-dimensional reconstruction method for outdoor scenes based on multi-view stereo and sparse views of grids, which has the following beneficial effects:
[0055] The present invention proposes an explicit-implicit hybrid radiation field representation method for sparse inputs. This radiation field can provide a global viewport coverage range, effectively capture the areas missed in traditional multi-view stereo vision methods, and thus ensure the comprehensive coverage of the scene and the integrity of geometric texture information during the three-dimensional reconstruction process.
[0056] The present invention designs a global-local feature adaptive fusion method to achieve the collaborative optimization of local texture details and global geometric constraints, effectively avoiding the problems of detail loss or structural distortion caused by a single cost volume in traditional methods. In addition, based on the multi-level representation characteristics of the voxel grid, the estimation error of the geometric shape can be significantly reduced, and the accuracy and visual realism of the three-dimensional reconstruction can be improved.
[0057] The present invention proposes a method for representing a borderless scene. By using a spatial equal-proportion contraction function, sampling points beyond the grid boundary are mapped to the finite grid range without changing the light direction, enabling the voxel grid to effectively cover the entire unbounded scene area and improving the ability to express distant view details.
[0058] The present invention uses the K-means clustering algorithm to optimize the spatial bounding box, compress redundant grid space, and adopts an optimized method of depth-guided ray sampling to reduce memory consumption while maintaining geometric accuracy. Brief Description of the Drawings
[0059] Figure 1 is the overall flowchart provided by the present invention;
[0060] Figure 2 is the overall framework diagram of the 3D reconstruction method provided by the present invention;
[0061] Figure 3 is the comparison diagram of sampling point distributions provided by the present invention;
[0062] Figure 4 is the schematic diagram of the spatial contraction function provided by the present invention;
[0063] Figure 5 is the schematic diagram of multi-resolution geometric feature extraction provided by the present invention;
[0064] Figure 6 is the comparison diagram of scene reconstruction results on the Free dataset provided by the present invention;
[0065] Figure 7 is the comparison diagram of scene reconstruction results on the Free dataset provided by the present invention. Detailed Description of the Embodiments
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] Embodiment
[0068] As Figure 1 shown, the flowchart of the 3D reconstruction method for outdoor scenes with sparse views based on multi-view stereo and grids proposed in an embodiment of the present invention specifically includes the following steps:
[0069] Step 1: Obtain the image dataset of the scene to be reconstructed, including image data, camera pose parameters, and sparse three-dimensional point cloud information of the scene, and determine the range of the scene bounding box.
[0070] Prepare the Free dataset, which is provided by F2-NeRF and consists of seven outdoor scenes. Compared with widely used outdoor scenes such as the LLFF and NeRF-360-V2 datasets, this dataset has narrow and long camera trajectories and focused foreground objects. The Free dataset contains image data, corresponding image camera poses obtained by COLMAP, and sparse three-dimensional point cloud information of the scene (including the coordinates, colors, and associated image indices of the points). During the training process, the evaluation of the Free dataset follows the ratio of dividing the training set and the test set in F2-NeRF, that is, one-eighth of the images are used for testing, and the rest are used for training.
[0071] When using a bounding box for a scene, the center of the bounding box is usually set as the coordinate origin. However, when the scene is a forward scene, the geometric distribution of the scene is highly concentrated in half of the bounding box space, resulting in a large number of invalid grids occupying memory resources and reducing the ray sampling efficiency. Therefore, the present invention proposes a method to re-determine the center and vertices of the bounding box based on the K-means clustering algorithm. K-means is an unsupervised clustering algorithm, and its core goal is to divide the data points into m clusters, maximizing the similarity of data points within the same cluster and the difference between different clusters. According to the characteristics of the forward scene, set the number of clusters m. For each point cloud , calculate its Euclidean distance from the centroid of each cluster
[0072]
[0073] Assign the point cloud to the closest cluster , update the centroid according to the mean value of the coordinates of the points within the cluster, and thus divide the point cloud into several clusters; select the main cluster with the highest point density, calculate its centroid as the center point of the scene, and extract the minimum and maximum values of the main cluster on the XYZ axes to generate the vertex coordinates of the bounding box. Optimizing the bounding box through K-means clustering can significantly compress the redundant grid space, improve the ray sampling efficiency and reconstruction quality, and provide an efficient data basis for subsequent feature fusion and volume rendering.
[0074] Step 2: Extract multi-scale image features of the input view, construct a cost volume based on the multi-view stereo (MVS) framework, regularize the cost volume to generate a neural encoding volume and a depth probability volume, and obtain the depth distribution prior of the scene through a probability depth estimation method. The specific process is as follows:
[0075] Ⅰ. Construction of cost volume: Select M pictures as input images, and use 2D CNN to extract 2D neural features from the input images, where represents the real number space, and represents the dimension of the matrix. As shown in Figure 2 , the 2D CNN consists of three downsampling convolutional layers, and finally the input image feature is obtained, and the matrix dimension size is . Obtain the camera parameters , where is the camera internal parameter matrix, and and represent the rotation and translation matrices of the camera respectively. Use the homography transformation matrix to transform each feature map to the frustum volume of the reference view. The homography transformation matrix is as follows:
[0076]
[0077] where represents the normal direction of the reference view, represents the camera parameters of the reference view, and represents the camera parameters of the input view. Through the homography transformation, the warped feature map of the input view at depth to the reference view is obtained: :
[0078]
[0079] where represents the pixel coordinates of the reference view. The warped feature map is used to generate the reference view cost volume through variance calculation, and the formula is as follows:
[0080]
[0081] Ⅱ. Depth distribution prediction: Use 3D CNN to regularize the reference view cost volume and process it into a depth probability volume and a neural encoding volume. The 3D CNN consists of a 3D U-Net to process the spatial information of the data. For a pixel point in the target view, represents the probability value of this pixel on the depth plane . As shown in the formula, the depth value of the pixel is defined as the expected value of the depth probability distribution along the depth direction:
[0082]
[0083] where represents the number of depth channels. By taking the expectation of each pixel, the depth distribution of the image is obtained.
[0084] Step 3, ray sampling and spatial constraint:
[0085] Ⅰ. Guiding ray sampling based on depth distribution: After obtaining the image depth distribution, the depth information corresponding to the pixel is transferred to the voxel grid space for ray sampling. Different from traditional NeRF, in the method of the present invention, each ray defines the nearest plane distance and the farthest plane distance based on the actual depth value, rather than using globally fixed nearest plane distance and farthest plane distance. Therefore, the sampling strategy based on pixel-level depth information can achieve more accurate sampling of the scene surface, reduce the sampling points in the blank areas in the three-dimensional space, and the ray sampling points are expressed as follows:
[0086]
[0087] where represents the ray origin, represents the ray direction, represents the metric distance.
[0088] Ⅱ. Parametrization of the metric distance: In unbounded scene reconstruction, the sampling strategy is crucial for efficiency and quality. Research shows that the non-uniform sampling strategy conforms to the characteristics of the scene distribution: the near-view region contains rich geometric and texture details, and high-density sampling is required to ensure the reconstruction accuracy; the far-view region has fewer details, and over-sampling will waste computing resources. However, the parametric metric distance equation of Mip-NeRF 360 has a sudden change in the sampling interval at the critical point, resulting in an uneven sampling distribution and prone to banding artifacts or texture breaks in the reconstruction results. Therefore, the present invention proposes a new mapping function between the metric distance and the normalized distance to optimize the sampling distribution:
[0089]
[0090] where are the near plane distance and the far plane distance defined according to the depth value respectively. As Figure 3 shown, in the left figure, the sampling points of Mip-NeRF 360 are too dense in the near-view region and too sparse in the far-view region, and the sampling interval changes suddenly at the critical point. While in the right figure, the sampling point distribution optimized by the present invention satisfies the characteristics of dense near-view and sparse far-view, eliminates the sudden change in the sampling interval, and significantly improves the smoothness and numerical stability.
[0091] Ⅲ. Representation of unbounded scenes: As Figure 4As shown in (a), in an unbounded scene, the dynamic range of the original ray length is extremely large and lacks geometric constraints. Traditional NeRF uses NDC transformation to map the forward unbounded scene to a cube space. For 360° scenes, Mip-NeRF 360 maps the distant space to a bounded sphere through a contraction function, but there is a problem that background straight lines are distorted into curves, resulting in difficulties in ray-voxel grid intersection detection. MERF proposes a piecewise function mapping scheme, but light discontinuity occurs at the boundary, resulting in uneven sampling distribution and reconstruction artifacts. Therefore, in the present invention, while maintaining the original propagation direction vector of the light, the scene space coordinates are globally linearly scaled. Based on the ray-voxel intersection detection theorem, the scale factor of the diagonal length of the scene bounding box and the side length of the grid space is calculated, and the sampling points are isometrically compressed into a finite grid space as defined below:
[0092]
[0093]
[0094] where represent the incident point coordinates and the exit point coordinates of the light entering the bounding box respectively, represents the ray index, represents the coordinates of the first sampling point of the ray. represents the scaling ratio of the ray. As Figure 4 shown in (b), the proposed space contraction function maps the unbounded scene to a finite voxel grid space by performing calculations on the sampling points without changing the light direction. This feature enables the use of standard ray-AABB intersection tests. In addition, through space contraction, the distribution of sampling points is optimized, which can not only improve the reconstruction quality of distant scene details, but also effectively solve the problem of difficult training convergence caused by the too large difference in the value range of sampling point coordinates in the unbounded scene.
[0095] Step 4, extract the local features of the target image and the global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through an attention network;
[0096] Traditional MVSNeRF constructs a single cost using limited views and often faces problems such as insufficient reconstruction accuracy of the image edge shape and incorrect geometric shapes in non-occluded areas. Therefore, the present invention proposes an explicit-implicit radiance field based on dynamic fusion of global-local features.
[0097] Ⅰ. Global Feature Query: To reduce the memory occupancy of the voxel grid, the present invention constructs a global spatial radiation field with a continuous grid based on tensor decomposition. Specifically, the radiation field is expressed as a set of compact vectors and matrices. To represent discrete voxels in the continuous radiation field representation, the global consistency confidence feature vector and appearance feature vector of the sampling points are obtained by linear interpolation and bilinear interpolation for the discrete voxels adjacent to the query point. The formula is as follows. Multiply the feature vector by the color matrix and combine with the ray direction Through the decoding function to obtain the global feature and the global consistency confidence .
[0098]
[0099] Among them, the decoding function adopts a multi-layer perceptron network, represents the appearance feature vector obtained by interpolation, represents the confidence feature vector obtained by interpolation, represents the dimension of the spatial tensor decomposition, represents the tensor decomposition channel index, represents the concatenation operation. The global consistency confidence is used to evaluate the geometric consistency, texture richness, and occlusion situation of the grid position in multiple views, and dynamically adjust the weights of the global feature and the local feature. The high-confidence region indicates that the geometric information of this grid position is highly consistent and reliable in multiple views (such as unoccluded and texture-rich regions). At this time, the global feature should be given priority and the possible noise in the local view should be suppressed. While the low-confidence region indicates that there may be occlusion, low texture, or inconsistency between views at this position. At this time, the local feature of the target view should be relied on to avoid the noise interference in the global feature.
[0100] Ⅱ. Local Feature Extraction: The 2D plane can provide strong semantic information for 3D perception image synthesis and 3D reconstruction. To achieve the generalization ability between scenes and solve the rasterization problem of the low-resolution voxel grid, the present invention proposes to extract multi-resolution features from the target image as local features . As Figure 5 shown, for any queried three-dimensional point, project it to different resolution levels of the target image, extract the corresponding 2D plane features by bilinear interpolation, and concatenate the interpolation features at three scales to construct a feature vector with a length of as the geometric local feature , where represents the single-scale feature dimension.
[0101] Ⅲ. Global-Local Feature Dynamic Fusion: For any spatial sampling point Interpolate from neighboring voxels to obtain the global features of the query points within the voxel , as Figure 2 shown, a lightweight attention network is adopted, using the global features local features as the input, and at the same time using as the gating factor to dynamically adjust the fusion weights. Generate the fusion weights . Based on this weight, dynamic fusion of global-local features is achieved. The specific process is as follows:
[0102] First, through the attention network , generate the fusion weights :
[0103]
[0104] where the Sigmoid function is used to map the weights to [0,1] to achieve smooth transition. Finally, the global and local features are weighted and summed to obtain the final fused features :
[0105]
[0106] Dynamically fuse the global features of the voxel grid and the local features of the target view through the attention module, and preferentially select the information in the high-confidence region; at the same time, update the feature weights in real time based on the geometric consistency of the voxel grid, and use the global features to fill the missing geometric information in the low-confidence region to avoid the hole or artifact problems caused by insufficient local features in traditional methods.
[0107] Step 5: Decode the fused features and the neural encoded volume features, output the volume density and color values of the sampling points, calculate the pixel RGB values based on the volume rendering equation, and finally perform end-to-end optimization by combining photometric consistency supervision and sparse geometric supervision of the grid;
[0108] To infer more appearance-aware information from the cost volume, trilinear interpolation is used for the queried sampling points to extract the corresponding features from the neural encoded volume, denoted as , and combine with the fused features and the ray direction to decode through the MLP network to obtain the volume density and color value of the sampling points, which are expressed as follows:
[0109]
[0110] After obtaining the color value and volume density prediction values of the sampling points, perform cumulative summation on the sampling points along each ray, and calculate the ray direction through the volume rendering equation The corresponding pixel RGB values :
[0111]
[0112] where represents the cumulative transmittance, represents the distance between adjacent sampling points, represents the number of sampling points for a single ray.
[0113] Due to the differentiability of the overall network structure, the multi-view stereo stage can be jointly optimized and trained with the neural radiance field constructed by the voxel grid. The training uses the mean squared error loss function to calculate the photometric consistency loss, which is defined as:
[0114]
[0115] where represents the total number of sampling rays. In addition, to encourage the sparsity of the tensor factor parameters and make the norms of these two density tensors as small as possible, the standard regularization is used to improve the quality of the extrapolated views and remove floating points in the final rendering. The expression is:
[0116]
[0117] where represents the matrix, represents the vector, represents the total number of voxel grid parameters. The photometric consistency loss and are weighted and summed to obtain the final loss function:
[0118]
[0119] where, represents the weight coefficient, which is set to .
[0120] The experimental comparison is carried out on the Free dataset, which contains 7 outdoor scenes with a resolution of 1461×935. Both model training and testing are carried out on a single RTX 4090 GPU. For the optimization of each scene, 10,000 iterations of training are performed, using the Adam optimizer with an initial learning rate of 5e -4 and an exponential scheduler. During the training process, Batch_size uses 4096 rays, the number of sampling points is 8, the number of input views is 3, and the settings of the remaining network parameters all follow the parameter settings in MVSNeRF. The grid adopts a coarse-to-fine reconstruction strategy, starting from a grid with = 128 3Starting from the initial low-resolution grid, the grid resolution is linearly upsampled in a fixed number of rounds to finally obtain =300 3 the grid resolution. PSNR, SSIM, and LPIPS are used as evaluation metrics for the rendering results.
[0121] On the Free dataset, the method proposed in this invention is compared with baseline methods, including the grid-based method TensoRF, the sparse view reconstruction methods MVSNeRF, ENeRF, FreeNeRF, and BoostMVSNeRF. The results are shown in Table 1, and the method of this invention performs best in the rendering quality of this dataset. Figure 6 and Figure 7 shows the qualitative comparison results. Due to the poor quality of low-resolution grid reconstruction, TensoRF has problems. For MVSNeRF and its variants, due to the locality of multi-view geometric constraints and the sparsity of view coverage, there are problems such as incorrect geometry and severe blurring at the boundaries in the rendered images. In contrast, the method proposed in this invention uses voxel grids to construct multi-view global features, fills in the missing geometric information in low-confidence regions through a global-local feature dynamic fusion mechanism, effectively solves the problems of edge blurring and geometric distortion under sparse views, and obtains better rendering quality. In addition, combined with the proposed spatial constraint processing, the method of this invention is superior in the ability to express distant view details.
[0122] Table 1 Quantitative comparison results of the Free dataset with the latest methods
[0123] Comparison method PSNR↑ SSIM↑ LPIPS↓ TensoRF 21.11 0.481 0.316 MVSNeRF 17.26 0.433 0.529 ENeRF 24.42 0.797 0.218 FreeNeRF 26.09 0.774 0.278 BoostMVSNeRF 25.64 0.835 0.193 The present invention 26.17 0.836 0.182
[0124] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A three-dimensional reconstruction method for sparse-view outdoor scenes based on multi-view stereo and meshes, characterized in that It includes the following steps: Step 1: Obtain the image dataset of the scene to be reconstructed, including image data, camera pose parameters, and sparse three-dimensional point cloud information of the scene, and enclose the scene with a bounding box; Step 2: Extract multi-scale image features of the input view, construct a cost volume based on the multi-view stereo framework, perform regularization processing on the cost volume to generate a neural encoding volume and a depth probability volume, and obtain the prior of the depth distribution of the scene through a probabilistic depth estimation method; Step 3: Ray sampling and spatial constraint: Perform adaptive sampling on the rays, and remap the unbounded scene rays using an equal-proportion spatial contraction function to limit the space within a finite grid range; Step 4: Extract the local features of the target image and the global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through an attention network; Step 5: Decode the fused features and the neural encoding volume features, output the volume density and color values of the sampling points, calculate the pixel RGB values based on the volume rendering equation, and finally perform end-to-end optimization by combining photometric consistency supervision and sparse geometric supervision of the grid.
2. The three-dimensional reconstruction method of sparse view outdoor scenes based on multi-view stereo and grid according to claim 1, characterized in that: In the said Step 1, the bounding box is used to define the geometric boundary of the scene. In this process, it is necessary to determine the size range of the bounding box. First, divide the sparse point cloud in the dataset into several clusters based on the K-means clustering method, select the main cluster with the highest point density, calculate its centroid as the center point of the scene, extract the minimum and maximum values of the main cluster on the XYZ axes to generate the vertex coordinates of the compact bounding box, and then obtain the bounding box range.
3. The three-dimensional reconstruction method of sparse-view outdoor scenes based on multi-view stereo and grid according to claim 1, characterized in that: In the said Step 2, the process of constructing the cost volume and regularization is as follows: Select M pictures as input images and use 2D CNN for the input views Extract 2D neural features to obtain a feature map , where represents the real number space, and represent the dimensions of the matrix; Obtain camera parameters , where is the camera intrinsic matrix, and represent the rotation and translation matrices of the camera respectively. Then use the homography transformation matrix to transform all feature maps into the conical solid space of the reference view. The homography transformation matrix is as follows: ; wherein represents the normal direction of the reference view, represents the camera parameters of the reference view, represents the camera parameters of the input view, and the depth is obtained through a homography transformation at the input view to the warped feature map of the reference view : ; Among them represents the pixel coordinates of the reference view, and the deformed feature map is used to generate the cost volume at the reference view through variance calculation , and the formula is as follows: ; Among them represents variance calculation, and uses 3D CNN to regularize the reference view cost volume to generate a depth probability volume and a neural encoding volume, where the 3D CNN is composed of a 3D U-Net to process the spatial information of the data; As shown in the formula, using the obtained probability volume , taking the expectation along the depth direction, the initial depth value of the corresponding pixel is obtained : ; wherein represents the number of depth channels, represents the probability of the pixel on the depth plane and the expectation is calculated for each pixel to obtain the depth distribution of the image.
4. The three-dimensional reconstruction method of sparse-view outdoor scenes based on multi-view stereo and grid according to claim 3, wherein: In step 3, adaptive sampling guides the light to perform more accurate scene surface sampling through the image depth distribution information, reduces the sampling points in the invalid areas in the three-dimensional space, and performs non-uniform sampling through the designed metric distance mapping function. Therefore, the light sampling points are expressed as follows: ; Among them represents the light origin represents the light direction represents the metric distance. By parameterizing the metric distance, a non-uniform sampling strategy is implemented to make the distribution of sampling points match the scene characteristics. The metric distance and the normalized distance The mapping function of is expressed as follows: ; wherein are a near-plane distance and a far-plane distance defined according to depth values, respectively; In addition, the spatial contraction function is designed based on the ray-voxel intersection detection theorem, and samples on any ray path are equidistantly contracted into a finite grid space, so as to achieve global linear scaling of the unbounded scene space. The specific definition is as follows: ; ; where respectively represent the incident point coordinates and the exit point coordinates where the light enters the bounding box, represents the ray index, represents the coordinates of the first sampling point of the ray, represents the scaling ratio of the ray. By performing calculations, the ray is compressed into a compact finite space.
5. The three-dimensional reconstruction method of sparse view outdoor scenes based on multi-view stereo and grid according to claim 4, characterized in that: In the said Step 4, the process of dynamic fusion of global-local features is as follows: First, extract the global features, construct a global spatial radiation field with a continuous grid based on tensor decomposition, obtain the global consistency confidence feature vector and appearance feature vector of the discrete voxels adjacent to the query point through interpolation, multiply the feature vectors with the color matrix and combine with the ray direction through the decoding function to obtain the global feature and the global consistency confidence . The calculation formula is as follows: ; Among them, the decoding function adopts a multi-layer perceptron network, represents the appearance feature vector obtained by interpolation, represents the confidence feature vector obtained by interpolation, represents the dimension of spatial tensor decomposition, represents the tensor decomposition channel index, represents the concatenation operation, global consistency confidence is used to evaluate the geometric consistency, texture richness, and occlusion situation of the grid position in multiple views, and dynamically adjusts the weights of global features and local features; Next, extract local features. For any three-dimensional point to be queried, project it onto the target image with the maximum resolution, extract two-dimensional plane features at each scale through bilinear interpolation, cascade the features, and generate a feature vector with a length of as the geometric local feature , where represents the single-scale feature dimension; Then a lightweight attention network is adopted , using the global feature and the local feature along with the confidence as inputs to generate the fusion weights . Based on these weights, an adaptive fusion of the global-local features is achieved. The weights and the calculation formula for the fused features are as follows: ; ; The Sigmoid function is used to limit the weights to [0, 1] to achieve smooth transition. It represents the finally obtained fused features.
6. The method for three-dimensional reconstruction of sparse-view outdoor scenes based on multi-view stereo and grid according to claim 5, characterized in that: In the said step 5, the decoding process is based on a multi-layer perceptron (MLP) to predict the volume density of the sampling points and the color values ; where the th sampling point is represented as: ; Among them represents the neural coding body feature, which is used to infer more appearance perception information from the cost body; The volume rendering equation is expressed as the cumulative summation of the volume density at the sampling points along each ray and the color values to obtain the RGB value of the pixel corresponding to the ray direction as follows: The volume rendering equation is as follows: ; Among them represents the cumulative transmittance, represents the distance between adjacent sampling points, represents the number of light sampling points; The photometric consistency supervision and the sparse geometric supervision process of the grid are used to calculate the color loss and the voxel grid Regularization is performed, and finally a weighted sum is taken to obtain the final loss function: ; wherein is a weight coefficient.
Citation Information
Patent Citations
Nerve radiation field three-dimensional reconstruction method based on space mixed attention
CN118918250A
Generalized neural radiation field reconstruction method based on multi-modal information fusion
CN119359934A
Sparse view angle surface reconstruction method based on cross-view angle information complementation
CN119693546A
News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data
CN119904592A
Neo 360: neural fields for sparse view synthesis of outdoor scenes
US20240171724A1
Cited By
NeRF modeling method based on multi-scale voxel fusion and total variation regularization
CN121236302A
Three-dimensional real scene model construction method and system for inclined wall dam construction site
CN121330209A
Method and system for reconstructing physical field based on near-wall velocity field and sparse observation points
CN121746165A
Color image and depth image fusion processing method and computer equipment
CN121746445A
Multi-view three-dimensional reconstruction method based on video diffusion model
CN122066871A