3D reconstruction of outdoor scenes with sparse views based on multi-view stereo and mesh

Through the sparse view reconstruction method of multi-view stereoscopic and mesh, the problem of high-quality reconstruction of boundless outdoor scenes under sparse view is solved, and a high-precision and low-consumption three-dimensional reconstruction effect is achieved.

CN120259559BActive Publication Date: 2025-08-15CHANGCHUN UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510732701.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-15
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the prior art, under the conditions of sparse view input, it is difficult to realize high-quality three-dimensional reconstruction of boundary-free outdoor scenes, and there are problems such as blurred boundary, artifacts, high memory consumption and lack of fine structures in rendered images.

Method used

A sparse view outdoor scene three-dimensional reconstruction method based on multi-view stereoscopic and mesh is adopted. By acquiring image data and sparse three-dimensional point clouds, the cost body and depth probability body are constructed, and combined with light sampling and spatial constraints, feature fusion and end-to-end optimization are performed to achieve high-precision reconstruction of bounded-free scenes.

Benefits of technology

Effectively covering boundless scenes improves reconstruction accuracy and visual reality, reduces memory consumption, avoids details loss and structural distortion, and improves the ability to express long-term details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259559B_ABST
    Figure CN120259559B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of three-dimensional reconstruction technology in computer graphics, and is particularly concerned with a method for three-dimensional reconstruction of sparse-view outdoor scenes based on multi-view stereo and grids. The method aims to address the problems of existing reconstruction technologies, such as input dependence on dense views, blurring artifacts easily generated by rendering results, and high memory usage for large-scale scenes. The method is particularly suitable for high-precision reconstruction of outdoor unbounded scenes with sparse view input. The implementation steps include: obtaining multi-view images of the scene to be reconstructed and the corresponding camera poses, and determining the scene center based on a K-means clustering algorithm; extracting input view features, generating a cost volume based on a multi-view stereo framework, and obtaining a scene depth distribution prior in combination with a probabilistic depth estimation method; guiding voxel grid sampling based on the depth distribution, and using a spatial contraction function to map light rays that exceed the grid boundary to a limited grid range; extracting local image features and adaptively fusing them with global voxel grid features; and finally generating a reconstructed image through volume rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional reconstruction in computer graphics, and in particular relates to a three-dimensional reconstruction method of a sparse view outdoor scene based on multi-view stereo and grid. Background Art

[0002] With the advancement of computer vision, 3D reconstruction and novel view synthesis have been widely used in fields such as virtual reality, robotic navigation, autonomous driving, and medical image analysis. As a novel 3D scene reconstruction technology, Neural Radiance Fields (NeRF) has demonstrated remarkable capabilities in novel view synthesis and photorealistic rendering. However, the traditional NeRF network structure uses a scene-by-scene independent optimization strategy, and prior knowledge cannot be shared between different scenes. Therefore, a large number of views must be input for training and optimization during the reconstruction process. Furthermore, NeRF's scene representation based on multi-layer perceptrons (MLPs) suffers from slow training and rendering speeds. These issues limit the practical application of NeRF in large-scale scene reconstruction.

[0003] For 3D reconstruction from sparse view input, existing multi-view stereo (MVS)-based methods are limited by viewport coverage and the number of input views, resulting in blurring and artifacts in the boundary regions of the rendered image. Recent research has improved training and rendering efficiency by introducing explicit mesh-based representations. However, large-scale, boundless scenes still face challenges such as high memory consumption, insufficient mesh resolution leading to loss of fine structure, and loss of background details. Therefore, achieving fast, efficient, and high-quality reconstruction of boundless outdoor scenes from sparse view input is key to large-scale 3D reconstruction based on neural radiance fields.

[0004] The existing public technology publication number is "CN119359934A", and its name is "A generalizable neural radiation field reconstruction method based on multimodal information fusion". This method generates a multimodal neural encoder through progressive complementary fusion, and converts the encoder and the original RGB pixel body into volume density and radiation brightness; then uses the transformer network to fuse the contextual features of the sampled light across views, and decodes the generated density and radiation values. This method improves the accuracy of generalized neural radiation field surface reconstruction while solving the shape radiation ambiguity problem. However, the large-scale parameter calculation of the multimodal feature fusion network and transformer network constructed by this method significantly increases the memory and computing power overhead, especially when processing high-resolution large-scale scenes. In addition, sparse geometric supervision may be unevenly distributed in texture-free areas, resulting in local over-smoothing or missing details in the reconstruction results. Summary of the Invention

[0005] (1) Technical problems solved

[0006] In response to the shortcomings of the existing technology, the present invention provides a sparse view three-dimensional reconstruction method for outdoor scenes based on multi-view stereo and grid, so as to overcome the dependence of traditional neural radiation fields on dense view input and improve the surface reconstruction accuracy of borderless outdoor scenes under sparse view input conditions.

[0007] (2) Technical solution

[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0009] The method for 3D reconstruction of sparse view outdoor scenes based on multi-view stereo and grid includes the following steps:

[0010] Step 1: Obtain an image dataset of the scene to be reconstructed, which includes image data, camera pose parameters, and sparse 3D point cloud information of the scene, and use a bounding box to enclose the scene;

[0011] Step 2: Extract multi-scale image features of the input view, construct a cost volume based on the multi-view stereo (MVS) framework, regularize the cost volume to generate a neural encoding volume and a depth probability volume, and obtain the depth distribution prior of the scene through a probabilistic depth estimation method;

[0012] Step 3, light sampling and spatial constraints: adaptively sample the light and use a proportional spatial contraction function to remap the unbounded scene light to confine the space to a finite grid.

[0013] Step 4: Extract the local features of the target image and the global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through the attention network;

[0014] In step 5, the fused features and neural encoding volume features are decoded, the volume density and color value of the sampling point are output, the pixel RGB value is calculated based on the volume rendering equation, and finally end-to-end optimization is performed by combining photometric consistency supervision and sparse geometric supervision of the grid.

[0015] Furthermore, in step 1, the bounding box is used to define the geometric boundary of the scene. This process requires determining the size of the bounding box range. First, the sparse point cloud in the dataset is divided into several clusters based on the K-means clustering method, and the main cluster with the highest point density is selected. Its center of mass is calculated as the center point of the scene, and the minimum and maximum values of the main cluster on the XYZ axis are extracted to generate the coordinates of the compact bounding box vertices, thereby obtaining the bounding box range.

[0016] Furthermore, in step 2, the cost body construction and regularization process are as follows:

[0017] Select M pictures as input images, use 2D CNN to input view Extract 2D neural features and obtain feature maps ,in represents the real number space, and Indicates the dimension of the matrix; obtains the camera parameters ,in is the camera intrinsic parameter matrix, and Represent the rotation and translation matrices of the camera respectively, and then use the homography transformation matrix to transform all feature maps Transform to the reference view's conical stereo space, homography transformation matrix As shown below:

[0018]

[0019] in represents the normal direction of the reference view, represents the reference view camera parameters, Represents the input view camera parameters, and the depth is obtained through homography transformation Input view Warp feature map to the reference view :

[0020]

[0021] in Represents the pixel coordinates of the reference view, and the cost volume at the reference view is generated by calculating the variance of the deformed feature map , the formula is as follows:

[0022]

[0023] in Represents variance calculation, using 3D CNN for the reference view cost volume Regularize and generate a deep probability body and neural encoding body, where 3D CNN is composed of 3D U-Net to process the spatial information of the data; as shown in the formula, the obtained probability body is used , find the expectation along the depth direction, and you will get the initial depth value of the corresponding pixel. :

[0024]

[0025] in Indicates the number of depth channels, Indicates that the pixel is in the depth plane The probability on , find the expectation for each pixel and obtain the depth distribution of the image.

[0026] Furthermore, in step 3, adaptive sampling is to guide the light to perform more accurate scene surface sampling by using the image depth distribution information, reduce the invalid area sampling points in the three-dimensional space, and perform non-uniform sampling through the designed metric distance mapping function, so that the light sampling points It is expressed as follows:

[0027]

[0028] in represents the origin of the light, Indicates the direction of light, Represents the metric distance, and implements a non-uniform sampling strategy by parameterizing the metric distance, so that the distribution of sampling points matches the characteristics of the scene, and the metric distance and normalized distance The mapping function is expressed as follows:

[0029]

[0030] in They are the near plane distance and far plane distance defined by the depth value respectively;

[0031] In addition, the spatial contraction function is designed based on the ray-voxel intersection detection theorem, which converts the sampling points on any ray path into Isometric compression into a finite grid space achieves global linear scaling of the unbounded scene space. The specific definition is as follows:

[0032]

[0033]

[0034] in Respectively represent the coordinates of the incident point and the exit point of the light entering the bounding box, represents the ray index, Indicates the coordinates of the first sampling point of the light, Indicates the scaling ratio of the light, by performing Calculation, the light is compressed into a compact limited space.

[0035] Furthermore, the global-local feature dynamic fusion process in step 4 is as follows:

[0036] First, global features are extracted, and a global spatial radiation field with a continuous grid is constructed based on tensor decomposition. The global consistency confidence feature vector and appearance feature vector are obtained by interpolation of discrete voxels adjacent to the query point. The feature vector is combined with the color matrix Multiply and combine with the light direction Through the decoding function Get global features and global consistency confidence , the calculation formula is as follows:

[0037]

[0038] The decoding function Using a multilayer perceptron network, represents the interpolated appearance feature vector, represents the confidence feature vector obtained by interpolation, represents the dimension of spatial tensor decomposition, represents the tensor decomposition channel index, Represents cascade operation, global consistency confidence Used to evaluate the geometric consistency, texture richness, and occlusion of grid positions in multiple views, and dynamically adjust the weights of global and local features;

[0039] Then extract local features, for any queried three-dimensional point, project it to the multi-resolution target image, extract the two-dimensional plane features of each scale by bilinear interpolation, cascade the features, and generate a length of The eigenvector of ,in Represents the single-scale feature dimension;

[0040] Then use a lightweight attention network , with global features With local features and confidence As input, generate fusion weights Based on this weight, the adaptive fusion of global and local features is achieved. The calculation formula of the fusion feature is as follows:

[0041]

[0042]

[0043] The Sigmoid function is used to limit the weight to [0,1] to achieve a smooth transition. Represents the final fusion feature.

[0044] Furthermore, in step 5, the decoding process is based on the multi-layer perceptron MLP to predict the volume density of the sampling point and color values ; among them The sampling points are expressed as:

[0045]

[0046] in Represents the neural encoding body features, which are used to infer more appearance perception information from the cost body;

[0047] The volume rendering equation is expressed as the volume density of the sampling point along each ray and color values Perform cumulative summation to get the light direction Corresponding pixel RGB value , where the volume rendering equation is as follows:

[0048]

[0049] in represents the cumulative transmittance, Represents the distance between adjacent sampling points, Indicates the number of light sampling points;

[0050] Photometric consistency supervision and sparse geometric supervision of the mesh are used to calculate the color loss and voxel grids Regularization, and finally weighted summation to obtain the final loss function:

[0051]

[0052] in is the weight coefficient.

[0053] (3) Beneficial effects

[0054] Compared with the existing technology, the present invention provides a sparse view 3D reconstruction method for outdoor scenes based on multi-view stereo and grid, which has the following advantages:

[0055] The present invention proposes an explicit-implicit hybrid radiance field representation method for sparse input. The radiance field can provide global viewport coverage and effectively capture areas missed by traditional multi-view stereo vision methods, thereby ensuring comprehensive coverage of the scene and the integrity of geometric texture information during the three-dimensional reconstruction process.

[0056] This paper designs an adaptive global-local feature fusion method that achieves the coordinated optimization of local texture details and global geometric constraints, effectively avoiding the detail loss and structural distortion problems caused by traditional methods using a single cost volume. Furthermore, the multi-level representation characteristics of the voxel grid significantly reduce geometric shape estimation errors, improving the accuracy and visual realism of 3D reconstruction.

[0057] The present invention proposes a method for representing unbounded scenes. It uses a spatial proportional shrinkage function to map sampling points beyond the grid boundary into a limited grid range without changing the direction of light, so that the voxel grid can effectively cover the entire unbounded scene area and improve the ability to express distant details.

[0058] The present invention uses the K-means clustering algorithm to optimize the spatial bounding box, compresses the redundant grid space, and adopts the optimization method of depth-guided ray sampling to reduce memory consumption while maintaining geometric accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The overall flow chart provided by the present invention;

[0060] Figure 2 This is the overall framework diagram of the 3D reconstruction method provided by the present invention;

[0061] Figure 3 A comparison diagram of the sampling point distribution provided by the present invention;

[0062] Figure 4 A schematic diagram of the space contraction function provided by the present invention;

[0063] Figure 5 Schematic diagram of multi-resolution geometric feature extraction provided by the present invention;

[0064] Figure 6 A comparison chart of scene reconstruction results on the Free dataset provided by the present invention;

[0065] Figure 7 This is a comparison chart of scene reconstruction results on the Free dataset provided by this invention. DETAILED DESCRIPTION

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0067] Example

[0068] like Figure 1 FIG. 1 is a flowchart of a method for 3D reconstruction of an outdoor scene using sparse views based on multi-view stereo and grids, according to an embodiment of the present invention. The method specifically includes the following steps:

[0069] Step 1: Obtain an image dataset of the scene to be reconstructed, including image data, camera pose parameters, and sparse 3D point cloud information of the scene, and determine the range of the scene bounding box;

[0070] Prepare the Free dataset, provided by F2-NeRF, consisting of seven outdoor scenes. Compared to widely used outdoor scenes such as LLFF and NeRF-360-V2, this dataset features narrow and long camera tracks and focused foreground objects. The Free dataset contains image data, corresponding image camera poses obtained by COLMAP, and sparse 3D point cloud information about the scene (containing point coordinates, colors, and associated image indices). During training, evaluation of the Free dataset follows the F2-NeRF split between training and testing, using one-eighth of the images for testing and the remaining for training.

[0071] When using a bounding box for a scene, the center of the bounding box is usually set as the coordinate origin. However, when the scene is a forward scene, the scene geometry distribution is highly concentrated in half of the bounding box, resulting in a large number of invalid grids occupying memory resources and reducing the efficiency of light sampling. Therefore, the present invention proposes a method based on the K-means clustering algorithm to redefine the center and vertex of the bounding box. K-means is an unsupervised clustering algorithm, whose core goal is to divide data points into m clusters so that the similarity of data points in the same cluster is maximized and the difference between different clusters is maximized. According to the characteristics of the forward scene, the number of clusters m is set, and for each point cloud , calculate its centroid with each cluster Euclidean distance :

[0072]

[0073] Point Cloud Assign to the closest cluster The point cloud is divided into clusters by updating the centroid based on the mean coordinates of the points within the cluster. The main cluster with the highest point density is selected, and its centroid is calculated as the scene center. The minimum and maximum values of the main cluster along the X, Y, and Z axes are extracted to generate the coordinates of the bounding box vertices. Optimizing the bounding box through K-means clustering significantly compresses redundant grid space, improving light sampling efficiency and reconstruction quality, and providing an efficient data foundation for subsequent feature fusion and volume rendering.

[0074] Step 2: Extract multi-scale image features of the input view, construct a cost volume based on the multi-view stereo (MVS) framework, regularize the cost volume to generate a neural encoding volume and a depth probability volume, and obtain the depth distribution prior of the scene through the probabilistic depth estimation method; the specific process is as follows:

[0075] Ⅰ. Construct the cost volume: Select M images as input images, use 2D CNN to construct the input images Extract 2D neural features, where represents the real number space, Indicates the dimension of the matrix. Figure 2 As shown, 2D CNN consists of three layers of downsampling convolution layers, and finally obtains the input image features , the matrix dimension size is . Get camera parameters ,in is the internal parameter array of the camera, and Represent the rotation and translation matrices of the camera respectively, and use the homography transformation matrix to transform each feature map Transform to the reference view's conical stereo space, homography transformation matrix As shown below:

[0076]

[0077] in represents the normal direction of the reference view, represents the reference view camera parameters, Represents the input view camera parameters. Depth is obtained through homography transformation Input view Warp feature map to the reference view :

[0078]

[0079] in Represents the pixel coordinates of the reference view, and the distorted feature map is calculated by variance To generate the reference view cost volume , the formula is as follows:

[0080]

[0081] II. Depth distribution prediction: Use 3D CNN to regularize the reference view cost volume and process it into a depth probability volume And neural code body, where 3D CNN is composed of 3D U-Net to process the spatial information of the data. For the pixel points in the target view , Indicates that the pixel is in the depth plane As shown in the formula, the depth value of the pixel Defined as the expected value of the depth probability distribution along the depth direction:

[0082]

[0083] in Represents the number of depth channels. By taking the expectation for each pixel, we get the depth distribution of the image.

[0084] Step 3, light sampling and spatial constraints:

[0085] Ⅰ. Guide light sampling based on depth distribution: After obtaining the depth distribution of the image, the depth information corresponding to the pixel point is transferred to the voxel grid space for light sampling. Unlike traditional NeRF, the method of the present invention defines the nearest plane distance and the farthest plane distance for each ray based on the actual depth value, rather than using the globally fixed nearest plane distance and farthest plane distance. Therefore, the sampling strategy based on pixel-level depth information can achieve more accurate scene surface sampling, reduce the number of blank area sampling points in the three-dimensional space, and reduce the number of light sampling points. It is expressed as follows:

[0086]

[0087] in represents the origin of the light, Indicates the direction of light, Represents a metric distance.

[0088] II. Metric distance parameterization: In the reconstruction of unbounded scenes, the sampling strategy is crucial to efficiency and quality. Studies have shown that the non-uniform sampling strategy conforms to the characteristics of scene distribution: the near-view area contains rich geometric and texture details, and high-density sampling is required to ensure reconstruction accuracy; the distant area has fewer details, and oversampling will waste computing resources. However, the parameterized metric distance equation of Mip-NeRF 360 has a sudden change in the sampling interval at the critical point, resulting in an uneven sampling distribution, and the reconstruction results are prone to banding artifacts or texture breaks. To this end, the present invention proposes a new metric distance and normalized distance The mapping function of , to optimize the sampling distribution:

[0089]

[0090] in They are the near plane distance and far plane distance defined by the depth value. Figure 3 As shown in the left figure, the Mip-NeRF 360 sampling point distribution is too dense in the near-field area and too sparse in the far-field area, with a sudden change in the sampling interval at the critical point. The right figure shows the optimized sampling point distribution of the present invention, which maintains the characteristics of dense near-field and sparse far-field while eliminating the sudden change in the sampling interval, significantly improving smoothness and numerical stability.

[0091] III. Unbounded scene representation: Figure 4As shown in (a), the dynamic range of the original ray length in the unbounded scene is extremely large and lacks geometric constraints. Traditional NeRF uses NDC conversion to map the forward unbounded scene to the cubic space. For 360° scenes, Mip-NeRF 360 maps the distant space to a bounded sphere through a shrinkage function, but there is a problem of background straight lines being distorted into curves, which makes it difficult to detect the intersection of rays and voxel grids. MERF proposes a piecewise function mapping scheme, but it produces ray discontinuities at the boundaries, resulting in uneven sampling distribution and reconstruction artifacts. To this end, the present invention performs global linear scaling on the scene space coordinates while maintaining the original propagation direction vector of the ray. Based on the ray and voxel intersection detection theorem, the proportional factor of the diagonal length of the scene bounding box and the side length of the grid space is calculated, and the sampling points on any ray path are Isometric compression into a finite grid space is defined as follows:

[0092]

[0093]

[0094] in Respectively represent the coordinates of the incident point and the exit point of the light entering the bounding box, represents the ray index, Indicates the coordinates of the first sampling point of the light. Indicates the scale of the light. Figure 4 As shown in (b), the proposed space contraction function does not change the direction of the light by performing The computation maps the boundless scene into a finite voxel grid space. This property enables the use of standard ray-AABB intersection tests. Furthermore, spatial contraction optimizes the distribution of sampling points, improving the reconstruction quality of distant scene details while effectively addressing the difficulty in training convergence caused by the large variance in sampling point coordinates in boundless scenes.

[0095] Step 4: Extract the local features of the target image and the global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through the attention network;

[0096] Traditional MVSNeRF uses limited views to construct a single cost, which often faces problems such as insufficient image edge shape reconstruction accuracy and incorrect geometric shapes in non-occluded areas. To this end, the present invention proposes an explicit-implicit radiation field based on dynamic fusion of global-local features.

[0097] I. Global feature query: To reduce the memory usage of voxel grids, the present invention constructs a global spatial radiation field with a continuous grid based on tensor decomposition. Specifically, the radiation field is expressed as a set of compact vectors and matrices. In order to represent discrete voxels in the continuous radiation field representation, the discrete voxels adjacent to the query point are linearly interpolated and bilinearly interpolated to obtain the global consistency confidence feature vector and appearance feature vector of the sampling point. The formula is as follows: the feature vector is combined with the color matrix Multiply and combine with the light direction Through the decoding function Get global features and global consistency confidence .

[0098]

[0099] The decoding function Using a multilayer perceptron network, represents the interpolated appearance feature vector, represents the confidence feature vector obtained by interpolation, represents the dimension of spatial tensor decomposition, represents the tensor decomposition channel index, Represents cascade operation. Global consistency confidence This feature evaluates the geometric consistency, texture richness, and occlusion of a grid location across multiple views, dynamically adjusting the weights of global and local features. High-confidence regions indicate that the geometric information at that grid location is highly consistent and reliable across multiple views (e.g., unobstructed, texture-rich areas). In these cases, global features should be prioritized to suppress potential noise from local views. Low-confidence regions, on the other hand, indicate potential occlusion, low texture, or inter-view inconsistency at that location. In these cases, local features of the target view should be relied upon to avoid noise interference from global features.

[0100] II. Local feature extraction: 2D plane features can provide strong semantic information for 3D perception image synthesis and 3D reconstruction. In order to achieve generalization between scenes and solve the rasterization problem of low-resolution voxel grids, this paper proposes to extract multi-resolution features from the target image as local features. .like Figure 5 As shown in the figure, for any queried 3D point, it is projected to different resolution levels of the target image, the corresponding 2D plane features are extracted by bilinear interpolation, and the interpolation features of the three scales are cascaded to construct a length of The eigenvector of ,in Represents the single-scale feature dimension.

[0101] III. Dynamic fusion of global and local features: For any spatial sampling point , interpolate the global features of the query point within the voxel from the neighboring voxels ,like Figure 2 As shown, a lightweight attention network is used to local features As input, As a gating factor, the fusion weight is dynamically adjusted. Generate fusion weight Based on this weight, the dynamic fusion of global and local features is achieved. The specific process is as follows:

[0102] First, through the attention network , generate fusion weights :

[0103]

[0104] The Sigmoid function is used to map the weights to [0,1] to achieve a smooth transition. Finally, the global and local features are weighted summed to obtain the final fusion feature. :

[0105]

[0106] The attention module dynamically fuses the global features of the voxel grid with the local features of the target view, giving priority to information in high-confidence areas. At the same time, the feature weights are updated in real time based on the geometric consistency of the voxel grid, and global features are used to fill the missing geometric information in low-confidence areas, avoiding the problems of holes or artifacts caused by insufficient local features in traditional methods.

[0107] Step 5: Decode the fused features and neural encoding volume features, output the volume density and color value of the sampling point, calculate the pixel RGB value based on the volume rendering equation, and finally perform end-to-end optimization by combining photometric consistency supervision and sparse geometric supervision of the grid;

[0108] In order to infer more appearance perception information from the cost body, trilinear interpolation is used to extract corresponding features from the neural encoding body for the query sampling points, which is expressed as ,Will and fusion features Combined with light direction Obtain the volume density of the sampling point through MLP network decoding and color values , which is expressed as follows:

[0109]

[0110] After obtaining the color value and volume density prediction value of the sampling point, the sampling points are accumulated and summed along each ray, and the light direction is calculated by the volume rendering equation Corresponding pixel RGB value :

[0111]

[0112] in represents the cumulative transmittance, Represents the distance between adjacent sampling points, Indicates the number of sampling points for a single ray.

[0113] Due to the differentiability of the overall network structure, the multi-view stereo stage can be jointly optimized and trained with the neural radiance field constructed from the voxel grid. The training uses the mean square error loss function to calculate the photometric consistency loss, which is defined as:

[0114]

[0115] in In order to encourage the sparsity of tensor factor parameters and make the norm of the two density tensors as small as possible, the standard Regularization is used to improve the quality of the extrapolated view and remove floating points in the final rendering. The expression is:

[0116]

[0117] in represents the matrix, represents a vector, Represents the total number of voxel grid parameters, which will be used to lose photometric consistency and The weighted sum of the losses gives the final loss function:

[0118]

[0119] in, Represents the weight coefficient, set in the experiment .

[0120] The experimental comparison is conducted on the Free dataset, which contains 7 outdoor scenes with a resolution of 1461×935. Model training and testing are performed on a single RTX 4090 GPU. For each scene optimization, 10,000 iterations of training are performed using the Adam optimizer with an initial learning rate of 5e. -4 and exponential scheduler. During the training process, the batch_size is 4096 rays, the number of sampling points is 8, the number of input views is 3, and the rest of the network parameter settings follow the parameter settings in MVSNeRF. The grid adopts a coarse-to-fine reconstruction strategy, from =128 3Starting with an initial low-resolution grid, linear upsampling is performed for a fixed number of rounds to increase the grid resolution, and finally the =300 3 The grid resolution is 10000 pixels. PSNR, SSIM, and LPIPS are used as evaluation indicators for rendering results.

[0121] On the Free dataset, our proposed method was compared with baseline methods, including the grid-based method TensoRF, and the sparse view reconstruction methods MVSNeRF, ENeRF, FreeNeRF, and BoostMVSNeRF. The results are shown in Table 1. Our proposed method performs best in terms of rendering quality on this dataset. Figure 6 and Figure 7 The qualitative comparison results show that TensoRF has poor reconstruction quality due to the low-resolution grid; MVSNeRF and its variants have problems such as incorrect geometric shapes and severe blurring at the boundaries of the rendered images due to the locality of multi-view geometric constraints and the sparseness of view coverage. In contrast, the method proposed in this paper uses voxel grids to construct multi-view global features, and through a dynamic fusion mechanism of global-local features to fill in the missing geometric information in low-confidence areas, effectively solving the edge blur and geometric distortion problems under sparse views, and achieving better rendering quality. In addition, combined with the proposed spatial constraint processing, the method of this invention is superior in the ability to express distant details.

[0122] Table 1 Quantitative comparison results with the latest methods on the Free dataset

[0123] Comparison Method PSNR↑ SSIM↑ LPIPS↓ TensoRF 21.11 0.481 0.316 MVSNeRF 17.26 0.433 0.529 ENeRF 24.42 0.797 0.218 FreeNeRF 26.09 0.774 0.278 BoostMVSNeRF 25.64 0.835 0.193 The present invention 26.17 0.836 0.182

[0124] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A sparse view outdoor scene 3D reconstruction method based on multi-view stereo and grid, characterized by: The following steps are involved: Step 1: Obtain an image dataset of the scene to be reconstructed, which includes image data, camera pose parameters, and sparse 3D point cloud information of the scene, and use a bounding box to enclose the scene; Step 2: Extract multi-scale image features of the input view, construct a cost volume based on the multi-view stereo framework, regularize the cost volume to generate a neural encoding volume and a depth probability volume, and obtain the scene's depth distribution prior through a probabilistic depth estimation method; Step 3, light sampling and spatial constraints: adaptively sample the light and use a proportional spatial contraction function to remap the unbounded scene light to confine the space to a finite grid. Step 4: Extract the local features of the target image and the global voxel grid features corresponding to the sampling points, and dynamically fuse the global and local features through the attention network; Step 5: Decode the fused features and neural encoding volume features, output the volume density and color value of the sampling point, calculate the pixel RGB value based on the volume rendering equation, and finally perform end-to-end optimization by combining photometric consistency supervision and sparse geometric supervision of the grid; In step 3, adaptive sampling uses image depth distribution information to guide light for more accurate scene surface sampling, reduces invalid area sampling points in three-dimensional space, and performs non-uniform sampling through the designed metric distance mapping function. Therefore, the light sampling point x is expressed as follows: x=o+lr Where o represents the origin of the light, r represents the direction of the light, and l represents the metric distance. The non-uniform sampling strategy is implemented by parameterizing the metric distance so that the distribution of sampling points matches the characteristics of the scene. The mapping function between the metric distance l and the normalized distance s is expressed as follows: l=L n +(L f -L n )s 4 Among them L n ,L f They are the near plane distance and far plane distance defined by the depth value respectively; In addition, the spatial contraction function is designed based on the ray-voxel intersection detection theorem, which converts the sampling point x on any ray path into j Isometrically shrinking to a finite grid space, thereby achieving global linear scaling of the unbounded scene space, the specific definition is as follows: contract(x j )=q enter,j +k j ·(x j -x n,j ) where q enter ,q exit They represent the coordinates of the incident point and the exit point of the light entering the bounding box, j represents the light index, x n,j Represents the coordinates of the first sampling point of the light, k represents the scaling ratio of the light, and by performing contract calculation on the sampling points, the light is compressed into a compact finite space; The global-local feature dynamic fusion process in step 4 is as follows: First, the global features are extracted. A global spatial radiation field with a continuous grid is constructed based on tensor decomposition. The global consistency confidence feature vector and appearance feature vector are obtained by interpolation of discrete voxels adjacent to the query point. The feature vector is multiplied by the color matrix B and combined with the light direction r to obtain the global feature F through the decoding function S. global and global consistency confidence C V , the calculation formula is as follows: The decoding function S uses a multi-layer perceptron network, A f represents the interpolated appearance feature vector, A c represents the confidence feature vector obtained by interpolation, m represents the dimension of spatial tensor decomposition, τ represents the tensor decomposition channel index, Represents cascade operation, global consistency confidence C V Used to evaluate the geometric consistency, texture richness, and occlusion of grid positions in multiple views, and dynamically adjust the weights of global and local features; Then extract local features. For any queried 3D point, project it to the multi-resolution target image, extract 2D plane features of each scale through bilinear interpolation, concatenate the features, and generate a feature vector of length 3×C as the geometric local feature F. local , where C represents the single-scale feature dimension; Then use a lightweight attention network With the global feature F global With local features F local and confidence C V As input, a fusion weight ω is generated. Based on this weight, adaptive fusion of global and local features is achieved. The weight ω and fusion feature calculation formula are as follows: F fused =ω·F global +(1-ω)·F local The Sigmoid function is used to limit the weight to [0,1] to achieve a smooth transition. fused Represents the final fusion feature.

2. The method for 3D reconstruction of outdoor scenes based on multi-view stereo and grid-based sparse views according to claim 1, characterized in that: In step 1, the bounding box is used to define the geometric boundary of the scene. This process requires determining the size of the bounding box range. First, the sparse point cloud in the dataset is divided into several clusters based on the K-means clustering method. The main cluster with the highest point density is selected, and its center of mass is calculated as the center point of the scene. The minimum and maximum values of the main cluster on the XYZ axis are extracted to generate the coordinates of the compact bounding box vertices, and then the bounding box range is obtained.

3. The method for 3D reconstruction of outdoor scenes based on multi-view stereo and grid-based sparse views according to claim 1, characterized in that: In step 2, the cost body construction and regularization process are as follows: Select M pictures as input images, use 2D CNN to input view Extract 2D neural features and obtain feature maps in Represents the real space, H×W×3 and H / 4×W / 4×C represent the dimensions of the matrix; obtain the camera parameters [K,R,t], where K is the camera intrinsic matrix, R and t represent the rotation and translation matrices of the camera respectively, and then use the homography transformation matrix to transform all feature maps F i Transform to the reference view's conical stereo space, homography transformation matrix As shown below: Where n1 represents the normal direction of the reference view, [K1, R1, t1] represents the reference view camera parameters, [K i ,R i ,t i ] represents the input view camera parameters, and the distortion feature map F from the input view i to the reference view at depth z is obtained by homography transformation i,z : Where (u, v) represents the pixel coordinates of the reference view. The cost volume V at the reference view is generated by calculating the variance of the deformed feature map. The formula is as follows: V(u,v,z)=Var(F i,z (u,v)) Var represents the variance calculation. 3D CNN is used to regularize the reference view cost volume V to generate a depth probability volume P and a neural encoding volume. The 3D CNN is composed of a 3D U-Net to process the spatial information of the data. Using the obtained probability volume P, the expectation is calculated along the depth direction to obtain the initial depth value L(u,v) of the corresponding pixel point: Where D represents the number of depth channels, P i (u,v) indicates the pixel is on the depth plane d i The probability on , find the expectation for each pixel and obtain the depth distribution of the image.

4. The method for 3D reconstruction of outdoor scenes based on multi-view stereo and grid-based sparse views according to claim 3, characterized in that: In step 5, the decoding process is based on the multi-layer perceptron MLP to predict the volume density σ of the sampling point i and color value c i ; The i-th sampling point is expressed as: c i ,σ i =MLP(F fused ,f vol ,r) where f vol Represents the neural encoding body features, which are used to infer more appearance perception information from the cost body; The volume rendering equation is expressed as the volume density σ of the sampling point along each ray i and color value c i Perform cumulative summation to obtain the pixel RGB value corresponding to the light direction r The volume rendering equation is as follows: in represents the cumulative transmittance, δ i Represents the distance between adjacent sampling points, and N represents the number of light sampling points; The photometric consistency supervision and the sparse geometric supervision of the mesh are used to calculate the color loss L col And voxel grid L1 regularization, and finally weighted summation to get the final loss function: Loss=L col +λ·L1 Where λ is the weight coefficient.

Citation Information

Patent Citations

  • Generalized neural radiation field reconstruction method based on multi-modal information fusion

    CN119359934A

  • Nerve radiation field three-dimensional reconstruction method based on space mixed attention

    CN118918250A

  • News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data

    CN119904592A