Method for improving 3D scene reconstruction and new view synthesis precision
By combining multi-resolution hash grid coding and a priori MLP modules, the problems of high computational complexity and insufficient accuracy of the NeRF method in complex scenes are solved, and efficient 3D scene reconstruction and new view synthesis are achieved.
Patent Information
- Application Number
- CN202510642236.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing NeRF methods have high computational complexity and slow rendering speed when processing complex scenes, and the lack of explicit geometric encoding leads to insufficient perspective synthesis accuracy.
Combining multi-resolution hash grid coding and prior MLP modules, hierarchical geometric features are extracted through the multi-resolution hash grid coding branch, and heuristic feature layers and segmentation mask layers are introduced to optimize the model representation and generate high-fidelity synthetic images.
It significantly improves the accuracy and efficiency of 3D scene reconstruction and new view synthesis, increases the peak signal-to-noise ratio (PSNR) by 3.07%, and performs more robustly in complex scenes.
Smart Images

Figure CN120689498A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scene reconstruction, and in particular to a method for improving the accuracy of 3D scene reconstruction and new view synthesis. Background Art
[0002] Neural Radiance Fields (NeRF) is a groundbreaking new perspective synthesis technique that utilizes a multi-layer perceptron (MLP) to reconstruct high-quality 3D scenes from sparse 2D input data. By estimating 3D radiance fields and camera pose, NeRF achieves accurate and realistic 3D scene reconstruction. However, NeRF is computationally expensive, prompting researchers to continuously explore methods to optimize its performance. In recent years, efficient NeRF variants such as Instant-ngp have introduced techniques such as hash coding, significantly accelerating rendering speed, enabling near-instant perspective synthesis, and expanding the application scope of NeRF in real-time environments. These innovations have promoted the application of NeRF in dynamic scene modeling. For example, Dynamic NeRF can capture moving objects in complex environments. Furthermore, NeRF is increasingly being used in medical imaging. For example, high-resolution 3D models can be constructed from CT and MRI scan data to assist in diagnosis and surgical planning.
[0003] NeRF relies on a multi-layer perceptron (MLP) to approximate the implicit function of a 3D scene. However, the inherent limitations of MLPs result in a bottleneck in view synthesis speed. Since each spatial sampling point requires forward propagation calculations, this results in high computational overhead and long processing time. To address this issue, researchers have proposed grid-based NeRF methods that significantly improve rendering efficiency by optimizing dense voxel grid representations. Instant-ngp, among others, introduces a multi-hash encoding scheme, replacing the traditional voxel grid with a hash map to more efficiently extract 3D features. This method accelerates the rendering process by reducing memory usage and computational cost, enabling real-time scene reconstruction. Hash encoding reduces the complexity of 3D space traversal and improves training and inference efficiency. In contrast, methods such as Plenoxels and DVGO focus on optimizing voxel-based representations or using grids of varying resolutions to improve rendering quality. However, these methods often incur higher memory consumption and computational overhead, limiting their scalability and efficiency in large-scale or real-time applications.
[0004] 2D features extracted independently from multiple input viewpoints and integrated into a multi-layer perceptron (MLP) give NeRF an intuitive adaptive mechanism. However, due to the lack of explicit geometric encoding, these features are limited in their performance when processing complex scenes. To address this problem, subsequent studies have demonstrated the effectiveness of introducing geometric prior information. For example, MVSNeRF constructs a cost volume and uses a 3D convolutional neural network (CNN) to reconstruct a neural encoding volume containing the features of each voxel. GeoNeRF further enhances this framework by introducing a cascaded cost volume and integrating an attention mechanism. NeuRay is used to identify visible 3D points by calculating a visibility feature map from the cost volume or depth map. However, these cost volume-based methods are highly sensitive to the choice of reference view. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for improving the accuracy of 3D scene reconstruction and new view synthesis, aiming to combine multi-resolution hash grid coding and introduce a priori MLP modules to further enhance the model's ability in complex scene reconstruction.
[0006] To achieve the above object, the present invention provides a method for improving the accuracy of 3D scene reconstruction and new view synthesis, comprising the following steps:
[0007] Perform 3D sampling along the camera rays projected from the image pixels to obtain sampling points;
[0008] Processing the sampling points through a multi-resolution hash grid coding branch to extract hierarchical geometric features;
[0009] Introducing a priori-based sampling sub-network branch to optimize the representation through additional prior features to obtain the expression features;
[0010] Input the representation features into the sampling subnetwork to predict volume density and color;
[0011] The rendering module further simulates the propagation of light in the three-dimensional scene to generate a high-fidelity synthetic image.
[0012] The specific method of processing the sampling points through the multi-resolution hash grid coding branch to extract the hierarchical geometric features is as follows:
[0013] Obtain the indices of the surrounding voxel vertices at different resolution levels through the multi-resolution hash grid encoding branch;
[0014] The corresponding features are extracted from the hash table using a hash function and linearly interpolated according to their relative positions in the voxel grid at each resolution level;
[0015] The feature results of each resolution level are spliced together to form the multi-resolution hash feature of the input point.
[0016] Wherein, the heuristic feature layer and the segmentation mask layer are introduced into the prior-based sampling subnetwork.
[0017] The goal of introducing the heuristic feature layer and the segmentation mask layer is to incorporate additional high-level information into the model.
[0018] The segmentation mask layer guides the model to focus on specific areas, reduces background interference, and enhances the realism and clarity of the generated results.
[0019] This paper presents a method for improving the accuracy of 3D scene reconstruction and novel view synthesis. This method utilizes multi-resolution hash grids to capture fine 3D details. Furthermore, we introduce a prior-based MLP module to compensate for the missing prior knowledge and combine it with the NeRF-based rendering pipeline to generate high-fidelity output. This method combines multi-resolution hash grid encoding and a prior-guided MLP architecture to improve the accuracy of 3D scene reconstruction and rendering. The multi-resolution hash grid encoding module uses hash table embedding to capture geometric details at different scales and enhances spatial perception and computational efficiency through cross-layer feature processing. Furthermore, the prior-guided MLP uses a strategy similar to the attention mechanism to leverage prior information, thereby improving the model's ability to capture the real environment. Experimental results show that MiResMLP outperforms most existing methods in the novel view synthesis task, achieving a peak signal-to-noise ratio (PSNR) improvement of up to 3.07 compared to representative methods in this field, while also outperforming the current state-of-the-art in terms of performance and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1This is a schematic diagram of the framework of a method provided by the present invention for improving the accuracy of 3D scene reconstruction and new view synthesis, wherein scene is the scene; camerapose is the camera pose, each pose Ti is composed of five constants (x, y, z, θ, φ), and the pose and the sampled image captured by the pose constitute the complete information of the sampling point; Multi-Resolution HashGrid is a multi-hash grid encoding branch that processes the sampling points and extracts hierarchical geometric features; Prior-based MLP is a sampling subnetwork branch based on priors, composed of an heuristic feature layer and a segmentation mask layer, which optimizes the representation through additional prior features to obtain the expression features; View MLP volume density sampling subnetwork obtains volume density, while the sampling subnetwork for obtaining color consists of only one linear layer, which is not marked in the figure, so the sampling subnetwork branch can obtain volume density and color. Volume Rendering is a rendering module that simulates the propagation of light in a three-dimensional scene to generate high-fidelity synthetic images.
[0022] Figure 2 This is a diagram of an encoder mapping 3D points to a fixed-size voxel grid in a multi-resolution hash grid. The input coordinate is composed of (x, y, z, θ, φ), with only (x, y, z) shown in the figure. Trainable Feature Vectors are feature vectors at different resolutions. The hash function is used to obtain the corresponding features (Trainable Feature Vectors) of the surrounding voxel vertices at different resolutions based on the input coordinates. After interpolation and concatenation, the multi-resolution hash features are formed and used as the input vector of the NeRF model.
[0023] Figure 3 Figure 1 is a schematic diagram of the overall process of using a multi-resolution hash grid for feature extraction. For a given input coordinate, the initial step is to obtain the indices of the surrounding voxel vertices at different resolution levels (where yellow and red represent two different resolutions). Through the hash function, the corresponding features are retrieved from the hash table and linearly interpolated based on the relative position of the input coordinate in the voxel grid at each resolution level. The results at each level are then concatenated to form a multi-resolution hash feature for the input coordinate, which is used as the input to the NeRF model.
[0024] Figure 4Figure 2 is a schematic diagram of the visual comparison of synthesized scenes in Blender, where (a) the Scene column represents: scene pictures, (b) the GT column, the groundtruth column, represents: details of the real scene, (c) the NeRF column represents: details predicted by the NeRF model, (d) the column represents: details predicted by the Instant-ngp model, (e) the EGRA-ngp column represents: details predicted by the EGRA-ngp model, and (f) the Ours column represents the details predicted by the model we designed.
[0025] Figure 5 It is a schematic diagram of the quantitative comparison results based on the peak signal-to-noise ratio (PSNR), where Table 1. The PSNR of different models and scenes. represents the performance of the peak signal-to-noise ratio (PSNR) indicator of different models in different scenes; Blender and LLFF represent two different datasets; Chair, Drums, Ficus, Hotdog, Lego, Material, Mic, and Ship represent image data of different items in the Blender dataset; Room, Fern, Leaves, Fortress, Orchids, Flower, T-Rex, and Horns represent image data of different items in the LLFF dataset; SRN, NV, LLFF, NeRF, Instant, EGRA, and Ours represent different models.
[0026] Figure 6 It is a schematic diagram of the quantitative comparison results of the structural similarity index (SSIM), where Table 2. The SSIM of different models and scenes. represents the performance of the structural similarity index (SSIM) indicator of different models in different scenarios; Blender and LLFF represent two different data sets; Chair, Drums, Ficus, Hotdog, Lego, Material, Mic, and Ship represent image data of different items in the Blender data set; Room, Fern, Leaves, Fortress, Orchids, Flower, T-Rex, and Horns represent image data of different items in the LLFF data set; SRN, NV, LLFF, NeRF, Instant, EGRA, and Ours represent different models.
[0027] Figure 7It is a schematic diagram of the quantitative comparison results of the Learning Perceptual Image Patch Similarity (LPIPS), where Table 3. The LPIPS of different models and scenes. represents the performance of the Learning Perceptual Image Patch Similarity (LPIPS) indicator in different models and scenes; Blender and LLFF represent two different datasets; Chair, Drums, Ficus, Hotdog, Lego, Material, Mic, and Ship represent image data of different items in the Blender dataset; Room, Fern, Leaves, Fortress, Orchids, Flower, T-Rex, and Horns represent image data of different items in the LLFF dataset; SRN, NV, LLFF, NeRF, Instant, EGRA, and Ours represent different models.
[0028] Figure 8 Table 4 is a diagram of the peak signal-to-noise ratio of ablation experiments, where Tablet 4. The PSNR of ablation experiments is the performance of the peak signal-to-noise ratio (PSNR) indicator of the ablation experiments; Blender represents a dataset; Chair, Drums, Ficus, Hotdog, Lego, Material, Mic, and Ship represent image data of different items in the Blender dataset; Baseline represents the basic NeRF model, Hash represents the hash grid coding module, and Prior represents the prior MLP model; H&P represents the combination of the hash grid coding module and the prior MLP model.
[0029] Figure 9 is a schematic diagram of the structural similarity index of ablation experiments, where Tablet 5. The SSIM of ablation experiments is the performance of the structural similarity index (SSIM) indicator of the ablation experiment; Blender represents a dataset; Chair, Drums, Ficus, Hotdog, Lego, Material, Mic, and Ship represent image data of different items in the Blender dataset; Baseline represents the basic NeRF model, Hash represents the hash grid encoding module, and Prior represents the prior MLP model; H&P represents the combination of the hash grid encoding module and the prior MLP model.
[0030] Figure 10It is a schematic diagram of quantitative results, where Table6.The impact of heuristic layer countin the prior-based MLP on network performance. It represents the impact of the number of heuristic feature layers in the prior MLP model on network performance. Metrics: indicators, n=1 means that the number of heuristic feature layers in the prior MLP model is 1, n=2 means that the number of heuristic feature layers in the prior MLP model is 2, n=3 means that the number of heuristic feature layers in the prior MLP model is 3, PSNR means peak signal-to-noise ratio, SSIM means structural similarity index, and LPIPS means learned perceptual image block similarity.
[0031] Figure 11 This is a flow chart of a method for improving the accuracy of 3D scene reconstruction and new view synthesis provided by the present invention.
[0032] Figure 12 It is a flowchart of a specific method of processing the sampling points through a multi-resolution hash grid coding branch to extract hierarchical geometric features. DETAILED DESCRIPTION
[0033] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0034] See also Figures 1 to 12 The present invention provides a method for improving the accuracy of 3D scene reconstruction and new view synthesis, comprising the following steps:
[0035] S1 performs 3D sampling along the camera rays projected from the image pixels to obtain sampling points;
[0036] S2 processes the sampling points through a multi-resolution hash grid coding branch to extract hierarchical geometric features;
[0037] In an embodiment of the present invention, a multi-resolution hash grid coding technique is introduced into the traditional NeRF framework. This coding method can effectively capture detailed information at different resolution levels, especially in low-dimensional data structures to achieve sparse coding, so that the model can focus more on key local features in complex structures. Unlike traditional methods that map 3D points to fixed-size voxel grids, this method does not directly store features at the vertices of the cube, but organizes and stores them through hash tables, such as Figure 2 shown.
[0038] Specific method:
[0039] S21 obtains the indices of its surrounding voxel vertices at different resolution levels through the multi-resolution hash grid encoding branch;
[0040] In the embodiment of the present invention, for a given input coordinate x, it is first necessary to obtain the indices of the surrounding voxel vertices at different resolution levels, which can be expressed as: in, Figure 3 The yellow and red colors represent two different resolution levels.
[0041] S22 uses a hash function to extract the corresponding features from the hash table and performs linear interpolation based on the relative positions in the voxel grid at each resolution level;
[0042] In the embodiment of the present invention, a hash function h:V→Y is used to extract corresponding features from the hash table, and linear interpolation is performed according to the relative positions in the voxel grid at each resolution level.
[0043] S23 stitches the feature results of each resolution level to form a multi-resolution hash feature of the input point.
[0044] In the embodiment of the present invention, finally, the feature results of each resolution level are concatenated to form a multi-resolution hash feature y of the input point x, and used as the input of the NeRF model.
[0045] Specifically, the selection of each resolution level follows a geometric progression, ranging from the coarsest to the finest resolution interval [Nmin, Nmax], and the resolution of the lth layer is defined by the following formula:
[0046]
[0047] For a vertex v = (vx, vy, vz), we use the hash function and resolution settings in Instant-ngp. The specific formula is as follows:
[0048]
[0049] in, Denotes a bitwise XOR operation, with πx = 1, πy = 265, 435, 761, and πz = 805, 459, 861. πy and πz are large prime numbers used for hashing. Furthermore, the total number of parameters in a multi-resolution hash grid is constrained by the following formula: L·T·F, where L represents the number of resolution levels, T is the hash table size, and F is the feature dimension at each resolution. To balance computational efficiency and storage capacity, we set L = 2, T = 219, and F = 2, resulting in a total of 876 parameters.
[0050] S3 introduces a sampling sub-network branch based on priors, and optimizes the representation through additional prior features to obtain the expression features;
[0051] In an embodiment of the present invention, a heuristic feature layer and a segmentation mask layer are introduced into the prior sampling subnetwork branch (MLP) structure. The design goal of the heuristic feature layer is to incorporate additional high-level information into the model. These features are different from basic geometric information and usually contain pre-calculated information, which helps the model understand complex scenes. For example, in image segmentation tasks, heuristic features (such as edge detection) can help the model to more accurately parse image content, thereby improving its generalization ability on unseen data. This layer processes these features through a simple linear transformation, and its mathematical expression is as follows
[0052] h heuristic =ReLU(W heuristic ·X heuristic +b heuristic )#(6)
[0053] Among them, h heuristic represents the heuristic features of the input, W heuristic and b heuristic are the weight and bias of the feature influence, respectively. Through this linear layer, the network can introduce heuristic features into the subsequent feature processing process, thereby improving the ability to understand complex scenes.
[0054] The segmentation mask layer can be considered a "soft attention" mechanism that guides the model to focus on specific areas, reducing background interference, thereby enhancing the realism and clarity of the generated results. This design draws on the attention mechanism and noise suppression theory, and its feature processing is as follows:
[0055] m=σ(W T ·h+b)#(7)
[0056] Among them, h is the input feature vector, W and b are the weight matrix and bias term of the linear layer respectively, σ() is the Sigmoid activation function, and m is the output of the mask layer, which is used to indicate the location where the model should focus.
[0057] The advantage of this design is that it alleviates the vanishing gradient problem during deep network training while allowing low-level data to be passed directly to higher layers, preserving low-level input information. Geometric position information in shallow layers of the network helps extract local features, but these features can be lost in deeper layers. Therefore, introducing a prior-based MLP module into the NeRF model preserves the integrity of spatial information across multiple layers of the network, thereby enhancing the model's ability to capture detailed information.
[0058] S4 inputs the representation features into the sampling subnetwork to predict volume density and color;
[0059] The S5 rendering module further simulates the propagation of light in a three-dimensional scene to generate a high-fidelity synthetic image.
[0060] To better understand the present technical solution, the following examples are provided for further explanation:
[0061] Experiments verify the effectiveness of a method to improve the accuracy of 3D scene reconstruction and novel view synthesis, evaluate the performance of MiResMLP in rendering synthesis and real scene optimization tasks, and compare it qualitatively and quantitatively with the baseline method NeRF.
[0062] DataSets
[0063] Forward-Facing (LLFF): This dataset contains 8 sets of complex real-world scenes, each consisting of 20 to 64 forward-facing handheld views with an image resolution of 756 × 1008. Experiments are evaluated on every 8 images to provide a fair comparison with existing methods.
[0064] Blender Dataset: This dataset contains eight complex scenes rendered in Blender. Each scene is captured from 20 to 64 different viewpoints, providing 100 training views to improve the generalization of the model. All images are accompanied by accurate camera pose information and rendered at a resolution of 800×800. In experiments, the images are downsized to 400×400 to reduce computational overhead. The different scenes in the dataset simulate various real-world lighting conditions and complex geometric structures to increase the difficulty and adaptability of the model training.
[0065] Evaluation metrics: For real datasets, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and perceptual image similarity (LPIPS) are used as evaluation metrics. PSNR measures reconstruction quality by calculating the ratio of the mean square error of the image's peak signal to the noise. A higher value indicates a higher similarity between the reconstructed image and the original. SSIM measures image similarity across three dimensions: brightness, contrast, and structure. A value closer to 1 indicates higher quality reconstructed images. LPIPS uses a deep learning model to simulate how the human eye perceives image differences. A lower value indicates a smaller subjective visual difference between the reconstructed image and the original.
[0066] ImplementationDetails
[0067] The experiments were conducted on a PC equipped with a single NVIDIA Geforce RTX 3090 GPU (24GB). MiResMLP was used to train the basic NeRF model. The parameters of the basic NeRF model were optimized using MSE loss and Adam optimizer. Both the coarse network and the fine network used 8-layer multilayer perceptrons (MLPs) as the main feature extraction and generation modules, with each layer containing 256 neurons. The resolution of each layer was based on the calculation results of Instant-NGP
[31] . The input features included encoding of 3D position and viewpoint information through position encoding. The hash grid table size of each layer was set to 16, that is, each hash grid table contained 2^4 = 16 embedding vectors. The random sample size ($N_{rand}$) was set to 1024, the initial number of sample points was 64, and 64 additional importance sampling points were used for fine sampling to enhance the model's ability to capture high-density areas. The model enabled the viewpoint processing branch to ensure the consistency and coherence of multi-viewpoint generation. To prevent performance degradation during training, the noise standard deviation was set to 1.0. The initial learning rate during training is set to 10^{-2}. For the NeRFSynthetic and Blender datasets, 8x downsampling is applied according to the LLFF sampling scheme, with viewpoints left out every 8 frames to ensure that the model can learn more complete 3D structures from sparse viewpoints.
[0068] Experimental results
[0069] Experimental results were obtained using the same process as the baseline NeRF algorithm. Our approach was compared with the standard NeRF model across all scenes on the Blender and LLFF datasets. Notably, evaluation metrics were also monitored throughout the training process. Detailed results can be found in Tables 1-3. Of particular note, our approach, MiResMLP, demonstrates significant performance improvements, specifically in terms of PSNR scores. This demonstrates that our approach effectively enhances NeRF's ability to capture details in complex scenes. Furthermore, our approach is more accurate and robust during training.
[0070] Qualitative comparison
[0071] Figure 4 A visual comparison of scenes synthesized in Blender is presented. The results demonstrate that our model is able to recover finer appearance and geometric details, such as the surface texture of Lego bricks, the surface of sails, and the color and texture of flowers. While NeRF synthesized views are rich in textural detail, they struggle with exposure variations and inconsistent masks, resulting in inconsistent colors. By introducing a prior-based MLP module, our method not only effectively reconstructs surface details but also achieves more visually realistic color representation.
[0072] Quantitative comparison
[0073] Quantitative comparison results based on Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learning-Perceptual Patch Similarity (LPIPS) are presented in Tables 1 to 3, covering both datasets. We also provide rendering results of our method under the same experimental conditions. Our observations show that our NeRF model performs on par with most baseline methods based on numerical metrics. Specifically, our method outperforms other methods on synthetic datasets and performs nearly on par with state-of-the-art neural implicit models on real scenes. Furthermore, our trained model shows fewer errors compared to other methods on both synthetic and real scenes.
[0074] We first compare our approach with the classic NeRF, Instant-NGP, and EGRA-NGP models. Specifically, Instant-NGP uses a multi-hash encoding scheme, while EGRA-NGP builds on Instant-NGP by dynamically allocating rays based on prior information, thereby focusing on textures and sharp edges. Although the dynamic ray allocation strategy (EGRA-NGP) may bring some benefits, the rendering results still appear too smooth. In contrast, our approach guides the learning of mesh features, focusing on the surface of effective features, while integrating prior information, providing rich geometric and surface details, compensating for the scene details missing from mesh features.
[0075] For each scene, we bold the best result in the table. Our approach achieves significant performance improvements over baseline methods. Specifically, on the Blender dataset, our approach improves the overall PSNR by 12% over the baseline NeRF. Compared to state-of-the-art methods such as Instant-NGP and EGRA-NGP, our approach achieves improvements of 4.3% and 3.6%, respectively. Despite the improvements of these methods, our approach maintains a significant advantage. Similarly, on the LLFF dataset, our model consistently outperforms other methods.
[0076] Ablation experiments
[0077] We conducted ablation experiments to investigate the independent contributions of the reconstructed MLP architecture and hash trellis coding to model performance. The results are shown in Tables 4 and 5. In these tables, "Hash" represents the hash trellis coding module, and "Prior" represents the prior MLP model. To isolate the impact of other parameters, we used a simple element-wise addition test ("H&P(Add)") to evaluate network performance. We found that retaining only the reconstructed MLP or only the hash trellis coding resulted in decreased performance, but both configurations still outperformed the baseline NeRF model. However, the best performance was achieved when both features were retained. These results demonstrate that the reconstructed MLP architecture and hash trellis coding complement each other, jointly improving network performance.
[0078] Furthermore, to validate the effectiveness of our network, we analyzed the impact of the number of heuristic layers in the prior MLP on network performance. Quantitative results, shown in Table 6, demonstrate that performance improves as the number of reference views increases within a certain angular range. However, increasing the number of heuristic layers also results in longer processing time. Therefore, we set the number of heuristic layers to 1.
[0079] Through qualitative and quantitative experiments, we validate the effectiveness of our proposed method. By incorporating multi-resolution hashing grids and prior features into the classic NeRF framework, our approach effectively overcomes the limitations of previous NeRF methods in synthesizing realistic scene views, such as the lack of detail. By modeling with a multi-resolution grid that incorporates geometric properties and leveraging prior features to maximize the information extracted from 2D images, NeRF significantly enhances its overall representation capabilities in real-world scenarios. Our method also demonstrates high applicability in real-world scenarios.
[0080] This method integrates a multi-resolution hash grid encoding module into the traditional MLP framework. This improved approach not only offers significant advantages in complex real-world scenes but also results in more realistic and clear 3D scene reconstruction. Furthermore, we design a priori MLP module in which heuristic and mask layers jointly handle the extraction of prior information features and high-precision scene reconstruction. During training, the prior information complements the hash grid, resolving the suboptimal problem caused by hash collisions. This method enhances the extraction of high-frequency details and achieves high-quality scene reconstruction.
[0081] The above disclosure is merely a preferred embodiment of a method for improving the accuracy of 3D scene reconstruction and new view synthesis according to the present invention. It is certainly not intended to limit the scope of the present invention. A person skilled in the art will understand that implementing all or part of the processes of the above embodiment and making equivalent changes in accordance with the claims of the present invention still fall within the scope of the invention.
Claims
1. A method for improving the accuracy of 3D scene reconstruction and new view synthesis, characterized in that: The following steps are involved: Perform 3D sampling along the camera rays projected from the image pixels to obtain sampling points; Processing the sampling points through a multi-resolution hash grid coding branch to extract hierarchical geometric features; Introducing a priori-based sampling sub-network branch to optimize the representation through additional prior features to obtain the expression features; Input the representation features into the sampling subnetwork to predict volume density and color; The rendering module further simulates the propagation of light in the three-dimensional scene to generate a high-fidelity synthetic image.
2. The method for improving the accuracy of 3D scene reconstruction and new view synthesis according to claim 1, characterized in that: The specific method of processing the sampling points by the multi-resolution hash grid coding branch to extract the hierarchical geometric features is as follows: Obtain the indices of the surrounding voxel vertices at different resolution levels through the multi-resolution hash grid encoding branch; The corresponding features are extracted from the hash table using a hash function and linearly interpolated according to their relative positions in the voxel grid at each resolution level; The feature results of each resolution level are spliced together to form the multi-resolution hash feature of the input point.
3. The method for improving the accuracy of 3D scene reconstruction and new view synthesis according to claim 1, characterized in that ; A heuristic feature layer and a segmentation mask layer are introduced into the prior-based sampling subnetwork.
4. The method for improving the accuracy of 3D scene reconstruction and new view synthesis according to claim 3, characterized in that ; The goal of introducing the heuristic feature layer and the segmentation mask layer is to incorporate additional high-level prior information into the model.
5. The method for improving the accuracy of 3D scene reconstruction and new view synthesis according to claim 3, characterized in that ; The segmentation mask layer guides the model to focus on specific areas, reduces background interference, and enhances the realism and clarity of the generated results.