Lightweight 3D Reconstruction Method for Large Scenes Based on Hybrid Representations of NeRF and 3DGS

Through the simplified NeRF and 3DGS hybrid model, combining independent colors and observation colors, the depth map generated by 3DGS is used to guide NeRF dynamic sampling, which solves the problems of slow rendering speed and high storage occupancy in 3D reconstruction, and achieves efficient rendering and high-quality 3D reconstruction.

CN120182507BActive Publication Date: 2025-07-22SHENZHEN SENSING DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510649999.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-22
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction methods have problems such as slow rendering speed, high storage occupancy, and difficulty in dealing with complex lighting reflections and transparency.

Method used

Using a simplified NeRF and 3DGS hybrid model, by dividing colors into independent colors and observing colors, the depth map generated by 3DGS is used to guide NeRF dynamic sampling, and a mixed rendering image is generated in combination with α-mixing technology.

Benefits of technology

Significantly reduces storage occupancy, improves rendering quality and efficiency, and can handle complex lighting and reflection effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182507B_ABST
    Figure CN120182507B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight three-dimensional reconstruction method for large scenes based on a hybrid representation of NeRF and 3DGS, belonging to the technical field of three-dimensional reconstruction. The method includes the steps of: collecting multi-view images of the scene and objects in the scene and preprocessing them; constructing a hybrid model, including a simplified 3DGS model, a simplified NeRF model, and a hybrid unit; constructing the total loss of the hybrid model; and training the hybrid model to obtain a lightweight three-dimensional reconstruction model for reconstructing new-view images of the scene. In the present invention, the color is decomposed into an independent color expressed by Gaussian sputtering and an observed color expressed by NeRF. It not only retains the advantages of NeRF in efficient storage and complex lighting and reflection effects, but also utilizes the fast rendering ability of 3DGS, significantly reducing the storage occupancy while improving the rendering quality. The present invention also uses the depth map generated by the 3DGS model to guide the dynamic sampling of NeRF, which can effectively improve the training and rendering efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a lightweight three-dimensional reconstruction method for large scenes based on a hybrid representation of NeRF and 3DGS. Background Art

[0002] In computer vision, three-dimensional reconstruction refers to the process of reconstructing three-dimensional information (such as shape, color) based on single-view or multi-view images. Traditional three-dimensional reconstruction methods rely on discrete representation methods such as voxels and point clouds. These methods approximate the real scene by dividing the continuous three-dimensional space into discontinuous small units. However, these methods have inherent defects. For example, voxel models lose details due to the limitation of grid resolution, and the sparsity of point cloud models leads to inaccurate reconstruction in some areas.

[0003] In recent years, deep learning has made remarkable progress in the field of three-dimensional reconstruction. Neural Radiance Fields (NeRF) is a representative method that implicitly represents the light distribution and transmission characteristics in a scene by constructing a neural network. Its input is a continuous 5D vector, including any spatial position (x, y, z) in the scene and the viewing direction at that position (spherical coordinate representation method, including azimuth angle θ and elevation angle φ), and the output is the volume density at that position and the color (r, g, b) along the viewing direction (θ, φ) at that position. The general spatial position is represented by sampling points, which are sampled from the rays from the camera to a certain pixel on the imaging plane. Assuming there are 1024 pixels on the imaging plane, there are 1024 rays. NeRF can render high-quality real scenes, but its rendering speed is slow, and it requires a large number of neural network operations and iteration times.

[0004] 3D Gaussian Splatting (3DGS), as a representation method that combines explicit and implicit, balances the contradiction between rendering quality and training time. 3DGS uses a large number of three-dimensional Gaussian distributions to represent the scene. Each three-dimensional Gaussian distribution (also called a 3D Gaussian sphere) has parameters such as position, scaling scale, rotation quaternion, opacity, and spherical harmonic coefficients. However, this method is difficult to handle complex light reflection, transparency, and shadow effects, and requires 48 spherical harmonic coefficients to describe the scene, resulting in complex model optimization and high storage occupancy.

[0005] Glossary of Terms:

[0006] OpenCV is a cross-platform computer vision and machine learning software library distributed under the Apache 2.0 license (open source) and can run on Linux, Windows, Android, and Mac OS operating systems.

[0007] COLMAP (COLLISION-MAPpping) is a powerful open-source 3D reconstruction tool that supports automated multi-view geometry reconstruction, including core functions such as feature extraction, camera pose estimation, sparse point cloud reconstruction, and dense point cloud generation.

[0008] SSlM (Structural Similarity) is a commonly used metric in image quality assessment to measure the similarity between a reconstructed image and a reference image (usually the original image), and 1 - SSlM is used to represent the structural similarity loss.

[0009] L1 (Mean Absolute Error, MAE) loss calculates the sum of the absolute errors between the predicted values and the true values.

[0010] L2 (Mean Squared Error, MSE) loss calculates the sum of the squared errors between the predicted values and the true values. Summary of the Invention

[0011] The object of the present invention is to provide a lightweight 3D reconstruction method for large scenes based on a hybrid representation of NeRF and 3DGS, which not only retains the advantages of NeRF in efficient storage and complex lighting and reflection effects, but also can significantly reduce storage occupancy and improve rendering quality.

[0012] To achieve the above object, the technical solution adopted by the present invention is as follows: A lightweight 3D reconstruction method for large scenes based on a hybrid representation of NeRF and 3DGS, comprising the following steps;

[0013] S1, Collect multi-view images of the scene and the objects within the scene;

[0014] S2, After preprocessing the images, obtain N images, generate a sparse point cloud of the scene and the camera pose of each image, and store the images in the image set in the acquisition order, where the nth image is P n , 1 ≤ n ≤ N;

[0015] S3, Construct a hybrid model, including a simplified 3DGS model, a simplified NeRF model, and a hybrid unit;

[0016] The simplified 3DGS model is used to initialize each point of the sparse point cloud as a three-dimensional Gaussian distribution, and generate a first rendering image and a depth image of each image. The color of the three-dimensional Gaussian distribution is represented by 0th-order SH coefficients, and the first rendering image corresponding to image P n is , and the depth image is D n , and the color of the pixel points in the first rendering image is an independent color;

[0017] The simplified NeRF model dynamically samples to generate sampling points on the ray from the camera to the pixel point based on the depth value of the pixel point in the depth map, inputs the coordinates of the sampling points and the viewing direction from the camera to the sampling points, outputs the color of the sampling points, and then uses the alpha blending technique to render to obtain the second rendered image, image P n The corresponding second rendered image is , and the color of the pixel point in the second rendered image is the viewing color;

[0018] The mixing unit is used to and perform weighted summation on the independent color and the viewing color of the pixel points in to generate a mixed rendered image

[0019] S4. Construct the total loss Loss of the hybrid model;

[0020] ,

[0021] wherein, is the L1 loss, is to calculate the structural similarity, is the L2 loss, and λ1, λ2, and λ3 are the weights of the L1 loss, the SSIM loss, and the L2 loss respectively, and the sum of the three is 1;

[0022] S5. Use the image set to train the hybrid model to convergence by minimizing the total loss to obtain a lightweight three-dimensional reconstruction model;

[0023] S6. Use the lightweight three-dimensional reconstruction model to perform new view image reconstruction of the scene.

[0024] Preferably, S2 is specifically: use OpenCV to detect and remove blurred images, name the images in the format set according to the acquisition order, and then use COLMAP to read the images and generate the sparse point cloud of the scene and the camera poses corresponding to each image.

[0025] Preferably, in S3, the construction method of the simplified 3DGS model is;

[0026] Obtain a 3DGS model, where the attributes of one three-dimensional Gaussian distribution include position, opacity, diagonal scaling matrix, rotation matrix, and color; represent the color with 0th-order SH coefficients to obtain a simplified 3DGS model.

[0027] Preferably, the method for generating the first rendered image of the image is;

[0028] Project each three-dimensional Gaussian distribution onto a two-dimensional imaging plane to obtain a two-dimensional Gaussian distribution;

[0029] Construct a fast rasterizer to sort the two-dimensional Gaussian distributions;

[0030] Using the alpha blending technique, according to and the opacity σ i generate the color of each pixel point in the imaging plane to obtain the first rendered image.

[0031] Preferably, in S3, for the depth map D n the calculation method of the depth value D(u) of a pixel point u is as follows;

[0032] According to the camera pose of P n generate a ray from the camera center to the pixel point u. If this ray intersects Q three-dimensional Gaussian distributions, then D(u) is obtained according to the following formula;

[0033] ,

[0034] ,

[0035] In the formula, d q (u) is the distance from the q-th three-dimensional Gaussian distribution on the ray to the camera center, σ q is the opacity of the q-th three-dimensional Gaussian distribution, T q is the cumulative occlusion value of the first q - 1 three-dimensional Gaussian distributions for the q-th three-dimensional Gaussian distribution, 1 ≤ q ≤ Q, σ j is the opacity of the j-th three-dimensional Gaussian distribution, 1 ≤ j ≤ q - 1, T Q+1 is the cumulative occlusion value of the first Q three-dimensional Gaussian distributions for the pixel point u.

[0036] Preferably, the simplified NeRF model includes an input layer, three hidden layers, and an output layer;

[0037] The input layer is used to input the position encoding of the sampling points, the camera position, and the viewing direction;

[0038] The hidden layer is 128-dimensional;

[0039] The output layer is 3-dimensional and is used to output the color of the sampling points.

[0040] Preferably, in S3, the method of dynamic sampling of the simplified NeRF model is as follows:

[0041] If reconstructing P n , read the camera pose of P n generate a ray from the camera center to each pixel point on the imaging plane. For a pixel point u in the imaging plane, randomly sample N u sampling points within the range (D(u) - k, D(u) + k) from the camera center on its ray at the sampling distance k u ;

[0042] ,

[0043] ,

[0044] ,

[0045] In the formula, N set is the preset number of sampling points, k set is the preset sampling distance, D(u) is the depth value of the pixel point u in the depth map D n , D n (u) is the normalized value corresponding to D(u), D min , D max are respectively n the minimum and maximum values of the depth values in D

[0046] Preferably, when training the hybrid model, a preset number of iterations is set. The total loss is calculated once for each iteration, and the parameters of the hybrid model are adjusted using the Adam optimizer.

[0047] The idea of the present invention is as follows:

[0048] (1) Design two simplified models, namely the simplified 3DGS model and the simplified NeRF model. Among them, in the simplified 3DGS model, the color of the three-dimensional Gaussian distribution in the original 3DGS model is represented by 0th-order SH coefficients, and only 3 floating-point numbers are required to represent the color, while the traditional 3DGS model requires 48 spherical harmonic coefficients (SH coefficients), thus greatly reducing the number of three-dimensional Gaussian distributions and optimization parameters. At the same time, the simplified 3DGS model of the present invention provides a first rendering image and a depth map generated based on 0th-order SH coefficients. The simplified NeRF model reduces the expression of opacity in the existing NeRF model. The existing NeRF model needs to generate the color and opacity at the spatial position during the rendering process, and the present invention only needs it to generate the color at a certain position, which is used as a color component in the hybrid rendering image. Of course, the NeRF model in the prior art can also be used to provide colors for the present invention, but the simplified NeRF model is faster and has a lower computational load.

[0049] (2) Use the depth map generated by the simplified 3DGS model to guide the dynamic sampling of the simplified NeRF model to generate sampling points. According to the depth value of a certain pixel point in the depth map, it is used to guide the generation of sampling points within a certain sampling range on the ray according to the dynamic sampling distance and the dynamic number of sampling points, and the color of the sampling points is generated for generating the second rendering image.

[0050] (3) Then, use the color of the pixel points in the first rendering image as the independent color and the color in the second rendering image as the observed color, and perform weighted summation to generate the mixed color and obtain the hybrid rendering image.

[0051] Compared with the prior art, the advantages of the present invention are as follows:

[0052] The present invention proposes a hybrid model combining Neural Radiance Field (NeRF) and 3D Gaussian Splatting for high-quality rendering and efficient storage of 3D models. Specifically:

[0053] (1) Reduced storage occupancy: Through geometric-optical decoupled representation, the color is divided into independent color and observed color related to light and viewing angle such as reflection and shadow. The independent color is represented by a simplified 3DGS model, which only requires 0th-order spherical harmonic coefficients (3 floating-point numbers), while the traditional method requires 48 spherical harmonic coefficients; the observed color is represented by a simplified NeRF model. The combined way of representing color can significantly reduce the storage requirement.

[0054] (2) Accelerated dynamic sampling: The depth map rendered by the 3DGS model guides the sampling on the ray in NeRF, avoiding redundant calculations in the empty area, reducing the number of parameters to be optimized, and thus improving the training and rendering efficiency.

[0055] (3) Improved rendering quality: Hybrid rendering pipeline: NeRF can simulate complex light reflection, transparency and shadow effects, while 3DGS is responsible for efficient shape and basic color representation. The present invention adopts a dual-path asynchronous rendering method to perform weighted superposition on the outputs of 3DGS and NeRF, and the obtained hybrid rendering map significantly improves the rendering quality.

[0056] In summary, the present invention proposes a hybrid model combining NeRF and 3DGS. Through the way of division of labor and cooperation, it not only retains the advantages of NeRF in efficient storage and complex light and reflection effects, but also utilizes the fast rendering ability of 3DGS, significantly reducing the storage occupancy and at the same time improving the rendering quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the schematic diagram of the present invention;

[0058] Figure 2 is the structural diagram of the simplified NeRF model;

[0059] Figure 3 is an image in a scene;

[0060] Figure 4 is Figure 3 the depth map of. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The present invention will be further described below in conjunction with the embodiments and the drawings.

[0062] Example 1: Refer to Figures 1 to 2, A lightweight 3D reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS, including the following steps;

[0063] S1, Collect multi-view images of the scene and the objects in the scene;

[0064] S2, After preprocessing the images, obtain N images, generate the sparse point cloud of the scene and the camera poses of each image, and store the images in the image set in the acquisition order, where the nth image is P n , 1 ≤ n ≤ N;

[0065] S3, Construct a hybrid model, including a simplified 3DGS model, a simplified NeRF model, and a hybrid unit;

[0066] The simplified 3DGS model is used to initialize each point of the sparse point cloud into a three-dimensional Gaussian distribution, and generate the first rendering image and the depth image of each image. The color of the three-dimensional Gaussian distribution is represented by the 0th-order SH coefficient, and the first rendering image corresponding to the image P n is , and the depth image is D n , and the color of the pixel points in the first rendering image is the independent color;

[0067] The simplified NeRF model dynamically samples to generate sampling points on the ray from the camera to the pixel point based on the depth value of the pixel point in the depth image, inputs the sampling point coordinates and the viewing direction from the camera to the sampling point, outputs the color of the sampling point, and then uses the α-blending technique to render to obtain the second rendering image. The second rendering image corresponding to the image P n is , and the color of the pixel points in the second rendering image is the viewing color;

[0068] The hybrid unit is used to and perform weighted summation of the independent color and the viewing color of the pixel points in to generate a hybrid rendering image

[0069] S4, Construct the total loss Loss of the hybrid model;

[0070] ,

[0071] In the formula, is the L1 loss, is to calculate the structural similarity, is the L2 loss, and λ1, λ2, and λ3 are the weights of the L1 loss, the SSIM loss, and the L2 loss respectively, and the sum of the three is 1;

[0072] S5, Use the image set to train the hybrid model to convergence by minimizing the total loss to obtain a lightweight 3D reconstruction model;

[0073] S6. Reconstruct the new perspective image of the scene using the lightweight 3D reconstruction model.

[0074] In this embodiment, S2 is specifically as follows: Use OpenCV to detect and remove blurred images, name the images in the format set according to the acquisition order, and then use COLMAP to read the images and generate the sparse point cloud of the scene and the camera pose corresponding to each image.

[0075] The construction method of the simplified 3DGS model is as follows: Obtain a 3DGS model, where the attributes of one three-dimensional Gaussian distribution include position, opacity, diagonal scaling matrix, rotation matrix, and color; Represent the color with 0th-order SH coefficients to obtain the simplified 3DGS model.

[0076] The method for generating the first rendering of the image is as follows: Project each three-dimensional Gaussian distribution onto the two-dimensional imaging plane to obtain a two-dimensional Gaussian distribution; Construct a fast rasterizer to sort the two-dimensional Gaussian distributions; Use the α-blending technique to generate the color of each pixel point in the imaging plane according to and the opacity σ i to obtain the first rendering.

[0077] In S3, for the depth value D(u) of a pixel point u in the depth map D n the calculation method is as follows;

[0078] According to the camera pose of P n generate a ray from the camera center to the pixel point u. If this ray intersects Q three-dimensional Gaussian distributions in total, then D(u) is obtained according to the following formula;

[0079] ,

[0080] ,

[0081] In the formula, d q (u) is the distance from the qth three-dimensional Gaussian distribution on the ray to the camera center, σ q is the opacity of the qth three-dimensional Gaussian distribution, T q is the cumulative occlusion value of the first q - 1 three-dimensional Gaussian distributions on the qth three-dimensional Gaussian distribution, 1 ≤ q ≤ Q, σ j is the opacity of the jth three-dimensional Gaussian distribution, 1 ≤ j ≤ q - 1, T Q+1 is the cumulative occlusion value of the first Q three-dimensional Gaussian distributions on the pixel point u.

[0082] The simplified NeRF model includes an input layer, three hidden layers, and an output layer. Among them, the input layer is used to input the position encoding of the sampling points, the camera position, and the viewing direction; the hidden layer is 128-dimensional; the output layer is 3-dimensional and is used to output the color of the sampling points. The method of dynamic sampling of the simplified NeRF model is as follows:

[0083] If reconstructing P n , read the camera pose of P n , generate rays from the camera center to each pixel point on the imaging plane. For a pixel point u in the imaging plane, randomly sample N u sample points within the range (D(u) - k, D(u) + k) from the camera center on its ray at the sampling distance k u ;

[0084] ,

[0085] ,

[0086] ,

[0087] where N set is the preset number of sampling points, k set is the preset sampling distance, D(u) is the depth value of the pixel point u in the depth map D n , D n (u) is the normalized value corresponding to D(u), and D min , D max are the minimum and maximum values of the depth values in D n respectively.

[0088] When training the hybrid model, preset the number of iterations, calculate the total loss once per iteration, and use the Adam optimizer to adjust the parameters of the hybrid model.

[0089] Example 2: Refer to Figures 1 to 4 , based on Example 1, a more detailed implementation method is given. A lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS includes the following steps;

[0090] S1, Collect multi-view images of the scene and the objects in the scene, which can be obtained by video shooting with a camera;

[0091] S2, Preprocess to generate N images and a sparse point cloud;

[0092] For the video obtained in S1, use the ffmpeg open-source software to segment the captured video, and segment it at a rate of 10 frames per second to obtain the captured images; use the OpenCV open-source software to screen the obtained captured images to remove blurred images; rename the obtained images in the order of shooting time, starting from 0000, adding 1 to each image, and a total of N images are obtained and stored in the image set, and are sequentially marked as P1~P N ;

[0093] Use the COLMAP open-source software based on SFM (Motion Pose Estimation) to process the N preprocessed images, generate the sparse point cloud of the scene and the camera pose corresponding to each image, and the internal parameter matrix of the camera uses the default pinhole imaging model matrix.

[0094] S3. Construct a hybrid model, including a simplified 3DGS model, a simplified NeRF model, and a hybrid unit.

[0095] S4. Construct the total loss Loss of the hybrid model.

[0096] S5. Train the hybrid model to obtain a lightweight three-dimensional reconstruction model, specifically including S51~S56;

[0097] S51. Read the sparse point cloud, camera pose, and image set generated in step S2, and initialize them as the training set; preset the number of iterations to 30,000 times;

[0098] S52. Initialize each point of the sparse point cloud with the simplified 3DGS model as a three-dimensional Gaussian distribution:

[0099] First, define the attributes of the three-dimensional Gaussian distribution, including position, opacity, diagonal scaling matrix, rotation matrix, and color;

[0100] Then read the sparse point cloud, and for a point i among them, obtain its RGB color and three-dimensional coordinates;

[0101] Take the three-dimensional coordinates as the position μ of the three-dimensional Gaussian distribution corresponding to this point i ;

[0102] Map the opacity parameter self._opacity to the interval [0,1] using the sigmod function to obtain the opacity σ i ;

[0103] Adopt the K-nearest neighbor algorithm to find the average distance d from the 5 points closest to point i to point i i , take d i as the scaling parameters in the three-axis directions of the coordinate system to obtain the diagonal scaling matrix S;

[0104] Define the quaternion (0,0,0,1) and construct the rotation matrix R;

[0105] Convert the RGB color of point i to a 0th-order SH representation as the color of a three-dimensional Gaussian distribution.

[0106] Then, generate the three-dimensional Gaussian distribution of point i according to the following formula;

[0107] ,

[0108] ,

[0109] In the formula, p is the three-dimensional coordinate of a point in the scene, e is the natural constant, is the covariance matrix of the Gaussian distribution, and T is the transpose operation.

[0110] S53. Initialize the simplified NeRF model. Its input is the sampled point coordinates and the viewing direction from the camera to the sampled point. It is necessary to map the input to 33-dimensional data, namely 27-dimensional sampled point position encoding, 3-dimensional camera position, and 3-dimensional viewing direction. Taking a sampled point on any ray as an example, if its three-dimensional coordinate is (x, y, z), then the position encoding is mapped to 27-dimensional data such as (x, y, z, sin(2 i πx), sin(2 i πy), sin(2 i πz), cos(2 i πx), cos(2 i πy), cos(2 i πz)). In this formula, i = (1, 2, 3, 4), i represents the i-th power of 2. The hidden layer is a fully connected layer, and the sigmod function is used for all activation functions. After passing through the sigmod function in the output layer, the color value is multiplied by 255 and mapped to the range of 0 - 255 to meet the requirements of RGB images.

[0111] S54. Generate the first rendering image, depth map, and second rendering image corresponding to image P n ;

[0112] Generate the first rendering image: Read the camera pose corresponding to P n , convert the world coordinates to camera coordinates, calculate the projection matrix, project each three-dimensional Gaussian distribution onto the two-dimensional imaging plane to obtain a two-dimensional Gaussian distribution; construct a fast rasterizer to sort the two-dimensional Gaussian distribution; use the α-blending technique to generate the color of each pixel point in the imaging plane according to and the opacity σ i to obtain the first rendering image;

[0113] Generate the depth map: As described in step S3, according to P nThe camera pose is used to generate a ray from the camera center to the pixel point u. This ray intersects Q three-dimensional Gaussian distributions. Then, according to the formula of D(u) and T, the depth value D(u) of the pixel point u is calculated, and finally a depth map is generated. q Generate the second rendering: Refer to

[0114] and Figure 3 and Figure 4 In the scene and objects, they are divided into foreground and background, and there is a large amount of empty space between the foreground and the background. Taking Figure 3 and Figure 4 as an example, the objects closer to the camera have a brighter representation in the depth map, while the farther ones tend to be darker. Based on this, the present invention designs a dynamic sampling method. For a pixel point u in the imaging plane, within the range of (D(u)-k, D(u)+k) from the camera center on its ray, N u random sampling points are sampled at a sampling distance of k; then, the color of the pixel point is obtained by using the α blending technique for the sampling points on the ray as the observed color described in the present invention, and the graph composed of the observed colors is the second rendering. u

[0115] S55, generate a mixed rendering.

[0116] To avoid NeRF learning too much color information and strengthen the expression ability of 3dGS, the pixel colors in the first rendering and the second rendering are added with different weights respectively to obtain the mixed rendering. Specifically, in the first 15,000 iterations at the start of training, the weight of NeRF is set to 0, and from 15,000 to 25,000 iterations, it gradually increases to the set threshold.

[0117] S56, train the mixed model. Calculate the total loss once for each training, backpropagate the gradient, and use the Adam optimizer to update all the parameters of the mixed model until the number of iterations is reached to obtain a lightweight three-dimensional reconstruction model; for the trained model, only retain the training weights of the simplified NeRF model and the three-dimensional model file of the simplified 3DGS model for subsequent image reconstruction from new viewpoints.

[0118] S6, perform new viewpoint image reconstruction of the scene using the lightweight three-dimensional reconstruction model. This reconstruction is actually a process of image rendering:

[0119] According to the new viewpoint to be reconstructed, generate the camera position and the imaging plane in this viewpoint. Project the three-dimensional Gaussian distribution onto the imaging plane using the simplified 3DGS model, and obtain the first rendering through means such as sorting and α blending techniques. Generate a depth map based on the camera position, the imaging plane, and the three-dimensional Gaussian distribution.

[0120] ​According to the new perspective to be reconstructed, the viewing directions from the camera center to each pixel point in the imaging plane can be generated. The simplified NeRF model is used to dynamically sample according to the depth value of each pixel point in the depth map to generate the sampling points of each ray. For each ray, according to the sampling points and the viewing directions, the RGB values of the colors of the sampling points in the viewing direction are output, and the second rendering image is obtained by rendering with the alpha blending technique. The independent colors of the pixel points in the first rendering image and the viewing colors of the corresponding pixel points in the second rendering image are weighted and summed to generate the blended color of the pixel points, and finally the blended rendering image is obtained.

[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A lightweight three-dimensional reconstruction method for large scenes based on a hybrid representation of NeRF and 3DGS, characterized in that, Including the following steps; S1, Collect multi-view images of the scene and the objects within the scene; S2. After preprocessing the image, N images are obtained, and a sparse point cloud of the scene and the camera pose of each image are generated. The images are stored in the image set in the acquisition order, where the n-th image is P n , 1 ≤ n ≤ N; S3, Construct a hybrid model, including a simplified 3DGS model, a simplified NeRF model, and a hybrid unit; The simplified 3DGS model is used to initialize each point of the sparse point cloud into a three-dimensional Gaussian distribution, and generate the first rendering and depth map of each image. The color of the three-dimensional Gaussian distribution is represented by the 0th-order SH coefficient, and the image P n The corresponding first rendering is , and the depth map is D n , and the color of the pixel points in the first rendering is an independent color; The simplified NeRF model dynamically samples to generate sampling points on the ray from the camera to the pixel point based on the depth value of the pixel point in the depth map, inputs the coordinates of the sampling points and the viewing direction from the camera to the sampling points, outputs the color of the sampling points, and then uses the alpha blending technique to render to obtain a second rendered image, image P n The corresponding second rendered image is , and the color of the pixel point in the second rendered image is the observed color; The mixing unit is used to and perform weighted summation on the independent colors and observed colors of the pixel points in to generate a mixed rendering graph ; S4, Construct the total loss Loss of the hybrid model; , wherein, is the L1 loss, is for calculating the structural similarity, is the L2 loss, and λ1, λ2, and λ3 are the weights of the L1 loss, the SSIM loss, and the L2 loss respectively, and the sum of the three is 1; S5, Use the image set to train the hybrid model to convergence by minimizing the total loss, obtaining a lightweight three-dimensional reconstruction model; S6, Use the lightweight three-dimensional reconstruction model to reconstruct new-view images of the scene; In S3, the construction method of the simplified 3DGS model is; Obtain a 3DGS model, where the attributes of a three-dimensional Gaussian distribution include position, opacity, diagonal scaling matrix, rotation matrix, and color; Represent the color with 0th-order SH coefficients to obtain the simplified 3DGS model; The simplified NeRF model includes an input layer, 3 hidden layers, and an output layer. The input layer is used to input the position encoding of the sampling points, the camera position, and the viewing direction. The hidden layer is 128-dimensional, and the output layer is 3-dimensional, used to output the color of the sampling points.

2. The lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS according to claim 1, wherein, S2 is specifically: Use OpenCV to detect and remove blurred images, name the images in the format set according to the acquisition order, and then use COLMAP to read the images and generate the sparse point cloud of the scene and the camera poses corresponding to each image.

3. The lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS according to claim 1, characterized in that, The method for generating the first rendering of the image is; Project each three-dimensional Gaussian distribution onto the two-dimensional imaging plane to obtain a two-dimensional Gaussian distribution; Construct a fast rasterizer to sort the two-dimensional Gaussian distributions; Using the alpha blending technique according to and the opacity σ i generate the color of each pixel in the imaging plane to obtain the first rendering.

4. The lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS according to claim 1, wherein In S3, depth map D n The calculation method of the depth value D(u) of a pixel point u in it is as follows; According to P n Generate a ray from the camera center to the pixel point u. The ray intersects Q three-dimensional Gaussian distributions in total. Then D(u) is obtained according to the following formula; , , where d q (u) is the distance from the q-th three-dimensional Gaussian distribution on the ray to the camera center, σ q is the opacity of the q-th three-dimensional Gaussian distribution, T q is the cumulative occlusion value of the first q - 1 three-dimensional Gaussian distributions on the q-th three-dimensional Gaussian distribution, 1 ≤ q ≤ Q, σ j is the opacity of the j-th three-dimensional Gaussian distribution, 1 ≤ j ≤ q - 1, T Q+1 is the cumulative occlusion value of the first Q three-dimensional Gaussian distributions on the pixel point u.

5. The lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS according to claim 1, wherein In S3, the method for dynamic sampling of the simplified NeRF model is: If reconstructing P n , read the camera pose of P n . Generate rays from the camera center to each pixel point on the imaging plane. For a pixel point u in the imaging plane, randomly sample N u sampling points at a sampling distance of k within the range (D(u) - k, D(u) + k) from the camera center along its ray; u ​ , , , Where N set is the preset number of sampling points, k set is the preset sampling distance, D(u) is the depth value of the pixel point u in the depth map D n , D n (u) is the normalized value corresponding to D(u), D min , D max are respectively the minimum and maximum values of the depth values in D n .

6. The lightweight three-dimensional reconstruction method for large scenes based on the hybrid representation of NeRF and 3DGS according to claim 1, wherein, When training the hybrid model, preset the number of iterations, calculate the total loss once per iteration, and use the Adam optimizer to adjust the parameters of the hybrid model.

Citation Information

Patent Citations

  • Three-dimensional model reconstruction method based on 3D Gaussian Splitting

    CN119091051A

  • METHODS AND SYSTEMS FOR GENERATING THREE-DIMENSIONAL RENDERINGS OF A SCENE USING A MOBILE SENSOR ARRAY, SUCH AS NEURAL RADIANCE FIELD (NeRF) RENDERINGS

    US20250104323A1