Three-dimensional scene reconstruction method without camera pose based on 3DGS
Through the 3DGS-based camera-free pose three-dimensional scene reconstruction method, the 3D Gaussian model and camera pose are optimized using depth maps and feature encoders, the problems of low three-dimensional reconstruction quality and inaccurate camera pose estimation in the prior art are solved, and higher quality three-dimensional reconstruction and more accurate camera pose estimation are achieved.
Patent Information
- Application Number
- CN202510216715.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
Existing three-dimensional reconstruction techniques are difficult to achieve high-quality three-dimensional reconstruction without camera postures, and camera posture estimates are inaccurate, especially when dealing with weak textures or repetitive areas.
Using a camera-free pose three-dimensional scene reconstruction method based on 3DGS, the image-free pose three-dimensional scene reconstruction method is used to obtain the depth map of the ordered picture sequence, construct the picture-based data set pre-trained feature encoder, initialize the 3D Gaussian model, and iteratively optimize the 3D Gaussian model and camera pose to achieve high-quality reconstruction of the three-dimensional scene.
It improves the quality of three-dimensional reconstruction and the accuracy of camera pose estimation, reduces artifacts and background collapse, and enhances the quality of pictures produced by new perspectives.
Smart Images

Figure CN120219613A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of three-dimensional reconstruction technology and view synthesis, and specifically relates to a method for three-dimensional scene reconstruction without camera pose based on 3DGS. Background Art
[0002] Implementing three-dimensional reconstruction from a two-dimensional image sequence is an important research task in the fields of computer graphics and computer vision, and has important application values in fields such as autonomous driving, cultural relic protection, and virtual game scene construction. Implementing three-dimensional reconstruction from a two-dimensional image sequence is to restore the geometric shape and texture information of the scene based on the two-dimensional image information of all viewpoints input. Three-dimensional scene reconstruction technology usually needs to obtain the camera pose through the Struct-from-Motion (SfM) algorithm. However, the SfM operation process is not only time-consuming, but may also fail when dealing with situations where the texture is weak or there are repetitive regions; in addition, the obtained camera pose is usually not accurate enough. Therefore, methods for three-dimensional scene reconstruction without camera pose have been continuously studied to address these challenges. Currently, the mainstream three-dimensional reconstruction technologies include three-dimensional reconstruction methods based on Neural Radiance Field (NeRF) and three-dimensional reconstruction methods based on 3DGS.
[0003] The neural radiance field models the geometric structure and color of the scene through an implicit representation method of a multi-layer perceptron, and combines volume rendering to achieve image rendering of a specified camera view, with good novel view synthesis ability; although the three-dimensional reconstruction method based on NeRF can render delicate and realistic images, the training time of such methods is long and the rendering speed is slow. Many studies have proposed to accelerate the training of NeRF from aspects such as introducing explicit geometric representation methods and combining effective acceleration techniques (improved algorithms such as KiloNeRF, InstantNGP, EfficentNeRF, etc.), effectively improving the training speed and rendering speed, but it still takes several minutes to render a scene. The three-dimensional reconstruction algorithms without camera pose based on NeRF can be roughly divided into two categories: SLAM-like methods and joint optimization methods; although SLAM-like methods can also achieve good reconstruction, they need to provide RGBD data as input or use the tracking system of SLAM to obtain accurate camera poses. Obtaining RGBD data requires a higher cost compared to obtaining images, and the reconstruction speed of such methods is also slow; the joint optimization method parameterizes the camera pose as an optimizable parameter, and jointly optimizes the camera pose and the neural radiance field during the training process. Although such methods can model the appearance of the scene, they usually assume situations such as providing inaccurate camera poses with small perturbations and the camera moving in a small range, and it is difficult to handle situations where the camera moves in a large range.
[0004] The 3DGS-based 3D reconstruction method uses explicit point clouds to represent the scene, and uses 3D Gaussian modeling to model the geometric shape of the scene, which can achieve high-quality real-time rendering of new perspectives. CF-3DGS (COLMAP-Free3DGaussian Splatting) based on 3DGS realizes 3D reconstruction without camera pose for ordered image sequences, and gradually adds new frames in a progressive training method. Before adding new frames, a local 3D Gaussian model is constructed and used to optimize the camera pose, so as to continuously train and expand the global 3D Gaussian to represent the scene; although this method has achieved good results in estimating camera pose and synthesizing new perspectives, the constructed local 3D Gaussian model lacks geometric structure information, and the depth map information is not fully utilized. At the same time, there are also some problems in synthesizing new perspectives, such as: the image has ghosts and the texture details are not accurate and detailed enough, and the camera pose optimization will be affected by lighting, noise, and weak texture. Summary of the invention
[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and to provide a camera-free 3D scene reconstruction method based on 3DGS, which can achieve high-quality 3D reconstruction of ordered frame sequences without relying on the Struct-From-Motion algorithm to obtain the camera pose, thereby improving the 3D reconstruction quality and the accuracy of estimating the camera pose.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A camera-free pose three-dimensional scene reconstruction method based on 3DGS comprises the following steps:
[0008] Obtain an ordered sequence of images taken by the camera, and input them into the monocular depth estimation network in sequence to predict the depth map of each frame of the image, and obtain a depth map sequence of the ordered image sequence;
[0009] Construct a picture pair dataset based on the ordered picture sequence, and pre-train a feature encoder on the picture pair dataset;
[0010] Initialize the camera pose of the first frame in the ordered image sequence and an empty training image set, restore the point cloud information according to the camera pose, depth map and color map of the first frame, initialize the 3D Gaussian model based on the point cloud information, and add the first frame in the ordered image sequence to the training image set;
[0011] Randomly select the color map and depth map of the image from the training image set as supervision information to iteratively optimize the 3D Gaussian model;
[0012] Generate a rendered color map according to the optimized 3D Gaussian model, input the rendered color map into a pre-trained feature encoder to obtain a feature map, and iteratively optimize the camera pose of the next frame of the ordered image sequence according to the rendered color map and the feature map, and add it to the training image set;
[0013] According to the order of the images in the ordered image sequence, adopt a progressive training method to alternately iteratively optimize the 3D Gaussian model and the camera pose of the new frame until all the images in the ordered image sequence are added to the training image set, obtaining the camera poses of all frames and the target 3D Gaussian model, and realizing three-dimensional scene reconstruction.
[0014] As a preferred technical solution, the monocular depth estimation network uses a pre-trained DPT model or ZoeDepth model.
[0015] As a preferred technical solution, the pre-training of a feature encoder is specifically:
[0016] Construct an image pair dataset: construct multiple image pairs for each camera view; each image pair contains three images, namely the ground truth image corresponding to the camera view, the noisy image corresponding to the camera view, and the ground truth image randomly selected from an adjacent view;
[0017] Use the VGG16 network as the feature encoder;
[0018] Input the three color maps of the image pair into the feature encoder to obtain three feature maps;
[0019] Calculate the Triplet loss function based on the three feature maps, and iteratively train the feature encoder until the maximum number of iterations or network convergence to obtain the final feature encoder;
[0020] The calculation formula of the Triplet loss function is:
[0021]
[0022] Among them, respectively represent the feature maps extracted by the ground truth image, the noisy image, and the ground truth image randomly selected from an adjacent view corresponding to the camera view P after being extracted by the feature encoder; margin is a pre-set difference threshold, set to 1; d() is the Euclidean distance, which is used to measure the distance between features.
[0023] As a preferred technical solution, the noise addition strategy in the noisy image adopts a region adaptive strategy for three times of noise processing;
[0024] The noise addition strategy is specifically:
[0025] Add two-dimensional Gaussian noise to the ground truth image corresponding to the camera view, and perform Gaussian blur according to an anisotropic Gaussian kernel to obtain the first initial noise map; then generate a first mask map based on a two-dimensional directional anisotropic Gaussian function, and then perform weighted summation of the first initial noise map and the ground truth image corresponding to the camera view according to the first mask map to obtain the first noise map, where the numerical value of the mask map serves as the corresponding weight coefficient;
[0026] Randomly select pixels on the first noise map to offset their pixel positions to obtain a second initial noise map, and generate a second mask map; perform weighted summation of the second initial noise map and the ground truth image corresponding to the camera view according to the second mask map to obtain a second noise map;
[0027] Perform Gaussian blur on the second noise map with an isotropic Gaussian kernel to obtain a third initial noise map, and generate a third mask map. Perform weighted summation of the third initial noise map and the ground truth image corresponding to the camera view according to the third mask map to obtain the final noise image.
[0028] As a preferred technical solution, the initialization of the 3D Gaussian model based on point cloud information is specifically as follows:
[0029] Initialize the external camera parameters of the first frame image in the ordered image sequence as the identity matrix, that is, initialize it as the unit camera pose, and add the first frame image to the training image set;
[0030] Obtain the depth map, camera internal parameters, and camera external parameters of the first frame image, and convert the pixel points of the depth map from the two-dimensional plane to a series of three-dimensional points through the inverse process of perspective projection;
[0031] Generate point cloud information based on the converted three-dimensional points and the color map of the first frame image;
[0032] Use the point cloud information to initialize the 3D Gaussian model.
[0033] As a preferred technical solution, the selection of the color map and depth map from the training image set as supervision information to iteratively optimize the 3D Gaussian model is specifically as follows:
[0034] Randomly select an image from the training image set as the target frame;
[0035] According to the camera pose of the target frame, use the 3D Gaussian model to generate the rendered color map and rendered depth map of the target frame;
[0036] Perform a linear transformation on the depth map of the target frame to obtain a corrected depth map, and calculate the color loss function and depth loss function to supervise the optimization of the 3D Gaussian model;
[0037] Iteratively execute the above steps until the maximum number of iterations is reached.
[0038] As a preferred technical solution, the method for generating the rendering color map is as follows:
[0039]
[0040] The method for generating the rendering depth map is as follows:
[0041]
[0042] where C is the rendering color map of the target frame, is the rendering depth map of the target frame, c i is calculated from the spherical harmonic coefficients and the viewing direction of the i-th 3D Gaussian in the 3D Gaussian model, α i is calculated according to the opacity ο and the two-dimensional covariance matrix of the i-th 3D Gaussian in the 3D Gaussian model, d i represents the depth value of the center of each 3D Gaussian;
[0043] The color loss function is calculated from the color map of the target frame and the rendering color map, and is expressed as:
[0044] L rgb =(1 - λ1)L1 + λ1L D-SSIM ,
[0045] The depth loss function is calculated based on the rendering depth map and the corrected depth map of the target frame, and is expressed as:
[0046]
[0047] D * = sD + b,
[0048] where L rgb is the color loss function, L1 is the L1 loss between the color map of the target frame and the rendering color map, L D-SSIM is the D-SSIM loss between the color map of the target frame and the rendering color map, λ1 is the weight coefficient, L depth is the depth loss function, is the rendering depth map of the target frame, D * is the corrected depth map of the target frame, D is the depth map of the target frame, s and b respectively represent the optimizable scaling coefficient and offset coefficient corresponding to the target frame;
[0049] The total loss function of the 3D Gaussian model optimization process is expressed as:
[0050] L = L rgb + λ2L depth ,
[0051] where λ2 is the weight coefficient of the depth loss function.
[0052] As a preferred technical solution, the camera pose of the next frame of the iteratively optimized ordered picture sequence is specifically as follows:
[0053] Obtain the next frame of the ordered picture sequence and its color map, and obtain the first feature map through a pre-trained feature encoder;
[0054] Perform a rigid transformation on the 3D Gaussian of the optimized 3D Gaussian model and transform it to the camera coordinate system of the previous frame of the next frame of the picture to obtain a transformed 3D Gaussian model;
[0055] Obtain the relative camera pose between the next frame of the picture and its previous frame of the picture. Use the transformed 3D Gaussian model to generate the rendered color map corresponding to the next frame of the picture, and input the rendered color map corresponding to the next frame of the picture into the pre-trained feature encoder to obtain the second feature map;
[0056] Calculate the color loss function and the feature map loss function, fix the parameters of the 3D Gaussian, and iteratively optimize the relative camera pose;
[0057] Calculate the camera pose of the next frame of the picture according to the optimized relative camera pose and the camera pose of the previous frame of the picture;
[0058] Add the next frame of the picture with the optimized camera pose to the training picture set.
[0059] As a preferred technical solution, the rigid transformation is expressed as:
[0060] G i =P i ⊙G world ,
[0061] where P i is the camera pose of the previous frame of the next frame of the picture, ⊙ is matrix multiplication, G world represents the 3D Gaussian model, and G i is the 3D Gaussian model in the camera coordinate system of the previous frame of the picture after rigid transformation;
[0062] The camera pose of the next frame of the picture is obtained by cumulative matrix multiplication of the relative camera poses between all adjacent pictures, and the calculation process is expressed as:
[0063] P i+1 =T i→i+1 ⊙…⊙T 2→3 ⊙T 1→2 ,
[0064] where P i+1 is the camera pose of the next frame of the picture, and T i→i+1 represents the relative camera pose between the next frame of the picture and its previous frame of the picture.
[0065] As a preferred technical solution, the feature map loss function is calculated from the first feature map and the second feature map, and is expressed as:
[0066]
[0067] where m j is the feature of the j-th pixel in the first feature map, is the feature of the j-th pixel in the second feature map, and ||||2 is the L2 norm.
[0068] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0069] On the one hand, in view of the problem that the prior art does not utilize depth information, the present invention introduces depth information guidance with multi-view consistency to provide rich depth information for optimizing the 3D Gaussian model, that is, applying the depth map to the optimization process of the 3D Gaussian model, which is beneficial to the modeling of geometric information and reduces the occurrence of artifacts and background collapse phenomena. On the other hand, the present invention uses the color map and the depth map as supervision information, directly applies the optimized 3D Gaussian model to optimize the relative camera pose, and then predicts the camera pose of the next frame of picture; since the optimized 3D Gaussian model has more accurate geometric structure information, it is beneficial to estimate a more accurate camera pose, and at the same time is beneficial to the 3D Gaussian model to better model the scene structure and model details, and better ensure the multi-view consistency of the scene content. On the other hand, the present invention introduces a new feature loss function when optimizing the relative camera pose, and the pre-trained feature encoder used in calculating the feature loss is robust to noise information and more sensitive to visual changes caused by camera movement. By reducing the feature loss, the influence of visual noise can be reduced, and the accuracy of camera pose estimation can be further improved. Therefore, the method proposed by the present invention can achieve higher-quality three-dimensional reconstruction, improve the quality of the synthesized pictures from new perspectives, and reduce the error of camera pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0071] Figure 1 FIG. is the overall flowchart of the three-dimensional scene reconstruction method without camera pose based on 3DGS in the embodiment of the present invention.
[0072] Figure 2This is the flow framework diagram of the 3D scene reconstruction method without camera pose based on 3DGS in the embodiments of the present invention.
[0073] Figure 3 This is the schematic flow diagram of the pre-trained feature encoder in the embodiments of the present invention.
[0074] Figure 4 This is the schematic flow diagram of iteratively optimizing the camera pose of the next frame in the ordered picture sequence in the embodiments of the present invention. Detailed implementation manners
[0075] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0076] In this application, referring to "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art understand explicitly and implicitly that the embodiments described in this application can be combined with other embodiments.
[0077] As Figure 1 、 2 shown, a 3D scene reconstruction method without camera pose based on 3DGS in this embodiment includes the following steps:
[0078] S1. Obtain an ordered picture sequence captured by a camera, and input it into a monocular depth estimation network in sequence to predict the depth map of each frame of the picture, and obtain a depth map sequence of the ordered picture sequence.
[0079] In this embodiment, the monocular depth estimation network adopts a pre-trained DPT model or ZoeDepth model. The DPT model and ZoeDepth model have good generalization capabilities. The DPT model can better predict the depth map of outdoor scene pictures, while the ZoeDepth model has excellent performance in depth prediction of indoor scenes. Therefore, using these two models can provide more accurate depth information.
[0080] S2. Construct a picture pair dataset according to the ordered picture sequence, and pre-train a feature encoder on the picture pair dataset.
[0081] Furthermore, in order to facilitate the calculation of the feature loss function during the camera pose optimization process and not affect the optimization speed of the 3D Gaussian, the present invention pre-trains a feature encoder, aiming to extract a feature map that is robust to noise, such as Figure 3 As shown, the pre-training process is as follows:
[0082] S2.1. First, construct a dataset of image pairs:
[0083] Construct multiple image pairs for each camera view; each image pair contains three images, namely the ground truth image corresponding to the camera view the noisy image corresponding to the camera view and the ground truth image randomly selected from an adjacent view It should be noted that since the present application studies an ordered image sequence, the images of adjacent views are also adjacent images, and either the previous frame or the next frame can be randomly selected.
[0084] Furthermore, the noise addition strategy in the noisy image uses a region adaptive strategy to perform three noise processes to increase the data diversity of the training set. Specifically:
[0085] First, add two-dimensional Gaussian noise to the ground truth image corresponding to the camera view and then perform Gaussian blur according to an anisotropic Gaussian kernel to obtain the first initial noise map I N1 ; then generate the first mask map M1 based on a two-dimensional directional anisotropic Gaussian function, and then perform weighted summation of the first initial noise map I N1 and the original image to obtain the first noise map. The numerical value of the mask map serves as the corresponding weight coefficient, and M1 is expressed as follows:
[0086] M1(i,j) = G(i - c i , j - c j ; σ i , σ j , A),
[0087] where (c i , c j ), (σ i , σ j ) represent the mean and standard deviation at the pixel position (i, j), respectively, and A represents the direction angle.
[0088] Then, randomly select pixels on the first noise map to offset their pixel positions to obtain the second initial noise map I N2 , and the offset pixel range is between [-2, 2]. Then generate the second mask map M2, and based on the second mask map, perform weighted summation of the second initial noise map I N2 and the original image Perform weighted summation to obtain a second noise map;
[0089] Finally, perform Gaussian blur on the second noise map using an isotropic Gaussian kernel to obtain a third initial noise map I N3 , generate a third mask map M3, and then re-combine the third initial noise map I N3 with the original image Perform weighted summation to obtain the final noise picture.
[0090] S2.2. Then, use the VGG16 network as the feature encoder; the VGG network can handle localization-related tasks well.
[0091] S2.3. Subsequently, input the three color maps of the picture pair into the feature encoder to obtain three feature maps;
[0092] S2.4. Finally, calculate the Triplet loss function based on the three feature maps, and continuously iteratively train the feature encoder by reducing the Triplet loss until the maximum number of iterations or the network converges to obtain the final feature encoder; the final feature encoder can extract similar feature maps for pictures from the same perspective and feature maps with large differences for pictures from different perspectives.
[0093] In this application, training is performed on the constructed dataset using the Triplet loss, making the feature maps extracted from pictures of different perspectives more distinct, while the feature maps extracted from pictures of the same perspective can overcome the interference of noise as much as possible and extract similar feature maps. The calculation formula of the Triplet loss function is:
[0094]
[0095] where respectively represent the ground truth picture, the noise picture corresponding to the camera perspective P, and the feature maps extracted from the ground truth picture randomly selected from an adjacent perspective by the feature encoder; margin is a pre-set difference threshold, set to 1 because it may become negative during the optimization process, which is because the training process will gradually reduce the and difference between them, while increasing the and difference between them. Therefore, when the difference between the two reaches -1, it can be understood that the loss is directly set to 0 without providing gradient optimization information, which is equivalent to achieving the optimization goal; d() is the Euclidean distance, used to measure the distance between features.
[0096] S3. Initialize the camera pose of the first frame of the ordered image sequence and an empty training image set. Restore the point cloud information based on the camera pose, depth map, and color map of the first frame, and initialize the 3D Gaussian model based on the point cloud information. Add the first frame of the ordered image sequence to the training image set.
[0097] Furthermore, the initial state of the training image set is empty. During the training process, after estimating the camera pose frame by frame in the order of the ordered image sequence, it is added to the training image set. After expanding the training image set, the 3D Gaussian model is trained, and the two steps of camera pose optimization and 3D Gaussian model optimization are repeated alternately. Therefore, initializing the training image set means adding the first frame of the image to the training image set, so as to first train the 3D Gaussian model according to the first frame of the image, then estimate the camera pose of the next frame of the image, and keep repeating until all images are added to the training image set and participate in the training of the 3D Gaussian model, so that the 3D Gaussian model can model the entire scene content. The 3D Gaussian model usually needs to be initialized based on the sparse point cloud, and without providing the camera pose, the sparse point cloud cannot be obtained; in order to initialize the 3D Gaussian model, the present invention first initializes the camera pose of the first frame of the image and the training image set, restores the point cloud information based on the camera pose, depth map, and color map of the first frame, and initializes the 3D Gaussian model based on the point cloud information, specifically including:
[0098] S3.1. Initialize the external camera parameters of the first frame of the ordered image sequence as the identity matrix, that is, initialize it as the unit camera pose, and add the first frame of the image to the training image set.
[0099] S3.2. Obtain the depth map D1, camera internal parameters K1, and camera external parameters P1 of the first frame of the image, and transform the pixel points of the depth map from the two-dimensional plane to a series of three-dimensional points through the inverse process of perspective projection.
[0100] More specifically, the inverse process of perspective projection is as follows:
[0101]
[0102] For all pixel points (x, y) in the image, project a ray from the center position of the camera through the pixel point. Since the depth value D represents the distance from the point to the camera center, the position of the three-dimensional point on the ray can be determined according to the depth value at the pixel position, and the spatial coordinates (x′, y′, z′) of the three-dimensional point can be calculated according to the principle of similar triangles, thus generating a series of three-dimensional points, where f x and f y represent the focal lengths in the x and y directions in the camera internal parameters.
[0103] S3.3. Generate point cloud information based on the converted three-dimensional points and the color map of the first frame of the image; the spatial coordinates of the points in the point cloud information are the coordinates of the corresponding three-dimensional points, and the colors of the points correspond to the colors of the corresponding pixels.
[0104] S3.4. Initialize the 3D Gaussian model using the point cloud information.
[0105] More specifically, the 3D Gaussian (3DGS) is represented by the mean μ, covariance matrix Σ, opacity ο, and spherical harmonic coefficients SH; the opacity ο affects the contribution of each Gaussian to the final color rendering; the spherical harmonic coefficients SH indicate that each Gaussian sphere has view-dependent colors, and when viewed from different perspectives, the colors of the Gaussian change with the perspective, defined as a 48-dimensional vector, which controls the color changes seen from different angles; c is calculated based on the spherical harmonic coefficients and the viewing direction; the specific method for initializing the 3D Gaussian model is as follows:
[0106] The mean μ is initialized to the coordinates of the point cloud, representing the center position of the Gaussian; and the covariance matrix is usually decomposed into a rotation matrix R and a scaling matrix S: ∑ = RSS T R T The rotation matrix and the scaling matrix are parameterized as a unit quaternion and a scaling vector respectively to represent. The covariance matrix represents the appearance shape of the Gaussian in space, and the scaling vector is initialized according to the distance between the point cloud and its surrounding neighboring points.
[0107] S4. Randomly select the color map and depth map of the image from the training image set as the supervision information to iteratively optimize the 3D Gaussian model.
[0108] Furthermore, there are some blurry or needle-like artifacts in the images rendered by the existing 3D reconstruction methods based on 3DGS. And many "floating objects" exist in the area closer to the camera, resulting in fog-like "spots" in some perspective-rendered images, and the phenomenon of blurring in the images; at the same time, there is also the phenomenon of background collapse, wrongly regarding the background as a foreground object, resulting in obvious errors in the depth map. In order to reduce the appearance of artifacts and background collapse, the present invention introduces depth information consistent with multiple perspectives to guide the optimization of the 3D Gaussian model, so as to learn more accurate scene geometry information, that is, select the color map and depth map from the training image set as the supervision information, align the scales of the depths of different perspectives, and iteratively optimize the 3D Gaussian model according to the depth consistent with multiple perspectives. Specifically:
[0109] S4.1. Randomly select an image from the training image set as the target frame; the adjacent frames of the new frame are more likely to be selected, which can ensure the reconstruction and rendering of the newly emerged scene content during the iterative optimization process. Since there is a newly added image in the progressive training process, when optimizing the 3D Gaussian model at the beginning, the adjacent frames near the newly added image can be preferentially selected, which can enable the 3D Gaussian model to better model the newly emerged scene content.
[0110] S4.2. According to the camera pose of the target frame, use the 3D Gaussian model to generate the rendered color map and the rendered depth map of the target frame.
[0111] More specifically, the generation method of the rendered color map is as follows:
[0112]
[0113] The generation method of the rendered depth map is as follows:
[0114]
[0115] Among them, C is the rendered color map of the target frame, is the rendered depth map of the target frame, c i is calculated from the spherical harmonic coefficients and viewing directions of the i-th 3D Gaussian in the 3D Gaussian model, α i is calculated according to the opacity ο and the two-dimensional covariance matrix of the i-th 3D Gaussian in the 3D Gaussian model, d i represents the depth value of the center of each 3D Gaussian.
[0116] S4.3. Perform a linear transformation on the depth map of the target frame to obtain a corrected depth map, and calculate the color loss function and the depth loss function to supervise the optimization of the 3D Gaussian model.
[0117] More specifically, the advantage of correcting the depth map of the target frame is that it can promote the multi-view consistency of the depth maps of multiple views, so as to further provide more accurate depth information and guide the optimization of the 3D Gaussian model. Since the depth map estimated by the monocular depth estimation network does not have multi-view consistency, while the depth map rendered based on the 3D Gaussian model has multi-view consistency, it is necessary to perform a linear transformation on the depth map of the target frame to obtain a corrected depth map, and optimize the linear transformation during the optimization process to make the depth maps of different views have view consistency.
[0118] Thus, the color loss function is calculated from the color map and the rendered color map of the target frame, and is expressed as:
[0119] L rgb =(1 - λ1)L1 + λ1L D-SSIM ,
[0120] The depth loss function is calculated based on the rendered depth map and the corrected depth map of the target frame, and is expressed as:
[0121]
[0122] D * = sD + b,
[0123] where L rgb is the color loss function, L1 is the L1 loss between the color map of the target frame and the rendered color map, L D-SSIM is the D-SSIM loss between the color map of the target frame and the rendered color map, λ1 is the weight coefficient, set to 0.2, L depth is the depth loss function, is the rendered depth map of the target frame, D * is the corrected depth map of the target frame, D is the depth map of the target frame, s and b respectively represent the optimizable scaling coefficient and offset coefficient corresponding to the target frame, for depth alignment.
[0124] S4.4. Iteratively execute the above steps until the maximum number of iterations is reached.
[0125] More specifically, the total loss function in the iterative optimization process of the 3D Gaussian model is expressed as:
[0126] L = L rgb + λ2L depth ,
[0127] where λ2 is the weight coefficient of the depth loss function, set to 0.05.
[0128] S5. Generate a rendered color map according to the optimized 3D Gaussian model, input the rendered color map into the pre-trained feature encoder to obtain a feature map, and iteratively optimize the camera pose of the next frame in the ordered picture sequence according to the rendered color map and the feature map, and add it to the training picture set.
[0129] Further, one way to optimize the camera pose is to directly parameterize the camera pose as a quaternion and a translation vector and optimize it according to the rendering result. However, this method is difficult to optimize when the camera moves significantly. The continuous sequence of picture frames is characterized by a small change in the camera movement between adjacent frames. Therefore, the present invention adopts the method of first estimating the relative camera pose between adjacent frames and then further obtaining the camera pose. In order to quickly estimate the accurate camera pose, the present invention is based on a 3D Gaussian model to complete the relative camera pose optimization. The relative camera pose is also parameterized as a quaternion and a translation vector for easy optimization. Usually, the pictures rendered by the 3D Gaussian model have noise information, and the lighting and other environments when taking two adjacent pictures may not be the same. These factors will affect the camera pose optimization. In order to reduce the influence brought by the noise factor, the present invention introduces a feature loss in the process of relative camera pose optimization. Specifically, as Figure 4 shown, the specific steps for iteratively optimizing the camera pose of the next picture are as follows:
[0130] S5.1. Obtain the next picture and its color map in the ordered picture sequence, and obtain the first feature map through the pre-trained feature encoder. The color map provides supervision information for the camera pose optimization of the next picture.
[0131] S5.2. Rigidly transform the 3D Gaussian of the optimized 3D Gaussian model to the camera coordinate system where the previous picture of the next picture is located, and obtain the transformed 3D Gaussian model (local 3DGS).
[0132] Specifically, the rigid transformation is expressed as:
[0133] G i =P i ⊙G world ,
[0134] where P i is the camera pose of the previous picture of the next picture, ⊙ is the matrix dot product, G world represents the 3D Gaussian model (global 3DGS), and G i is the 3D Gaussian model (local 3DGS) in the camera coordinate system of the previous picture after rigid transformation.
[0135] S5.3. Obtain the optimizable relative camera pose between the next picture and its previous picture, generate the rendered color map corresponding to the next picture using the transformed 3D Gaussian model, and input the rendered color map corresponding to the next picture into the pre-trained feature encoder to obtain the second feature map.
[0136] S5.4. Calculate the color loss function and the feature map loss function, fix the parameters of the 3D Gaussian, and iteratively optimize the relative camera pose.
[0137] Specifically, the color loss function is known above and will not be elaborated here; the feature map loss function is calculated from the first feature map and the second feature map, expressed as:
[0138]
[0139] where m j is the feature of the j-th pixel in the first feature map, is the feature of the j-th pixel in the second feature map, ||||2 is the L2 norm, used to calculate the modulus of the feature vector, thereby calculating the cosine similarity between feature vectors and calculating the feature loss.
[0140] S5.5. Calculate the camera pose of the next frame of image based on the optimized relative camera pose and the camera pose of the previous frame of image.
[0141] Specifically, the camera pose of the next frame of image is obtained by cumulative matrix multiplication of the relative camera poses between all adjacent images, and the calculation process is expressed as:
[0142] P i+1 = T i→i+1 ⊙…⊙T 2→3 ⊙T 1→2 ,
[0143] where P i+1 is the camera pose of the next frame of image, and T i→i+1 represents the relative camera pose between the next frame of image and its previous frame of image.
[0144] S5.6. Add the next frame of image with the optimized camera pose to the training image set.
[0145] S6. According to the order of the images in the ordered image sequence, adopt an incremental training method to alternately and iteratively optimize the 3D Gaussian model and the camera pose of the new frame of image until all frames in the ordered image sequence are added to the training image set, obtaining the camera poses of all frames and the target 3D Gaussian model, and realizing three-dimensional scene reconstruction.
[0146] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0147] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should all be considered as the scope described in this specification.
[0148] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A camera-free pose three-dimensional scene reconstruction method based on 3DGS, characterized in that: The steps include: Obtain an ordered sequence of images taken by the camera, and input them into the monocular depth estimation network in sequence to predict the depth map of each frame of the image, and obtain a depth map sequence of the ordered image sequence; Construct a picture pair dataset based on the ordered picture sequence, and pre-train a feature encoder on the picture pair dataset; Initialize the camera pose of the first frame in the ordered image sequence and an empty training image set, restore the point cloud information according to the camera pose, depth map and color map of the first frame, initialize the 3D Gaussian model based on the point cloud information, and add the first frame in the ordered image sequence to the training image set; Randomly select the color map and depth map of the image from the training image set as supervision information to iteratively optimize the 3D Gaussian model; Generate a rendering color map based on the optimized 3D Gaussian model, input the rendering color map into the pre-trained feature encoder to obtain a feature map, and iteratively optimize the camera pose of the next frame in the ordered image sequence based on the rendering color map and the feature map, and add it to the training image set; According to the order of pictures in the ordered picture sequence, a progressive training method is used to alternately iteratively optimize the 3D Gaussian model and optimize the camera pose of the new frame until all pictures in the ordered picture sequence are added to the training picture set, and the camera pose and target 3D Gaussian model of all frames are obtained to achieve three-dimensional scene reconstruction.
2. The camera-free pose 3D scene reconstruction method according to claim 1, characterized in that: The monocular depth estimation network adopts a pre-trained DPT model or a ZoeDepth model.
3. The camera-free pose 3D scene reconstruction method according to claim 1, characterized in that: The pre-training of a feature encoder is specifically as follows: Constructing a picture pair dataset: construct multiple picture pairs for each camera view; each picture pair contains three pictures, namely, the true value picture corresponding to the camera view, the noise picture corresponding to the camera view, and the true value picture randomly selected from the adjacent view; The VGG16 network is used as the feature encoder; Input the three color maps of the image pair into the feature encoder to obtain three feature maps; The Triplet loss function is calculated based on the three feature maps, and the feature encoder is iteratively trained until the maximum number of iterations or network convergence, to obtain the final feature encoder; The triplet loss function calculation formula is: in, They respectively represent the true value image corresponding to the camera perspective P, the noise image, and the feature map extracted by the feature encoder from the true value image randomly selected from the adjacent perspective; margin is the preset difference threshold, set to 1; d() is the Euclidean distance, which is used to measure the distance between features.
4. The camera-free pose 3D scene reconstruction method according to claim 3, characterized in that: The noise adding strategy in the noise picture adopts a regional adaptive strategy to perform three noise processing; The noise adding strategy is specifically as follows: Two-dimensional Gaussian noise is added to the true value image corresponding to the camera perspective, and Gaussian blur is performed according to an anisotropic Gaussian kernel to obtain a first initial noise image; then a first mask image is generated based on a two-dimensional directional anisotropic Gaussian function, and then the first initial noise image and the true value image corresponding to the camera perspective are weighted summed according to the first mask image to obtain a first noise image, wherein the numerical value of the mask image is used as the corresponding weight coefficient; A second initial noise map is obtained by randomly selecting pixels on the first noise map and shifting their pixel positions, and a second mask map is generated; a second initial noise map is weighted summed with a true value image of a corresponding camera perspective according to the second mask map to obtain a second noise map; A third initial noise map is obtained by Gaussian blurring with an isotropic Gaussian kernel on the second noise map, and a third mask map is generated. According to the third mask map, the third initial noise map and the true value image of the corresponding camera perspective are weighted summed to obtain a final noise image.
5. The camera-free pose 3D scene reconstruction method according to claim 1, characterized in that: The 3D Gaussian model is initialized based on point cloud information, specifically: Initialize the camera extrinsic parameters of the first frame in the ordered image sequence to the unit matrix, that is, initialize it to the unit camera pose, and add the first frame to the training image set; Get the depth map, camera intrinsic parameters and camera extrinsic parameters of the first frame, and convert the pixels of the depth map from a two-dimensional plane into a series of three-dimensional points through the inverse process of perspective projection; Generate point cloud information based on the converted 3D points and the color map of the first frame; Initialize the 3D Gaussian model using point cloud information.
6. The camera-free pose 3D scene reconstruction method according to claim 1, characterized in that: The method of selecting a color map and a depth map from a training image set as supervisory information to iteratively optimize the 3D Gaussian model is specifically as follows: Randomly select an image from the training image set as the target frame; According to the camera pose of the target frame, a rendering color map and a rendering depth map of the target frame are generated using a 3D Gaussian model; Perform linear transformation on the depth map of the target frame to obtain the corrected depth map, and calculate the color loss function and depth loss function to supervise the optimization of the 3D Gaussian model; The above steps are iterated until the maximum number of iterations is reached.
7. The camera-free pose 3D scene reconstruction method according to claim 6, characterized in that: The rendering color map is generated in the following way: The rendering depth map is generated in the following way: Among them, C is the rendering color map of the target frame, is the rendered depth map of the target frame, c i Calculated from the spherical harmonic coefficients and viewing direction of the i-th 3D Gaussian in the 3D Gaussian model, α i It is calculated based on the opacity ο of the i-th 3D Gaussian in the 3D Gaussian model and the two-dimensional covariance matrix, d i Represents the depth value of each 3D Gaussian center; The color loss function is calculated by the color map of the target frame and the rendering color map, and is expressed as: L rgb =(1-λ1)L1+λ1L D-SSIM , The depth loss function is calculated based on the rendered depth map and the corrected depth map of the target frame, and is expressed as: Among them, L rgb is the color loss function, L1 is the L1 loss between the color map of the target frame and the rendered color map, L D-SSIM is the D-SSIM loss between the target frame’s color map and the rendered color map, λ1 is the weight coefficient, and L depth is the depth loss function, is the rendered depth map of the target frame, D * is the corrected depth map of the target frame, D is the depth map of the target frame, s and b represent the optimizable scaling factor and offset factor corresponding to the target frame respectively; The total loss function of the 3D Gaussian model optimization process is expressed as: L=L rgb +λ2L depth , Among them, λ2 is the weight coefficient of the depth loss function.
8. The camera-free pose 3D scene reconstruction method according to claim 1, characterized in that: The iterative optimization of the camera pose of the next frame in the ordered image sequence is specifically as follows: Obtain the next frame of the image and its color map in the ordered image sequence, and obtain the first feature map through the pre-trained feature encoder; The 3D Gaussian of the optimized 3D Gaussian model is rigidly transformed to the camera coordinate system of the previous frame of the next frame, so as to obtain a transformed 3D Gaussian model; Obtain the relative camera pose between the next frame and the previous frame, generate a rendering color map corresponding to the next frame by transforming the 3D Gaussian model according to the relative camera pose, and input the rendering color map corresponding to the next frame into the pre-trained feature encoder to obtain a second feature map; Calculate the color loss function and feature map loss function, fix the parameters of the 3D Gaussian, and iteratively optimize the relative camera pose; Calculate the camera pose of the next frame based on the optimized relative camera pose and the camera pose of the previous frame; Add the next frame of the image with optimized camera pose to the training image collection.
9. The camera-free pose 3D scene reconstruction method according to claim 8, characterized in that: The stiffness change is expressed as: G i =P i ⊙G world , Among them, P i is the camera pose of the previous frame of the next frame, ⊙ is the matrix dot product, G world Represents a 3D Gaussian model, G i It is the 3D Gaussian model in the camera coordinate system of the previous frame after rigid transformation; The camera pose of the next frame is obtained by cumulative matrix multiplication of the relative camera poses between all adjacent pictures. The calculation process is expressed as: P i+1 =T i→i+1 ⊙…⊙T 2→3 ⊙T 1→2 , Among them, P i+1 is the camera pose of the next frame, T i→i+1 Represents the relative camera pose between the next frame and its previous frame.
10. The camera-free pose 3D scene reconstruction method according to claim 8, characterized in that: The feature map loss function is calculated from the first feature map and the second feature map, and is expressed as: Among them, m j is the feature of the jth pixel in the first feature map, is the feature of the j-th pixel in the second feature map, and ||||2 is the L2 norm.
Citation Information
Cited By
Large-scene three-dimensional reconstruction method based on three-dimensional Gaussian sputtering
CN120472121A
Object six-dimensional pose detection method and device and electronic equipment
CN121074126A