Three-dimensional Gaussian city scene reconstruction method and system based on point cloud depth diffusion

Through the method based on point cloud depth diffusion, the initialization problem of sparse lidar point clouds in three-dimensional Gaussian scene reconstruction is solved, and the modeling quality and performance are improved, especially in image distortion scenarios, which achieves better reconstruction effects.

CN120070757AActive Publication Date: 2025-05-30ZHEJIANG UNIV

Patent Information

Application Number
CN202510148634.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art in three-dimensional Gaussian scene reconstruction based on the reliance of sparse lidar point clouds leads to a lack of sufficient surface profile and geometry during initialization, and performance degradation in image distortion scenarios.

Method used

Using a method based on point cloud depth diffusion, the sparse lidar point cloud is projected to the image sequence, and the dense depth map is obtained by weighted summing, and the depth point cloud is uniformly sampled through voxel selection method to serve as the initialization position of the three-dimensional Gaussian model. At the same time, a dense depth loss function and overall depth loss coefficient are introduced to train a three-dimensional Gaussian model to improve modeling quality.

Benefits of technology

The modeling capability of three-dimensional scene reconstruction model is improved, the performance in image distortion scenarios is enhanced, the modeling speed and quality is balanced, and it can be compatible with the existing technology, and it is used as a plug-in method to extend it.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070757A_ABST
    Figure CN120070757A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional Gaussian city scene reconstruction method and system based on point cloud depth diffusion. The method comprises the following steps: acquiring a data set with labels of a target scene for reconstructing a scene model; processing the laser radar point cloud data in the data set by using a depth diffusion method to obtain a dense point cloud and a dense depth map; using a voxel selection method to select a proper number of dense point clouds as three-dimensional Gaussian initialization positions; when the three-dimensional Gaussian model is trained, a dense depth map is used for depth supervision, and a dense depth loss function and a loss coefficient acting on the overall loss function are additionally added; and after training, a trained explicit city scene three-dimensional Gaussian model is obtained. The method fully considers the characteristics of the three-dimensional Gaussian city scene reconstruction task, adapts to the characteristics, proposes a more suitable initialization method and depth supervision mode, and can obtain a better scene modeling effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D reconstruction, and particularly to a 3D Gaussian scene reconstruction method based on point cloud depth diffusion. Background Art

[0002] Recent progress in end-to-end autonomous driving has highlighted the importance of closed-loop evaluation. However, existing simulators inevitably have domain differences from the real world, which emphasizes the need for real-world closed-loop simulators and promotes the development of high-quality urban scene modeling methods. Methods based on Neural Radiance Fields (NeRF) and methods based on 3D Gaussian Splashes (3DGS) are leading technologies in this field, capable of providing realistic rendering effects for new viewpoints. In contrast, methods based on 3DGS have boosted the rendering speed to real-time levels through explicit modeling and efficient differentiable splashes, and are currently widely used.

[0003] LiDAR point clouds, with their precise depth priors, are widely used in 3DGS-based autonomous driving methods, mainly serving two purposes: Gaussian initialization. 3DGS heavily relies on the initialized point cloud. Compared with the Structure from Motion (SfM) points, LiDAR points are denser, which helps reduce textureless areas and thus promotes the densification process; depth supervision. Depth is crucial for 3DGS scene modeling because it directly determines the 3D spatial position of the Gaussian. LiDAR depth provides effective constraints during training, reducing artifacts and improving rendering quality. However, LiDAR point clouds are still too sparse and lack sufficient surface contours and geometric structures during initialization. Meanwhile, statistical data shows that on average only 0.68% of the image pixels receive depth supervision. Especially in scenes with image distortion, such as exposure or low-light conditions, relying solely on color and sparse depth supervision will lead to a significant drop in performance. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a 3D Gaussian urban scene reconstruction method based on point cloud depth diffusion.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] In a first aspect, the present invention provides a 3D Gaussian urban scene reconstruction method based on point cloud depth diffusion, which includes the following steps:

[0007] S1. Obtain a training dataset of the target scene for reconstructing the scene model, where the training dataset has a spatio-temporally continuous image sequence of the target scene, corresponding camera parameters, as well as LiDAR point cloud and SfM point cloud data;

[0008] S2. Project the sparse lidar point cloud into the image space of each image in the image sequence. Traverse the image space, take each pixel therein as a target pixel one by one, select the depths of multiple nearest neighboring point clouds near each target pixel, calculate weights according to the distances between each point cloud and the target pixel, and thus sum the depths of the selected multiple point clouds weighted by the calculated weights to obtain the depth value of the target pixel. After traversing all pixels in the image space, obtain the dense depth map corresponding to the image, and project the dense depth map into the world coordinate system through the camera pose parameters to obtain the depth point cloud;

[0009] S3. Use the voxel selection method to uniformly sample the depth point cloud, and jointly form a dense point cloud with the sampled point cloud, the SfM point cloud and the lidar point cloud in the dataset and add it to the training dataset as the initialization position of the three-dimensional Gaussian to balance the modeling speed and the modeling quality;

[0010] S4. Use the training dataset to train the three-dimensional Gaussian model. During the training process, additionally add a dense depth loss function to the depth rendering result, and use the weighted sum of the image reconstruction loss, the SSIM loss and the dense depth loss multiplied by the loss coefficient as the final loss. Input the dense depth map and supervise and optimize the three-dimensional Gaussian model based on the final loss to improve the modeling quality of the three-dimensional Gaussian model for the scene;

[0011] S5. Visualize the trained three-dimensional Gaussian model to obtain the explicit modeling result.

[0012] Preferably, the training dataset in step S1 includes multiple spatially and temporally continuous image samples collected from the target scene, and each image sample contains the corresponding camera intrinsic matrix camera extrinsic parameters SfM point cloud P SfM and lidar point cloud P LiDAR , where is the rotation matrix, is the translation vector.

[0013] Preferably, the specific steps for obtaining the dense depth map and the depth point cloud for each image in the image sequence in step S2 are as follows:

[0014] S21. Project the sparse lidar point cloud P LiDAR containing L points from the world coordinate system to the image coordinate system to obtain the point set where represents the pixel coordinates in the image coordinate system obtained by projecting the i-th point, represents the depth corresponding to the i-th point after projection;

[0015] S22. For each pixel p = (u, v, d) on the current image in the image sequence, use the top-k algorithm to select k lidar depth points in the image coordinate system that are closest to the pixel p from the point set Meanwhile, it is necessary to ensure that for any selected lidar depth point and any unselected lidar depth point the following is satisfied: Satisfy:

[0016]

[0017] where ||·|| 2 represents the Euclidean distance,

[0018] represents the complement of Λ p,k For the complement set;

[0019] S23. For each lidar depth point at a different distance from the pixel p calculate the corresponding weight using the following formula

[0020]

[0021] where τ is a constant;

[0022] For all weights perform normalization to obtain each final weight which is used to perform weighted summation on the depths of the k lidar depth points as the diffused depth D(u, v) of each pixel p = (u, v, d) in the dense depth map :

[0023]

[0024] S24. Project the dense depth map back to the world coordinate system through the camera intrinsic matrix K and the camera extrinsic parameters [R, T] to obtain the depth point cloud P depth corresponding to the current image.

[0025] Preferably, in step S23, the constant τ = 1.

[0026] Preferably, when the voxel selection method is used to uniformly sample the depth point cloud in step S3, the three-dimensional space of the depth point cloud is divided into voxel grids according to a preset voxel size, and 1 depth point at the center of each voxel is used to replace all the depth point clouds inside the voxel, so as to obtain the sampled point cloud after uniform sampling.

[0027] Preferably, when using the voxel selection method to uniformly sample the depth point cloud, the density of point cloud sampling in the depth point cloud P is controlled by adjusting the voxel size. depth in the middle.

[0028] Preferably, the calculation steps of the final loss in step S4 are as follows:

[0029] S41. Project the three-dimensional Gaussian model in each round of training iteration into the camera coordinate system through the external camera parameters [R, T] to obtain the rendered depth of the current iteration round. where H represents the pixel height of the image and W represents the pixel width of the image;

[0030] S42. Calculate the dense depth loss of the image through the following formula:

[0031]

[0032] where (u, v) are the image pixel coordinates, and || || 1 represents the L1 norm;

[0033] S43. Record the dense depth loss of each image in the training dataset to obtain the dense depth loss vector l ∈ R M , where M represents the number of images in the training set; after normalizing the dense depth loss vector l, the loss coefficient ξ ∈ R M is obtained:

[0034] ξ = γ × Sgmoid(l × β)

[0035] where γ and β are two control constants, and Sgmoid(·) represents the Sigmoid normalization operation;

[0036] S44. Calculate the image reconstruction loss SSIM loss for the i-th image in the training dataset, and then together with the dense depth loss in step S42

[0037]

[0038] where λ 1and λ 2 are adjustable weight hyperparameters.

[0039] Preferably, in step S43, the weight values γ = 2 and β = 100.

[0040] Preferably, the weight value λ in step S44 1 = 0.2 and λ 2 = 0.01.

[0041] Preferably, in S5, rasterization technology is used to visualize the trained three-dimensional Gaussian model.

[0042] In a second aspect, the present invention provides a three-dimensional Gaussian urban scene reconstruction system based on point cloud depth diffusion, which includes the following modules:

[0043] A data acquisition module for acquiring a training data set of a target scene for reconstructing a scene model, where the training data set is provided with a spatio-temporally continuous image sequence of the target scene, corresponding camera parameters, and lidar point cloud and SfM point cloud data;

[0044] A depth point cloud generation module for projecting a sparse lidar point cloud into the image space of each image in the image sequence, traversing the image space to take each pixel therein as a target pixel, selecting the depths of multiple point clouds closest to each target pixel, calculating weights according to the distances between each point cloud and the target pixel, and thus weighted summing the depths of the selected multiple point clouds according to the calculated weights to obtain the depth value of the target pixel. After traversing all pixels in the image space, a dense depth map corresponding to the image is obtained, and the dense depth map is projected into the world coordinate system through camera pose parameters to obtain a depth point cloud;

[0045] An initial position generation module for uniformly sampling the depth point cloud using a voxel selection method, jointly constructing a dense point cloud with the sampled point cloud, the SfM point cloud, and the lidar point cloud in the data set, and adding it to the training data set as the initial position of the three-dimensional Gaussian to balance the modeling speed and modeling quality;

[0046] A model training module for training a three-dimensional Gaussian model using the training data set. During the training process, an additional dense depth loss function is added to the depth rendering result, and the weighted sum of the image reconstruction loss, SSIM loss, and dense depth loss multiplied by a loss coefficient is used as the final loss. The dense depth map is input and the three-dimensional Gaussian model is supervised and optimized based on the final loss to improve the modeling quality of the three-dimensional Gaussian model for the scene;

[0047] A visualization module for visualizing the trained three-dimensional Gaussian model to obtain an explicit modeling result.

[0048] In a third aspect, the present invention provides a computer electronic device, which includes a memory and a processor;

[0049] The memory is used for storing a computer program;

[0050] The processor is used for implementing the three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion as described in any one of the above first aspects when executing the computer program.

[0051] The present invention proposes a depth initialization and supervision scheme that is more suitable for three-dimensional scene reconstruction tasks, designs a densification method for sparse lidar point clouds based on the depth diffusion method, and can bring better modeling capabilities to the three-dimensional scene reconstruction model. Compared with traditional three-dimensional reconstruction schemes, the present invention has the following beneficial effects:

[0052] First, the present invention proposes a feasible method for three-dimensional urban scene reconstruction based on point cloud depth diffusion.

[0053] Secondly, the present invention fully considers the characteristics of the three-dimensional reconstruction task, and specifically designs the lidar point cloud depth diffusion for point cloud initialization and depth map supervision, and designs corresponding dense depth loss functions and overall depth loss coefficients.

[0054] Finally, the three-dimensional urban scene reconstruction method proposed by the present invention has wide generality. As a plug-in method, it can be simply extended to existing general methods and can obtain good results. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a schematic flow chart of the three-dimensional urban scene reconstruction method based on point cloud depth diffusion of the present invention.

[0056] Figure 2 is a schematic diagram of the module composition of the three-dimensional urban scene reconstruction system based on point cloud depth diffusion of the present invention.

[0057] Figure 3 is a schematic diagram of the composition of the computer electronic device in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0059] On the contrary, the present invention encompasses any alternatives, modifications, equivalent methods, and solutions defined by the claims that are within the essence and scope of the present invention. Further, in order to enable the public to better understand the present invention, in the following detailed description of the present invention, some specific details are described in detail. Those skilled in the art can fully understand the present invention without the description of these details.

[0060] A three-dimensional urban scene reconstruction method based on point cloud depth diffusion in the present invention is used to assist the initialization and depth supervision of a three-dimensional Gaussian model through the densification of sparse lidar point clouds. It should be particularly noted that the specific type of the three-dimensional Gaussian model in the present invention is not limited, and the targeted urban scene reconstruction is a preferred task of the present invention, and it is not actually limited to the reconstruction of urban scenes. The structure of the three-dimensional Gaussian model in the present invention is allowed to be different. Refer to Figure 1 , in a preferred embodiment of the present invention, the three-dimensional scene reconstruction method based on point cloud depth diffusion specifically includes the following steps:

[0061] S1. Obtain a training data set of the target scene for reconstructing the scene model, where the training data set has a spatio-temporally continuous image sequence of the target scene, corresponding camera parameters, and lidar point clouds and SfM point cloud data.

[0062] In the above S1 step of this embodiment, the training data set includes multiple spatio-temporally continuous image samples collected from the target scene, and each image sample should be attached with a corresponding camera intrinsic matrix camera extrinsic parameters SfM point cloud P SfM and lidar point cloud P LiDAR , where is the rotation matrix, is the translation vector. It should be noted that the multiple spatio-temporally continuous image samples in the present invention are essentially equivalent to the video continuously collected by the imaging device during the process of traveling in the target scene, and the image frames can be extracted from the video to form the above spatio-temporally continuous image sequence in the training data set.

[0063] In addition, in the embodiment of the present invention, in order to test the new view generation effect of the scene reconstruction of the present invention, based on the above training data set, a corresponding test data set can also be set, generally by extracting frames from the training data set to form the test data set. In the present invention, the ultimate algorithm goal is to render the modeled three-dimensional model into a color image from the perspective of the cameras in the actual training set or the test set, respectively showing its scene reconstruction ability and new view generation ability.

[0064] S2. Project the sparse lidar point cloud onto the image space of each image in the image sequence. Traverse the image space, take each pixel therein as a target pixel one by one, select the depths of multiple point clouds closest to each target pixel, calculate the weights according to the distance between each point cloud and the target pixel, and thus obtain the depth value of the target pixel by weighted summation of the selected multiple point cloud depths. After traversing all pixels in the image space, obtain the dense depth map corresponding to the image, and project the dense depth map into the world coordinate system through the camera pose parameters to obtain the depth point cloud.

[0065] In the above S2 step of this embodiment, the acquisition of the dense depth map and the depth point cloud mainly has the following sub-steps:

[0066] The specific steps to obtain the dense depth map and the depth point cloud for each image in the image sequence are as follows:

[0067] S21. Project the sparse lidar point cloud P containing L points LiDAR from the world coordinate system to the image coordinate system to obtain the point set where represents the pixel coordinates in the image coordinate system obtained by projecting the i-th point, and

[0068] represents the depth corresponding to the i-th point after projection; S22. For each pixel p = (u, v, d) on the current image in the image sequence, use the top-k algorithm to select k lidar depth points closest to the pixel p in the image coordinate system from the point set At the same time, it is necessary to ensure that for any selected lidar depth point and any unselected lidar depth point satisfy:

[0069]

[0070] where ||·|| 2 represents the Euclidean distance,

[0071] represents Λ p,k For the complement set;

[0072] S23. For each lidar depth point at a different distance from the pixel p calculate the corresponding weight using the following formula

[0073]

[0074] where is the required The corresponding weight; Represents the distance between two pixels; σ 2 Is the variance of the Gaussian function, which controls the Diffusion range; τ is a constant. In the embodiment of the present invention, τ = 1.

[0075] For all Weights Are normalized to obtain each Final weight Used to perform weighted summation on the depths of k lidar depth points Of the depth To obtain the diffused depth D(u, v) of each pixel p=(u, v, d) in the dense depth map As:

[0076]

[0077] S24. Project the dense depth map Back to the world coordinate system through the camera intrinsic matrix K and the camera extrinsic parameters [R, T] to obtain the depth point cloud P corresponding to the current image depth .

[0078] S3. Use the voxel selection method to uniformly sample the depth point cloud, and jointly form a dense point cloud with the sampled point cloud, the SfM point cloud, and the lidar point cloud in the dataset and add it to the training dataset as the initialization position of the three-dimensional Gaussian to balance the modeling speed and the modeling quality.

[0079] When using the voxel selection method to uniformly sample the depth point cloud in step S3 of the above embodiment, the specific method is to divide the three-dimensional space of the depth point cloud into voxel grids according to a preset voxel size, and replace all the depth point clouds inside the voxel with 1 depth point at the center of each voxel, so as to obtain the sampled point cloud after uniform sampling. Since only one point is retained in each voxel, the density of the point cloud sampling in the depth point cloud P depth Is controlled by adjusting the voxel size. In the embodiment of the present invention, the voxel size is 0.1m×0.1m×0.2m when dividing the voxel grid.

[0080] S4. Use the training dataset to train the three-dimensional Gaussian model. During the training process, an additional dense depth loss function and a loss coefficient acting on the overall loss are added to the depth rendering result. The weighted sum of the image reconstruction loss, the SSIM loss, and the dense depth loss is multiplied by the loss coefficient as the final loss. The dense depth map is input and the three-dimensional Gaussian model is supervised and optimized based on the final loss to improve the modeling quality of the three-dimensional Gaussian model for the scene.

[0081] It should be noted that the above three-dimensional Gaussian model belongs to the prior art and can be implemented by GaussianPro in the embodiments of the present invention. Its input is an image sequence with spatio-temporal continuity of the target scene, and a dense point cloud jointly composed of a sampled point cloud, an SfM point cloud, and a lidar point cloud (for initializing the three-dimensional Gaussian model). The above sampled point cloud, SfM point cloud, and lidar point cloud can be directly superimposed to form a dense point cloud.

[0082] In the above S4 step of this embodiment, the calculation steps of the dense depth loss function, the overall loss, and the final loss are as follows:

[0083] S41. Project the three-dimensional Gaussian model in each round of training iteration into the camera coordinate system through the camera external parameters [R, T] to obtain the rendered depth of the current iteration round where H represents the pixel height of the image and W represents the pixel width of the image;

[0084] S42. Calculate the dense depth loss of the image through the following formula:

[0085]

[0086] where (u, v) are the image pixel coordinates, and || || 1 represents the L1 norm;

[0087] S43. Record the dense depth loss of each image in the training dataset to obtain a dense depth loss vector l ∈ R M , where M represents the number of images in the training set; after normalizing the dense depth loss vector l, a loss coefficient ξ ∈ R M is obtained:

[0088] ξ = γ × Sgmoid(l × β)

[0089] where γ and β are two control constants, and Sgmoid(·) represents the Sigmoid normalization operation. In the embodiments of the present invention, the weight values γ = 2 and β = 100 are set.

[0090] S44. Calculate the image reconstruction loss SSIM loss for the i-th image in the training dataset, and then together with the dense depth loss in step S42

[0091]

[0092] where λ 1 and λ 2is an adjustable weight hyperparameter. In the embodiments of the present invention, the weight value λ is set 1 = 0.2 and λ 2 = 0.01.

[0093] It should be noted that the present invention can actually be combined with other 3D reconstruction methods as a plug-in method. If combined with other general methods, the corresponding loss function unique to the method can also be added to the functional form of the final loss and the corresponding weight λ new .

[0094] S5. Visualize the trained 3D Gaussian model to obtain an explicit modeling result.

[0095] It should be noted that visualizing the 3D Gaussian model belongs to the prior art and can be implemented by using techniques such as rasterization in the embodiments of the present invention.

[0096] It should be noted that the method steps of S1 to S5 above can essentially be implemented in the form of a computer program.

[0097] Therefore, based on the same inventive concept, the present invention also provides a 3D Gaussian urban scene reconstruction system corresponding to the 3D Gaussian urban scene reconstruction method based on point cloud depth diffusion provided in the above embodiments. As Figure 2 shown, the system includes:

[0098] A data acquisition module, configured to acquire a training data set of a target scene for reconstructing a scene model, where the training data set carries an image sequence of the target scene that is spatio-temporally continuous, corresponding camera parameters, and lidar point cloud and SfM point cloud data;

[0099] A depth point cloud generation module, configured to project a sparse lidar point cloud into the image space of each image in the image sequence, traverse the image space, take each pixel therein as a target pixel, select the depths of multiple point clouds closest to each target pixel, calculate weights according to the distances between each point cloud and the target pixel, and thus weighted sum the depths of the selected multiple point clouds according to the calculated weights to obtain the depth value of the target pixel. After traversing all pixels in the image space, a dense depth map corresponding to the image is obtained, and the dense depth map is projected into the world coordinate system through the camera pose parameters to obtain a depth point cloud;

[0100] An initialization position generation module, configured to uniformly sample the depth point cloud using a voxel selection method, and jointly form a dense point cloud with the sampled point cloud, the SfM point cloud, and the lidar point cloud in the data set and add it to the training data set as the initialization position of the 3D Gaussian to balance the modeling speed and modeling quality;

[0101] The model training module is used to train a three-dimensional Gaussian model using the training dataset. During the training process, a dense depth loss function is additionally added to the depth rendering result, and the weighted sum of the image reconstruction loss, SSIM loss, and dense depth loss is multiplied by a loss coefficient as the final loss. The dense depth map is input, and the three-dimensional Gaussian model is supervised and optimized based on the final loss to improve the modeling quality of the three-dimensional Gaussian model for the scene.

[0102] The visualization module is used to visualize the trained three-dimensional Gaussian model to obtain an explicit modeling result.

[0103] In the three-dimensional Gaussian urban scene reconstruction system based on point cloud depth diffusion in the above embodiment, the specific processes executed by each module can also refer to the specific steps of S1 to S5 described above, which will not be elaborated here.

[0104] Similarly, based on the same inventive concept, the present invention also provides a computer electronic device corresponding to the three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion provided in the above embodiment, as Figure 3 shown, which includes a memory and a processor;

[0105] The memory is used to store computer programs;

[0106] The processor is used to implement the three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion as described above when executing the computer program;

[0107] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0108] It can be understood that the above storage medium can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. At the same time, the storage medium can also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.

[0109] It can be understood that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0110] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In each of the embodiments provided in the present application, the division of steps or modules in the described system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0111] Next, a three-dimensional Gaussian city scene reconstruction method shown in S1 to S5 above will be applied to a specific example to demonstrate its specific technical effects.

[0112] Embodiment

[0113] The implementation method of this embodiment is as described in S1 to S5 above, and the specific steps will not be elaborated in detail. Only the effects will be shown for the case data below. The present invention is implemented on a dataset with lidar point cloud annotations:

[0114] Waymo Open dataset: The entire dataset contains approximately 12 million LiDAR annotation boxes and approximately 12 million image annotation boxes, generating approximately 113k LiDAR object trajectories and approximately 2.5 million image trajectories. The dataset is divided into 1000 training sets and 150 test sets.

[0115] In this example, 11 challenging scenarios are selected from the Waymo Open dataset for testing, and the verification metrics are chosen as FPS, PSNR, SSIM, and LPIPS. FPS (Frames Per Second) is a metric for measuring the smoothness of video, animation, or game graphics. In this embodiment, it is used as the rendering speed metric when a 3D model is rendered into a 2D image of the scene. The higher the FPS value, the faster the rendering. PSNR (Peak Signal-to-Noise Ratio) is an objective metric for measuring image quality. It measures image quality by comparing the mean squared error (MSE) between the original image and the distorted image. The higher the PSNR value, the smaller the distortion and the higher the image quality. SSIM (Structural Similarity Index Measure) is a metric for measuring the structural similarity between two images. SSIM takes into account the perception of structural information by the human visual system and can better reflect the details and texture information of the image. The value of SSIM ranges from 0 to 1, and the closer the value is to 1, the better the image quality. The advantage of SSIM is that it takes into account brightness, contrast, and structure, which is in line with human visual perception. The disadvantage is that the calculation is relatively complex and requires a large amount of computing power. At the same time, for some images with small distortions, the SSIM metric may not give accurate evaluation results. LPIPS (Learned Perceptual Image Patch Similarity) is a deep learning-based method for image quality assessment. LPIPS evaluates image quality by training a deep neural network to simulate the perception of images by the human visual system. The core idea of the LPIPS algorithm is based on perceptual learning, that is, by training a deep neural network to be able to simulate and learn the features of human visual perception and evaluate image quality based on these features. The advantage of LPIPS is that it can well simulate the perception of images by the human visual system and is sensitive to the distortion of details and textures. At the same time, this method has good generalization ability and can evaluate the image quality in different scenarios.

[0116] In this embodiment, the existing 3D Gaussian model GaussianPro is selected as the basic model, and the method of the present invention is applied on this basic model (denoted as the method of the present invention). The test results of the two in the scene reconstruction task are shown in the following table:

[0117] Table 1: Average test results of scene reconstruction in 11 test scenarios of the Waymo Open dataset

[0118]

[0119] The above experiments show that the method of the present invention can well improve the model reconstruction performance of existing general methods, while only bringing a slight decrease in the rendering speed, which is negligible in actual situations. It can well realize the 3D reconstruction application in urban scenes and has good application value.

[0120] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting the means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion, characterized in that: The following steps are involved: S1. Acquire a training data set of a target scene for reconstructing a scene model, wherein the training data set contains a spatiotemporally continuous image sequence of the target scene, corresponding camera parameters, and a lidar point cloud and SfM point cloud data; S2. Project the sparse lidar point cloud to the image space of each image in the image sequence, traverse the image space and take each pixel therein as the target pixel one by one, select the depth of multiple point clouds closest to each target pixel, calculate the weight according to the distance between each point cloud and the target pixel, and then sum the depths of the selected multiple point clouds according to the calculated weights to obtain the depth value of the target pixel. After traversing all pixels in the image space, obtain the dense depth map corresponding to the image, and project the dense depth map to the world coordinate system through the camera pose parameters to obtain the depth point cloud; S3, uniformly sampling the depth point cloud using a voxel selection method, and forming a dense point cloud with the sampled point cloud and the SfM point cloud and the lidar point cloud in the data set, and adding the dense point cloud to the training data set as the initialization position of the three-dimensional Gaussian to balance the modeling speed and the modeling quality; S4, using the training data set to train the three-dimensional Gaussian model, adding an additional dense depth loss function to the depth rendering result during the training process, taking the weighted sum of the image reconstruction loss, the SSIM loss and the dense depth loss multiplied by the loss coefficient as the final loss, inputting the dense depth map and performing supervised optimization on the three-dimensional Gaussian model based on the final loss, thereby improving the modeling quality of the three-dimensional Gaussian model for the scene; S5. Visualize the trained three-dimensional Gaussian model to obtain explicit modeling results.

2. The three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion according to claim 1 is characterized in that: The training data set in step S1 includes multiple image samples collected from the target scene in time and space, and each image sample contains the corresponding camera intrinsic parameter matrix Camera extrinsics SfM Point Cloud P SfM and the LiDAR point cloud P LiDAR ,in is the rotation matrix, is the translation vector.

3. The three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion according to claim 1 is characterized in that The specific steps of obtaining a dense depth map and a depth point cloud for each image in the image sequence in step S2 are as follows: S21, the sparse lidar point cloud P containing L points LiDAR Projecting from the world coordinate system to the image coordinate system to obtain a point set in represents the pixel coordinates in the image coordinate system obtained by projecting the i-th point, Indicates the depth corresponding to the projection of the i-th point; S22, for each pixel p = (u, v, d) on the current image in the image sequence, use the top-k algorithm to select the pixel from the point set Select the k lidar depth points closest to the pixel p in the image coordinate system At the same time, it is necessary to ensure that for any selected lidar depth point and any unselected lidar depth point satisfy: where ||·||2 represents the Euclidean distance, Represents Λ p,k for The complement of S23, for each laser radar depth point at different distances from pixel p Use the following formula to calculate the corresponding weight Where τ is a constant; For all Weight Normalize to get each The final weight For k lidar depth points Depth Perform weighted summation as a dense depth map The diffusion depth D(u,v) of each pixel p=(u,v,d) in is: S24, dense depth map The depth point cloud P corresponding to the current image is obtained by projecting the camera intrinsic parameter matrix K and the camera extrinsic parameter [R, T] back to the world coordinate system. depth .

4. The three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion according to claim 1 is characterized in that: When the voxel selection method is used to uniformly sample the depth point cloud in step S3, the three-dimensional space of the depth point cloud is divided into voxel grids according to a preset voxel size, and one depth point at the center of each voxel replaces all the depth point clouds inside the voxel, thereby obtaining a sampling point cloud after uniform sampling.

5. The method for reconstructing a three-dimensional Gaussian urban scene based on point cloud depth diffusion according to claim 4, characterized in that: When the depth point cloud is uniformly sampled using the voxel selection method, the depth point cloud P is controlled by adjusting the voxel size. depth The density of point cloud sampling.

6. The method for reconstructing a three-dimensional Gaussian urban scene based on point cloud depth diffusion according to claim 1, characterized in that: The calculation steps of the final loss in step S4 are as follows: S41. Project the 3D Gaussian model in each training iteration to the camera coordinate system through the camera external parameter [R, T] to obtain the rendering depth of the current iteration. Where H represents the pixel height of the image, and W represents the pixel width of the image; S42. Calculate the dense depth loss of the image by the following formula: Where (u, v) is the image pixel coordinate, ||||1 represents the L1 norm; S43. Record the dense depth loss of each image in the training data set and obtain the dense depth loss vector l∈R M , where M represents the number of images in the training set; the loss coefficient ξ∈R is obtained by normalizing the dense depth loss vector l M : ξ=γ×Sgmoid(l×β) Where γ and β are two control constants, Sgmoid(·) represents the Sigmoid normalization operation; S44. Calculate the image reconstruction loss for the i-th image in the training data set SSIM loss Then with the dense depth loss in step S42 Together they constitute the overall loss, and the i-th element ξ(i) in the loss coefficient ξ is applied to the overall loss to obtain the final loss of the i-th image in the training dataset: Where λ1 and λ2 are adjustable weight hyperparameters.

7. The method for reconstructing a three-dimensional Gaussian urban scene based on point cloud depth diffusion according to claim 6, characterized in that: The weight values ​​γ=2 and β=100 in step S43; the weight values ​​λ1=0.2 and λ2=0.01 in step S44.

8. The three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion according to claim 1 is characterized in that: In S5, the trained three-dimensional Gaussian model is visualized using rasterization technology.

9. A three-dimensional Gaussian urban scene reconstruction system based on point cloud depth diffusion, characterized in that: Includes the following modules: A data acquisition module, used to acquire a training data set of a target scene for reconstructing a scene model, wherein the training data set contains a spatiotemporally continuous image sequence of the target scene, corresponding camera parameters, and a lidar point cloud and SfM point cloud data; The depth point cloud generation module is used to project the sparse lidar point cloud to the image space of each image in the image sequence, traverse the image space and take each pixel as the target pixel one by one, select the depth of multiple point clouds closest to each target pixel, calculate the weight according to the distance between each point cloud and the target pixel, and then sum the depths of the selected multiple point clouds according to the calculated weights to obtain the depth value of the target pixel. After traversing all pixels in the image space, a dense depth map corresponding to the image is obtained, and the dense depth map is projected to the world coordinate system through the camera posture parameters to obtain a depth point cloud; An initialization position generation module is used to uniformly sample the depth point cloud using a voxel selection method, and the sampled point cloud and the SfM point cloud and the lidar point cloud in the data set together form a dense point cloud and add it to the training data set as the initialization position of the three-dimensional Gaussian to balance the modeling speed and modeling quality; A model training module is used to train a three-dimensional Gaussian model using the training data set. During the training process, a dense depth loss function is additionally added to the depth rendering result. The weighted sum of the image reconstruction loss, the SSIM loss, and the dense depth loss multiplied by the loss coefficient is used as the final loss. The dense depth map is input and the three-dimensional Gaussian model is supervised and optimized based on the final loss to improve the modeling quality of the three-dimensional Gaussian model for the scene; The visualization module is used to visualize the trained three-dimensional Gaussian model and obtain explicit modeling results.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the three-dimensional Gaussian urban scene reconstruction method based on point cloud depth diffusion as described in any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Deep learning-based image laser data fusion method for building reconstruction

    CN115423978A

  • Automatic driving scene simulation method and system based on neural point rendering

    CN117150755A

  • Indoor complex scene high-fidelity real-time rendering method based on three-dimensional Gaussian representation

    CN118096988A

  • Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion

    CN118351252A

Cited By

  • Diffusion model and Gaussian splashing-based three-dimensional scene generation method and related equipment

    CN121033252A