3D Gaussian weak texture compensation and density control reconstruction method

By combining laser point cloud data and image data, initializing the 3D Gaussian algorithm parameters and performing deep image fusion processing, the calculation cost and model quality problems of 3D Gaussian reconstruction technology in large scene reconstruction is solved, achieving more efficient rendering and more accurate model reconstruction.

CN120070752APending Publication Date: 2025-05-30BEIJING GREEN VALLEY TECH CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510133367.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing 3D Gaussian reconstruction technology has high requirements for GPU devices during large-scenario reconstruction, and consumes a lot of rendering resources. The model is prone to artifacts and floating objects. In environments with weak textures and poor lighting conditions, it lacks good initialization locations, resulting in poor model quality.

Method used

By collecting laser point cloud data, image data and pose data, using the Colmap framework for pose recovery and sparse point cloud generation, combining laser point cloud data for sparse processing and merging, initializing 3D Gaussian algorithm parameters, generating depth images and performing truncation depth processing, and fusing information from different data sources to improve model integrity and robustness.

Benefits of technology

It effectively reduces the computational cost and rendering pressure of 3D Gaussian reconstruction technology, reduces the appearance of artifacts and floating objects, improves the quality of the model in weak texture and low-light environments, and ensures the reasonable distribution of the Gaussian ellipsoid in space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070752A_ABST
    Figure CN120070752A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D Gaussian weak texture compensation and density control reconstruction method, and the method comprises the steps: recovering a camera pose P through employing a Colmap frame, merging the obtained first sparse point cloud data with laser point cloud data, solving a vacancy problem possibly existing in the sparse point cloud, and especially in an area with insufficient environment illumination or texture loss, carrying out the reconstruction of the 3D Gaussian weak texture compensation and density control, and carrying out the reconstruction of the 3D Gaussian weak texture compensation and density control. More accurate position information is provided for the Gaussian ellipsoid by using depth information of the laser point cloud, and floating objects are reduced; a depth image is generated based on the laser point cloud data and the camera pose P, an obtained depth value Z is fused with the laser point cloud data and the image data, and reasonable distribution of Gaussian ellipsoids in the space is ensured; excessive Gaussian distribution in a dense region is reduced through a dynamic threshold value and a voxel point number limiting strategy, so that the training cost and the rendering pressure are reduced; and particularly, by applying point number limitation in voxels and a dynamic threshold rejection strategy, the Gaussian density is reduced, so that ellipsoids at dense positions are reduced, and the training cost and the rendering pressure are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a method for weak texture compensation and density control reconstruction of 3D Gaussian. Background Art

[0002] With the continuous development of three-dimensional reconstruction technology, especially the emergence of the Neural Radiance Field (NeRF) technology, the generation of high-quality real-scene three-dimensional models has become possible. Through the learning ability of neural networks, this technology can reproduce the details in the scene with extremely high precision, including complex light and shadow effects and material characteristics. However, this high-precision reconstruction technology is often accompanied by huge computational costs, which limits its popularity in real-time and large-scale applications. To overcome these limitations, an emerging method called 3D Gaussian Splatting has received extensive attention in recent years. This method stands out with its efficient computational method and excellent performance. Compared with the traditional NeRF technology, this method not only significantly reduces the computational cost but also can show more realistic visual effects in the rendering stage. Therefore, the 3D Gaussian Splatting technology is becoming an important breakthrough in the field of three-dimensional reconstruction, providing new possibilities and more efficient solutions for applications in fields such as real-time scene modeling, virtual reality (VR), augmented reality (AR), and film production.

[0003] However, there are still many problems during its reconstruction.

[0004] 1. When representing large scenes, a large number of Gaussian ellipsoids are required, so the GPU device requirements are very high during training, and too much resources are occupied during rendering.

[0005] 2. The model is prone to artifacts and floating objects, affecting the overall quality of the model.

[0006] 3. In environments with weak textures and poor lighting conditions, the 3D Gaussian reconstruction algorithm lacks a good initial position, resulting in a poor quality of the final model. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for weak texture compensation and density control reconstruction of 3D Gaussian, which solves the above-mentioned technical problems pointed out in the prior art.

[0008] The present invention provides a method for weak texture compensation and density control reconstruction of 3D Gaussian, including the following operating steps:

[0009] Collect and obtain the data to be processed; perform pose recovery through the Colmap framework based on the data to be processed to obtain the camera pose P; obtain the first sparse point cloud data based on the camera pose P;

[0010] The data to be processed includes laser point cloud data, image data, and pose data;

[0011] After sparsifying the laser point cloud data, it is merged with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters; the first 3D Gaussian algorithm parameters are initialized to obtain the initialized 3D Gaussian algorithm parameters;

[0012] Based on the laser point cloud data and the camera pose P, a depth image is calculated; the depth image is subjected to truncated depth processing to obtain the depth value Z; based on the depth value Z, the laser point cloud data, and the image data, fusion processing is performed to obtain the image to be processed;

[0013] Using the training image data and the image to be processed D, 3D Gaussian optimization operations are performed to update the initialized 3D Gaussian algorithm parameters to obtain the updated 3D Gaussian algorithm parameters;

[0014] Record the backpropagation gradient during the process of the gradient descent algorithm inversely updating the initialized 3D Gaussian algorithm parameters; based on the backpropagation gradient, a Gaussian distribution is selected;

[0015] Preferably, the first 3D Gaussian algorithm parameters include rotation parameters, scale parameters, color parameters, and position opacity parameters in the 3D Gaussian model;

[0016] The initialized 3D Gaussian algorithm parameters include rotation parameters, scale parameters, color parameters, position opacity parameters, and voxel index T in the 3D Gaussian model.

[0017] Preferably, after collecting and obtaining the data to be processed; based on the data to be processed, pose recovery is performed through the Colmap framework to obtain the camera pose P; after obtaining the first sparse point cloud data based on the camera pose P, it further includes calculating an offset based on the camera pose P and the pose data; determining whether the offset is greater than or equal to a preset offset threshold, and if so, screening out the data to be processed corresponding to the camera pose P.

[0018] Preferably, the step of merging the sparsified laser point cloud data with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters includes the following operation steps:

[0019] Perform voxel downsampling on the laser point cloud data to obtain the downsampled laser point cloud data;

[0020] Perform voxel downsampling on the first sparse point cloud data to obtain the downsampled first sparse point cloud data;

[0021] Fuse the laser point cloud data after applying the lower ointment with the first sparse point cloud data after downsampling to obtain fused data;

[0022] Obtain the first 3D Gaussian algorithm parameters based on the fused data.

[0023] Preferably, the conversion calculation process based on the depth value Z and the laser point cloud data to obtain the image to be processed includes the following operation steps:

[0024] Obtain the camera internal parameters and camera external parameters;

[0025] Calculate the point cloud P in the camera coordinate system based on the laser point cloud data through the camera external parameters c ;

[0026] Project the point cloud P in the camera coordinate system c onto the image plane through the camera internal parameters to obtain a projected image;

[0027] Obtain the depth value Z' to be filled; fill the depth value Z' to be filled into the projected image to obtain the image to be processed D; the depth value Z' to be filled is the depth value when the depth value Z is less than or equal to the depth value threshold.

[0028] Preferably, the calculation method of the point cloud P in the camera coordinate system c is:

[0029] P c = RP world + t;

[0030] In the formula, P world = (x, y, z) is the laser point cloud data, P c = [X, Y, Z] is the point cloud in the camera coordinate system, R is the rotation matrix of the camera external parameters; t is the translation vector of the camera external parameters;

[0031] The calculation method of the projected image is:

[0032]

[0033] where K is the camera internal parameter, f x and f y are the focal lengths in the camera internal parameters, c x and c y are the principal point positions in the camera internal parameters;

[0034] The image to be processed D is expressed as:

[0035] D = +∞

[0036]

[0037] In the formula, H and W are the height and width of the image.

[0038] Preferably, the 3D Gaussian optimization operation is performed using the training image data and the image D to be processed, and the initialized 3D Gaussian algorithm parameters are updated to obtain the updated 3D Gaussian algorithm parameters, including the following operation steps:

[0039] Obtain the training image data; the training image data includes the training image I i and the training pose P corresponding to the training image i ;

[0040] Based on the training image I i and the training pose P i Calculate the rendering view W through the initialized 3D Gaussian algorithm parameters i and the depth image D' corresponding to the training image I i ; i ;

[0041] Based on the training image I i , the rendering view W i , the depth image D' i and the image D to be processed, calculate the loss function L;

[0042] Based on the loss function L, use the gradient descent algorithm to update the initialized 3D Gaussian algorithm parameters in reverse to obtain the updated 3D Gaussian algorithm parameters.

[0043] Preferably, the calculation method of the loss function L is:

[0044] L = (1 - λ)L 1 (I i , W i ) + λL SSIM (I i , W i ) + λ depth L 1 (D i , D' i );

[0045] Among them, λ is the L 1 and SSIM loss weight, and λ depth is the depth loss weight;

[0046] The updated 3D Gaussian algorithm parameters include the updated rotation parameter, the updated scale parameter, the updated color parameter, the updated position opacity parameter, and the updated voxel index T'.

[0047] Preferably, obtaining the Gaussian distribution by screening based on the backpropagation gradient includes the following operating steps:

[0048] Screen and obtain a plurality of Gaussian distributions to be processed whose backpropagation gradient is greater than or equal to a preset gradient threshold;

[0049] Count the number of Gaussian points in each of the Gaussian distributions to be processed; screen the Gaussian distributions whose number of Gaussian points is greater than or equal to a preset Gaussian point number threshold as the target Gaussian distributions to be processed;

[0050] Calculate the Gaussian distribution V' based on the scale S of each Gaussian in the target Gaussian distribution to be processed.

[0051] Preferably, the calculation method of the Gaussian distribution V′ is:

[0052] V′ = {i ∈ V||S i - μ| ≤ δσ};

[0053] Wherein, δ is a constant coefficient, and V′ is the Gaussian distribution retained in the voxel.

[0054] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:

[0055] Analyzing the above-mentioned method for weak texture compensation and density control reconstruction of 3D Gaussian provided by the present invention, in specific applications, first, data to be processed is collected through multiple sensor devices (including lidar, IMU sensor, and camera). The lidar provides accurate point cloud data, the IMU provides pose information, and the camera provides image data. These data provide a basis for subsequent pose recovery, sparse point cloud generation, and depth map calculation. Further, the Colmap framework is used to recover the camera pose P from these data, and key feature points are extracted from the camera perspective for triangulation to obtain the first sparse point cloud data, providing the necessary spatial information (pose) and the preliminary structure of the three-dimensional space (sparse point cloud data) for subsequent modeling. Further, the lidar point cloud data is sparsified and merged with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters, and the 3D Gaussian model parameters are initialized, including rotation, scale, color, position opacity, and voxel index, to mark the position of each Gaussian distribution. By initializing these parameters, a more reasonable starting point can be provided for the Gaussian model, especially in areas with weak texture and weak illumination, which can reduce the reconstruction error caused by improper initialization. Through the initialization operation, the Gaussian model has a more reasonable starting point, and by supplementing the lidar point cloud information, the problem of possible gaps in the sparse point cloud can be solved, especially in areas with insufficient ambient light or missing texture. Compared with traditional methods, this solution uses the depth information of the lidar point cloud to provide more accurate position information for the Gaussian ellipsoid, effectively reducing the appearance of floating objects. Further, a depth image is generated based on the lidar point cloud data and the camera pose P, and the depth image is subjected to truncated depth processing to obtain the accurate depth value Z. The depth value Z, lidar point cloud data, and image data are fused to obtain the image to be processed, combining the information of different data sources, thereby improving the integrity of the model. The fusion of the generated depth image and the lidar point cloud data effectively constrains the positioning of the Gaussian model, ensuring the reasonable distribution of the Gaussian ellipsoid in space. By fusing the information of different data sources, the problem of insufficient information or noise that may be brought by a single data source can be solved, improving the robustness of the modeling. Further, based on the training image data and the image to be processed D, 3D Gaussian optimization operations are performed by updating the initialized 3D Gaussian algorithm parameters (including rotation, scale, color, position opacity, voxel index, etc.) to make the model gradually improve during the training process. Through optimization, the 3D Gaussian model becomes more accurate.Guided by the training image data, the optimized Gaussian parameters can better match the structures and textures in the actual scene, reducing the reconstruction errors caused by poor initialization or excessive sparsification. Further, by recording the backpropagation gradients during the gradient descent process and screening the Gaussian distributions based on these gradients, unnecessary Gaussian distributions are eliminated. Through the dynamic threshold and the number of points within voxel limit strategy, the excessive Gaussian distributions in the dense regions are reduced, thereby reducing the training cost and rendering pressure. In particular, by applying the number of points within voxel limit and dynamic threshold elimination strategy, the density of the Gaussian is reduced, resulting in fewer ellipsoids in the dense areas and reducing the training cost and rendering pressure. Brief Description of the Drawings

[0056] Figure 1 It is a schematic diagram of the main process of a weak texture compensation and density control reconstruction method for 3D Gaussian;

[0057] Figure 2 It is a schematic diagram of the operation steps for obtaining the first 3D Gaussian algorithm parameters in a weak texture compensation and density control reconstruction method for 3D Gaussian;

[0058] Figure 3 It is a schematic diagram of the operation steps for obtaining the image to be processed in a weak texture compensation and density control reconstruction method for 3D Gaussian;

[0059] Figure 4 It is a schematic diagram of the operation steps for obtaining the updated 3D Gaussian algorithm parameters in a weak texture compensation and density control reconstruction method for 3D Gaussian;

[0060] Figure 5 It is a schematic diagram of the operation steps for obtaining the Gaussian distribution in a weak texture compensation and density control reconstruction method for 3D Gaussian. Detailed Description of the Invention

[0061] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.

[0063] Embodiment 1

[0064] As Figure 1 shown, Embodiment 1 of the present invention provides a weak texture compensation and density control reconstruction method for 3D Gaussian, including the following operation steps:

[0065] Step S10: Collect and obtain the data to be processed; perform pose recovery on the basis of the data to be processed through the Colmap framework to obtain the camera pose P; obtain the first sparse point cloud data on the basis of the camera pose P;

[0066] The data to be processed includes lidar point cloud data, image data, and pose data;

[0067] It should be noted that in the above embodiments of the present application, multiple data acquisition devices with lidar, IMU sensors (inertial measurement units), and cameras are used for acquisition to obtain the lidar point cloud data obtained by lidar, the image data captured by the camera, and the pose data recorded by the IMU sensors installed in the data acquisition devices;

[0068] Performing pose recovery on the basis of the data to be processed through the Colmap framework to obtain the camera pose P and obtaining the first sparse point cloud data are common knowledge to those skilled in the art, and the present application will not elaborate;

[0069] Step S20: After sparsifying the lidar point cloud data, merge it with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters; initialize the first 3D Gaussian algorithm parameters to obtain the initialized 3D Gaussian algorithm parameters;

[0070] The first 3D Gaussian algorithm parameters include rotation parameters, scale parameters, color parameters, and position opacity parameters in the 3D Gaussian model;

[0071] The initialized 3D Gaussian algorithm parameters include rotation parameters, scale parameters, color parameters, position opacity parameters, and voxel index T in the 3D Gaussian model;

[0072] It should be noted that in the above embodiments of the present application, the rotation, scale, color, and position opacity in the 3D Gaussian model are initialized, so that the 3D Gaussian obtains a better initial position, and the missing parts of the Colmap sparse points in weak texture and weak illumination areas are complemented. Moreover, through the initialization operation, the initialized 3D Gaussian algorithm parameters add the voxel index T where it is located compared with the first 3D Gaussian algorithm parameters to mark its spatial position.

[0073] Step S30: Calculate a depth image on the basis of the lidar point cloud data and the camera pose P; perform truncated depth processing on the depth image to obtain a depth value Z; perform fusion processing on the basis of the depth value Z, the lidar point cloud data, and the image data to obtain an image to be processed;

[0074] Step S40: Use the training image data and the image to be processed D to perform 3D Gaussian optimization operations to update the initialized 3D Gaussian algorithm parameters to obtain the updated 3D Gaussian algorithm parameters;

[0075] Step S50: Record the backpropagation gradient during the process of the gradient descent algorithm to update the parameters of the initialized 3D Gaussian algorithm in reverse; and obtain a Gaussian distribution based on the backpropagation gradient.

[0076] It should be noted that in the above embodiments of the present application, due to the limitations of hardware devices, the rendering efficiency is affected by the number of Gaussians. Therefore, during the training process, it is necessary to prevent excessive encryption of Gaussians.

[0077] In the above embodiments of the present application, data to be processed is first collected by multiple sensor devices (including lidar, IMU sensors, and cameras). The lidar provides accurate point cloud data, the IMU provides pose information, and the camera provides image data. These data provide a basis for subsequent pose recovery, sparse point cloud generation, and depth map calculation. Further, the Colmap framework is used to recover the camera pose P from these data, and key feature points are extracted from the camera perspective for triangulation to obtain the first sparse point cloud data, providing the necessary spatial information (pose) and the preliminary structure of the three-dimensional space (sparse point cloud data) for subsequent modeling. Further, the lidar point cloud data is sparsified and merged with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters, and the 3D Gaussian model parameters are initialized, including rotation, scale, color, position opacity, and voxel index, to mark the position of each Gaussian distribution. By initializing these parameters, a more reasonable starting point can be provided for the Gaussian model, especially in areas with weak texture and weak illumination, which can reduce the reconstruction error caused by improper initialization. By the initialization operation, the Gaussian model has a more reasonable starting point, and by supplementing the lidar point cloud information, the possible vacancy problem of the sparse point cloud can be solved, especially in areas with insufficient ambient light or missing texture. Compared with traditional methods, this solution uses the depth information of the lidar point cloud to provide more accurate position information for the Gaussian ellipsoid, effectively reducing the appearance of floating objects. Further, a depth image is generated based on the lidar point cloud data and the camera pose P, and the depth image is subjected to truncated depth processing to obtain an accurate depth value Z. The depth value Z, the lidar point cloud data, and the image data are fused to obtain an image to be processed, combining the information of different data sources, thereby improving the integrity of the model. The fusion of the generated depth image and the lidar point cloud data effectively constrains the positioning of the Gaussian model, ensuring the reasonable distribution of the Gaussian ellipsoids in space. By fusing the information of different data sources, the problem of insufficient information or noise that may be brought by a single data source is solved, and the robustness of the modeling is improved. Further, based on the training image data and the image D to be processed, 3D Gaussian optimization operations are performed by updating the initialized 3D Gaussian algorithm parameters (including rotation, scale, color, position opacity, voxel index, etc.) so that the model is gradually improved during the training process. Through optimization, the 3D Gaussian model becomes more accurate. Guided by the training image data, the optimized Gaussian parameters can better match the structure and texture in the actual scene, reducing the reconstruction error caused by poor initialization or excessive sparsification. Further, by recording the backpropagation gradients during the gradient descent process and screening the Gaussian distributions based on these gradients, unnecessary Gaussian distributions are removed. Through the dynamic threshold and the number of points within the voxel limit strategy, the excessive Gaussian distributions in the dense area are reduced, thereby reducing the training cost and rendering pressure. Especially by applying the number of points within the voxel limit and the dynamic threshold removal strategy, the density of the Gaussian is reduced, the ellipsoids in the dense area are reduced, and the training cost and rendering pressure are reduced.

[0078] Specifically, after step S10, it further includes: calculating an offset based on the camera pose P and the pose data; determining whether the offset is greater than or equal to a preset offset threshold, and if so (if not, then perform subsequent operation steps), screening out the data to be processed corresponding to the camera pose P.

[0079] It should be noted that in the above embodiments of the present application, when the offset is greater than or equal to a preset offset threshold (the offset threshold is usually set to 0.2 m), it is considered that the camera pose P is an untrustworthy result, and the data to be processed is an unreasonable acquisition result. Then, the camera position is removed, and at the same time, the data to be processed corresponding to the camera position is screened out to prevent noise and errors during the execution of subsequent steps.

[0080] Specifically, as Figure 2 shown, in step S20, after sparsifying the laser point cloud data and merging it with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters, it includes the following operation steps:

[0081] Step S21: Perform voxel downsampling on the laser point cloud data (the voxel downsampling size of the laser point cloud data is 0.05 m) to obtain the downsampled laser point cloud data;

[0082] Step S22: Perform voxel downsampling on the first sparse point cloud data (the voxel downsampling size of the first sparse point cloud data is 0.02 m) to obtain the downsampled first sparse point cloud data;

[0083] Step S23: Fuse the downsampled laser point cloud data with the downsampled first sparse point cloud data to obtain fused data;

[0084] Step S24: Obtain the first 3D Gaussian algorithm parameters based on the fused data.

[0085] It should be noted that in the above embodiments of the present application, the laser point cloud data is first subjected to voxel downsampling processing, and the points in the original laser point cloud data are downsampled according to the size of the voxel grid, so as to reduce the data volume, reduce the computational complexity, and facilitate subsequent processing and calculation; and further perform voxel downsampling processing on the first sparse point cloud data, so as to compress the sparse point cloud data and make the data more compact; further, the laser point cloud data after voxel downsampling processing is fused with the first sparse point cloud data. The laser point cloud data provides dense geometric information of the environment, while the sparse point cloud data provides more detailed information (such as texture and visual features) obtained through the camera. The fusion of the two complements the missing sparse points and enhances the integrity of the Gaussian model, especially in the parts with weak illumination and missing texture; further, through the fused data, the parameters of the 3D Gaussian algorithm are further obtained, so as to improve the accuracy of parameter estimation and make the generated 3D Gaussian model better fit the actual scene.

[0086] Specifically, as Figure 3 shown, in step S30, the conversion calculation process based on the depth value Z and the laser point cloud data to obtain the image to be processed includes the following operation steps:

[0087] Step S31: Obtain the camera internal parameters and the camera external parameters;

[0088] Step S32: Calculate the point cloud P in the camera coordinate system based on the laser point cloud data through the camera external parameters c ;

[0089] The point cloud P in the camera coordinate system c is calculated as follows:

[0090] P c = RP world + t;

[0091] In the formula, P world =(x, y, z) is the coordinate of the laser point cloud data, and P c =[X, Y, Z] is the point cloud in the camera coordinate system (that is, the coordinate of the point obtained by converting the laser point cloud data to the camera coordinate system through the camera external parameters), R is the rotation matrix of the camera external parameters; t is the translation vector of the camera external parameters;

[0092] Step S33: Project the point cloud P in the camera coordinate system c onto the image plane through the camera internal parameters to obtain a projection image;

[0093] The calculation method of the projection image is:

[0094]

[0095] where K is the camera internal parameter, f x and f y are the focal lengths in the camera internal parameters, and c x and c y are the principal point positions in the camera internal parameters;

[0096] Step S34: Obtain the depth value Z' to be filled; fill the depth value Z' to be filled into the projection image to obtain a to-be-processed image D;

[0097] The to-be-processed image D is expressed as:

[0098] D = +∞

[0099]

[0100] where H and W are the height and width of the image;

[0101] The depth value Z' to be filled is the depth value when the depth value Z is less than or equal to the depth value threshold (usually 200m);

[0102] It should be noted that in the above embodiments of the present application, due to the limited acquisition effective distance of the lidar, therefore, in the embodiments of the present application, the depth value Z' to be filled with a depth value less than or equal to the depth value threshold (200m) is selected to fill the projection image. When filling the projection image, the depth value Z' to be filled is filled into the projection image pixel points at the corresponding positions, and the depth values of the remaining projection image pixel points are recorded as +∞;

[0103] In the above embodiments of the present application, the camera internal parameters and external parameters are first obtained as the basis for subsequent processing; then, the lidar point cloud data is converted from the world coordinate system to the camera coordinate system by using the external parameters (rotation matrix R and translation vector t) of the camera, providing a spatial reference for subsequent image generation and depth filling; the point cloud in the camera coordinate system is projected through the internal parameters of the camera to obtain the projection of the point cloud on the image plane, generating a depth map, where the depth information of each pixel corresponds to the distance of the 3D point cloud; further, by determining the depth value Z' to be filled and filling it into the projection image, the blank or missing parts of the image caused by the limitation of the lidar point cloud data are avoided.

[0104] Specifically, as Figure 4 shown, in step S40, a 3D Gaussian optimization operation is performed using the training image data and the to-be-processed image D to update the initialized 3D Gaussian algorithm parameters, obtaining the updated 3D Gaussian algorithm parameters, including the following operation steps:

[0105] Step S41: Obtain the training image data; the training image data includes the training image I iand the training pose P corresponding to the training image i ;

[0106] Step S42: Based on the training image I i and the training pose P i Calculate the rendering view W through the initialized 3D Gaussian algorithm parameters i and the depth image D' corresponding to the training image I i ; i ;

[0107] Step S43: Based on the training image I i , the rendering view W i , the depth image D' i and the image to be processed D to calculate the loss function L;

[0108] The calculation method of the loss function L is as follows:

[0109] L = (1 - λ)L 1 (I i , W i ) + λL SSIM (I i , W i ) + λ depth L 1 (D i , D' i );

[0110] where λ is the L 1 and SSIM loss weight, and λ depth is the depth loss weight;

[0111] Step S44: Based on the loss function L, use the gradient descent algorithm to update the initialized 3D Gaussian algorithm parameters in the reverse direction to obtain the updated 3D Gaussian algorithm parameters;

[0112] The updated 3D Gaussian algorithm parameters include the updated rotation parameter, the updated scale parameter, the updated color parameter, the updated position opacity parameter, and the updated voxel index T';

[0113] It should be noted that in the specific implementation process of the above embodiments of the present application, when using the loss function L to update the initialized 3D Gaussian algorithm parameters in the reverse direction through the gradient descent algorithm, the position of the Gaussian will be updated, and accordingly, the voxel index of each Gaussian will be updated (that is, the updated voxel index T' is obtained).

[0114] Specifically, as Figure 5 shown, in step S50, screening the Gaussian distribution based on the backpropagation gradient includes the following operation steps:

[0115] Step S51: Screening and obtaining a plurality of Gaussian distributions to be processed whose back-propagation gradient is greater than or equal to a preset gradient threshold;

[0116] It should be noted that in the above-mentioned embodiment of the present application, during the training process, it is necessary to record the back propagation gradient of each Gaussian, and then by screening out Gaussian distributions with larger gradients (i.e., the back propagation gradient is greater than or equal to the gradient threshold), it can be ensured that during the optimization process, those Gaussian distributions that have a greater impact on the loss are focused on;

[0117] Step S52: Count the number of Gaussian points in each of the Gaussian distributions to be processed; select Gaussian distributions whose number of Gaussian points is greater than or equal to a preset Gaussian point number threshold (usually 20 points) as target Gaussian distributions to be processed;

[0118] Step S53: Calculate the Gaussian distribution V′ based on the scale S of each Gaussian in the target Gaussian distribution to be processed;

[0119] The Gaussian distribution V′ is calculated as follows:

[0120] V′={i∈V||S i -μ|≤δσ};

[0121] in, δ is a constant coefficient and V′ is the Gaussian distribution preserved within the voxel.

[0122] It should be noted that the above-mentioned embodiment of the present application applies the point number limit in the voxel and the dynamic threshold elimination strategy during the encryption process, which reduces the density of Gaussians, reduces the ellipsoids in dense areas, and reduces the training cost and rendering pressure. The above control process reduces the number of Gaussians, especially in dense areas. This not only reduces the training cost, but also reduces the computational burden during rendering. Using the depth information provided by the laser point cloud, the position of the Gaussian ellipsoid is accurately constrained, which reduces the floating objects or inaccurate positions in the model and improves the accuracy and authenticity of the overall model.

[0123] In summary, a method for weak texture compensation and density control reconstruction of 3D Gaussian proposed in the embodiments of the present invention first collects data to be processed through multiple sensor devices (including lidar, IMU sensors, and cameras). The lidar provides accurate point cloud data, the IMU provides pose information, and the camera provides image data, providing a basis for subsequent pose recovery, sparse point cloud generation, and depth map calculation. Further, the Colmap framework is used to recover the camera pose P from this data, and key feature points are extracted from the camera perspective for triangulation to obtain the first sparse point cloud data, providing the necessary spatial information (pose) and the preliminary structure of the three-dimensional space (sparse point cloud data) for subsequent modeling. Further, the lidar point cloud data is sparsified and merged with the first sparse point cloud data to obtain the first 3D Gaussian algorithm parameters, and the 3D Gaussian model parameters are initialized, including rotation, scale, color, position opacity, and voxel index, to mark the position of each Gaussian distribution. By initializing these parameters, a more reasonable starting point can be provided for the Gaussian model, especially in areas with weak texture and weak illumination, reducing the reconstruction error caused by improper initialization. Through the initialization operation, the Gaussian model has a more reasonable starting point, and by supplementing the lidar point cloud information, the possible vacancy problem of the sparse point cloud can be solved, especially in areas with insufficient ambient light or missing texture. Compared with traditional methods, this solution uses the depth information of the lidar point cloud to provide more accurate position information for the Gaussian ellipsoid, effectively reducing the appearance of floating objects. Further, a depth image is generated based on the lidar point cloud data and the camera pose P, and the depth image is subjected to truncated depth processing to obtain an accurate depth value Z. The depth value Z, the lidar point cloud data, and the image data are fused to obtain the image to be processed, combining the information of different data sources to improve the integrity of the model. The fusion of the generated depth image and the lidar point cloud data effectively constrains the positioning of the Gaussian model, ensuring the reasonable distribution of the Gaussian ellipsoid in space. By fusing the information of different data sources, the information deficiency or noise problem that may be brought by a single data source is solved, and the robustness of the modeling is improved. Further, based on the training image data and the image D to be processed, 3D Gaussian optimization operations are performed by updating the initialized 3D Gaussian algorithm parameters (including rotation, scale, color, position opacity, voxel index, etc.) to gradually improve the model during the training process. Through optimization, the 3D Gaussian model becomes more accurate.Guided by the training image data, the optimized Gaussian parameters can better match the structures and textures in the actual scene, reducing the reconstruction errors caused by poor initialization or excessive sparsification. Further, by recording the backpropagation gradients during the gradient descent process and screening the Gaussian distributions based on these gradients, unnecessary Gaussian distributions are eliminated. Through the dynamic threshold and the voxel internal point number limit strategy, the excessive Gaussian distributions in the dense regions are reduced, thereby reducing the training cost and rendering pressure. In particular, by applying the voxel internal point number limit and the dynamic threshold elimination strategy, the density of the Gaussian is reduced, resulting in fewer ellipsoids in the dense areas and reducing the training cost and rendering pressure.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; those of ordinary skill in the art can modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A 3D Gaussian weak texture compensation and density control reconstruction method, characterized in that: The steps are as follows: Acquire data to be processed; perform posture recovery based on the data to be processed through the Colmap framework to obtain a camera posture P; obtain first sparse point cloud data based on the camera posture P; The data to be processed includes laser point cloud data, image data and position data; The laser point cloud data is subjected to sparse processing and then merged with the first sparse point cloud data to obtain first 3D Gaussian algorithm parameters; the first 3D Gaussian algorithm parameters are initialized to obtain initialized 3D Gaussian algorithm parameters; A depth image is calculated based on the laser point cloud data and the camera pose P; a depth image is truncated and processed to obtain a depth value Z; a fusion process is performed based on the depth value Z, the laser point cloud data and the image data to obtain an image to be processed; Performing a 3D Gaussian optimization operation using the training image data and the image D to be processed, updating the initialized 3D Gaussian algorithm parameters, and obtaining updated 3D Gaussian algorithm parameters; Record the back propagation gradient in the process of the gradient descent algorithm reversely updating the initialized 3D Gaussian algorithm parameters; and screen based on the back propagation gradient to obtain a Gaussian distribution.

2. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 1, characterized in that: The first 3D Gaussian algorithm parameters include rotation parameters, scale parameters, color parameters and position opacity parameters in the 3D Gaussian model; The initialized 3D Gaussian algorithm parameters include a rotation parameter, a scale parameter, a color parameter, a position opacity parameter and a voxel index T in the 3D Gaussian model.

3. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 2, characterized in that: The method further comprises: obtaining the data to be processed by the acquisition; performing posture recovery through the Colmap framework based on the data to be processed to obtain the camera posture P; obtaining the first sparse point cloud data based on the camera posture P, and further comprising calculating an offset based on the camera posture P and the posture data; determining whether the offset is greater than or equal to a preset offset threshold, and if so, filtering out the data to be processed corresponding to the camera posture P.

4. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 3, characterized in that: The step of merging the laser point cloud data with the first sparse point cloud data after sparse processing to obtain first 3D Gaussian algorithm parameters includes the following steps: Performing voxel downsampling processing on the laser point cloud data to obtain downsampled laser point cloud data; Performing voxel downsampling processing on the first sparse point cloud data to obtain downsampled first sparse point cloud data; Fusing the laser point cloud data after the medicine is applied and the first sparse point cloud data after the downsampling to obtain fused data; A first 3D Gaussian algorithm parameter is obtained based on the fused data.

5. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 4, characterized in that: The converting and calculating based on the depth value Z and the laser point cloud data to obtain the image to be processed includes the following steps: Get camera intrinsic parameters and camera extrinsic parameters; Based on the laser point cloud data, the camera coordinate system point cloud P is obtained by calculating the camera external parameters. c ; The camera coordinate system point cloud P c Projecting the camera intrinsic parameters onto the image plane to obtain a projected image; Get the depth value Z' to be filled; The depth value Z' to be filled is filled into the projection image to obtain the image D to be processed; the depth value Z' to be filled is a depth value whose depth value Z is less than or equal to a depth value threshold.

6. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 5, characterized in that: The camera coordinate system point cloud P c The calculation method is: P c =RP world +t; Where P world =(x, y, z) is the laser point cloud data, P c =[X, Y, Z] is the point cloud of the camera coordinate system, R is the rotation matrix of the camera extrinsic parameters; t is the translation vector of the camera extrinsic parameters; The projection image is calculated as follows: Where K is the camera internal parameter, f x and f y is the focal length of the camera internal parameter, c x and c y is the principal point position in the camera intrinsic parameters; The image D to be processed is expressed as: D=+∞ Where H and W are the image height and width.

7. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 6, characterized in that: The method of performing a 3D Gaussian optimization operation using the training image data and the image to be processed D, updating the initialized 3D Gaussian algorithm parameters, and obtaining updated 3D Gaussian algorithm parameters includes the following steps: Acquire training image data; the training image data includes training image I i And the training pose P corresponding to the training image i ; Based on the training image I i And the training pose P i The rendering viewing angle W is obtained by calculating the parameters of the initialized 3D Gaussian algorithm. i And the training image I i The corresponding depth image D' i ; Based on the training image I i , the rendering perspective W i , the depth image D' i And the image D to be processed is calculated to obtain a loss function L; Based on the loss function L, the initialized 3D Gaussian algorithm parameters are reversely updated through a gradient descent algorithm to obtain updated 3D Gaussian algorithm parameters.

8. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 7, characterized in that: The loss function L is calculated as follows: L=(1-λ)L1(I i ,W i )+λL SSIM (I i ,W i )+λ depth L1(D i ,D’ i ); Among them, λ is the L1 and SSIM loss weights, λ depth is the depth loss weight; The updated 3D Gaussian algorithm parameters include updated rotation parameters, updated scale parameters, updated color parameters, updated position opacity parameters and updated voxel index T'.

9. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 8, characterized in that: The step of screening based on the back propagation gradient to obtain a Gaussian distribution includes the following steps: Screening and obtaining a plurality of Gaussian distributions to be processed whose back-propagation gradient is greater than or equal to a preset gradient threshold; Counting the number of Gaussian points in each of the Gaussian distributions to be processed; selecting a Gaussian distribution whose number of Gaussian points is greater than or equal to a preset Gaussian point number threshold as a target Gaussian distribution to be processed; The Gaussian distribution V′ is obtained by calculating based on the scale S of each Gaussian in the target Gaussian distribution to be processed.

10. A 3D Gaussian weak texture compensation and density control reconstruction method according to claim 9, characterized in that: The Gaussian distribution V′ is calculated as follows: V′={i∈V||S i -μ|≤δσ}; in, δ is a constant coefficient and V′ is the Gaussian distribution preserved within the voxel.

Citation Information

Cited By

  • Weak texture scene-oriented three-dimensional model reconstruction method and system, intelligent terminal and storage medium

    CN120526064A

  • Three-dimensional model reconstruction method and system for weak texture scene, intelligent terminal and storage medium

    CN120526064B