Storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering
The storage optimization of three-dimensional Gaussian sputtering through improved adaptive density control and sensitivity-aware vector clustering technology solves the high storage requirements in three-dimensional scene reconstruction and improves rendering efficiency and data transmission performance.
Patent Information
- Application Number
- CN202510306262.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
Three-dimensional Gaussian sputtering has a high storage requirement in the process of three-dimensional scene reconstruction, especially because the parameters of a large number of three-dimensional Gaussian primitives occupy a large amount of storage space, which affects the rendering efficiency and data transmission.
The improved adaptive density control strategy is adopted to optimize the spatial distribution of three-dimensional Gaussian primitives through the synchronous position gradient, pixel weight gradient and gradual enhancement weight gradient synthesis, and the low opacity Gaussian primitives are deleted in the attribute compression module. The color characteristics, covariance matrix and opacity parameters of Gaussian primitives are encoded and compressed by sensitivity perceptual vector clustering technology, and finally the storage optimization is optimized using the DEFLATE compression format.
It effectively reduces storage demand, improves the efficiency and data transmission performance of three-dimensional scene reconstruction, and reduces the use of storage space.
Smart Images

Figure BDA0005313042250000031 
Figure BDA0005313042250000033 
Figure BDA0005313042250000035
Abstract
Description
Technical Field:
[0001] The present invention relates to the technical field of 3D scene reconstruction, and particularly to a storage optimization method for 3D scene reconstruction based on 3D Gaussian Splatting. Background Art:
[0002] Novel View Synthesis (NVS) aims to generate new 3D views from scenes captured in pictures or videos, and this field has made remarkable progress in the past few decades. Traditional multi-view stereo methods have been gradually replaced by deep learning, especially the Neural Radiance Field (NeRF) method.
[0003] NeRF and its improved methods directly render new view scenes by using a small neural network MLP and volume rendering, which are widely used in 3D reconstruction. NeRF itself has demonstrated very good novel view synthesis quality. However, the neural network in NeRF has high computational cost and long training time, seriously reducing the efficiency of novel view synthesis. To solve this problem, 3D Gaussian Splatting (3DGS) emerged.
[0004] 3DGS is a point-based rendering method that uses the sparse point cloud extracted from Structure from Motion (SfM) as initialization, expands each point cloud into an anisotropic 3D Gaussian basis element with parameters such as shape, color, and opacity, and uses a highly customized CUDA kernel algorithm for rasterization pipeline rendering. The parameters of 3DGS are differentiable, so during the process of matching images, it is optimized through differentiable rendering. However, 3DGS requires a large number of 3D Gaussian basis elements to maintain the high fidelity of the rendered image, and each 3D Gaussian basis element has up to 62 parameters, occupying a large amount of storage space, which poses a great challenge to rendering efficiency and data transmission. Summary of the Invention:
[0005] The technical problem to be solved by the present invention is to provide a storage optimization method for 3D scene reconstruction based on 3D Gaussian Splatting, which is used to solve the problem of high storage requirements in the process of 3D scene reconstruction using 3D Gaussian Splatting.
[0006] The technical solution adopted by the present invention to solve its technical problems is: to provide a storage optimization method for 3D scene reconstruction based on 3D Gaussian Splatting, including the following steps:
[0007] Capture 2D scene images from multiple perspectives;
[0008] Input the 2D scene images from multiple perspectives into the 3D Gaussian Splatting model to achieve 3D reconstruction of the scene;
[0009] Input the reconstructed model into the attribute compression module for efficient compression processing; wherein, the three-dimensional Gaussian sputtering model includes:
[0010] An improved adaptive density control strategy for gradually adjusting the gradient weight according to the increase in the number of iterations to achieve more precise density control;
[0011] The attribute compression module includes:
[0012] Delete the low-opacity Gaussian basis elements;
[0013] Use the sensitivity-aware vector clustering technique to encode and compress the color features, covariance matrix, and opacity parameters of the Gaussian basis elements.
[0014] The improved adaptive density control strategy includes:
[0015] Co-directional position gradient, only retain the magnitude of the gradient of each three-dimensional Gaussian basis element, regardless of the direction, and take the absolute value of each three-dimensional Gaussian basis element gradient component for accumulation;
[0016] Pixel weight gradient, use the number of pixels covered by each Gaussian basis element as a weight of the gradient;
[0017] Gradually enhanced weight gradient, propose to add the number of iterations to the cross-view gradient averaging to amplify the gradient contribution of the three-dimensional Gaussian basis elements with higher iteration numbers;
[0018] Finally, the co-directional position gradient, pixel weight gradient, and gradually enhanced weight gradient are combined to form the total gradient of the Gaussian basis elements.
[0019] The co-directional position gradient is specifically where L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian point participates in every 100 iterations, represents the co-directional position gradient of each Gaussian basis element.
[0020] The pixel weight gradient is specifically where L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian point participates in every 100 iterations, represents the number of pixels covered by the three-dimensional Gaussian basis element of the i-th Gaussian basis element under view k.
[0021] The gradually enhanced weight gradient is specifically Among them, since densification is performed every 100 times, d i represents the current iteration number during these 100 times of densification, and M represents the total number of views that the Gaussian points participate in every 100 iterations.
[0022] The total gradient of the Gaussian basis element is specifically where L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian points participate in every 100 iterations, represents the number of pixels covered by the three-dimensional Gaussian basis element of the i-th Gaussian basis element under view k, represents the co-directional position gradient of each Gaussian basis element, and d i represents the current iteration number during every 100 times of densification.
[0023] The attribute compression module includes:
[0024] Every 100 iterations, delete the three-dimensional Gaussian basis elements with opacity lower than 0.05;
[0025] Adopt the sensitivity-aware vector clustering technology to encode and compress the color features, covariance matrices and opacity parameters of the Gaussian points.
[0026] The sensitivity-aware technology is specifically where N represents the total number of views in the training set for scene reconstruction, P k is the total number of pixels in view k, E k is the total energy of view k, that is, the sum of the GRB components on all pixels. To measure the sensitivity of Ek to the change of q, E k is represented by the gradient with respect to q, and q represents the color features, covariance matrices and opacity parameters.
[0027] The vector clustering technology is specifically that for the three-dimensional Gaussian basis elements with higher sensitivity to color features, covariance matrices and opacity than the threshold, the K-means clustering algorithm is used for processing.
[0028] The compression method is specifically the DEFLATE compression format, which mainly combines the LZ77 algorithm and Huffman coding.
[0029] Compared with the prior art, the remarkable progress of the present invention lies in that the present invention adopts an adaptive density control strategy formed by comprehensively combining the same-direction position gradient, pixel weight gradient, and gradually enhanced weight gradient to optimize and adjust the spatial distribution of three-dimensional Gaussian basis elements, and effective results are obtained in reducing storage. During attribute compression, first, extremely low opacity is deleted, and then the sensitivity-aware vector clustering technology is used to encode the color features, covariance matrix, and opacity parameters of the Gaussian basis elements. Finally, the DEFLATE compression format is used for compression. The present invention effectively solves the problem of high storage requirements during the three-dimensional scene reconstruction using three-dimensional Gaussian sputtering. Description of the Drawings:
[0030] Figure 1 is a flowchart of the storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering according to an embodiment of the present invention;
[0031] Figure 2 is a framework diagram of the method for reconstructing a three-dimensional scene model and optimizing storage in an embodiment of the present invention Detailed Embodiments:
[0032] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0033] An embodiment of the present invention relates to a storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering. This method is based on the three-dimensional Gaussian sputtering algorithm and proposes an adaptive density control strategy formed by comprehensively combining the same-direction position gradient, pixel weight gradient, and gradually enhanced weight gradient to achieve three-dimensional reconstruction of the scene; subsequently, the reconstructed model is input into an attribute compression module for efficient compression processing. As Figure 1 shown, it includes the following steps:
[0034] Step 1, capture 2D scene images from multiple perspectives;
[0035] Step 2, input the 2D scene images from multiple perspectives into a three-dimensional Gaussian sputtering model to achieve three-dimensional reconstruction of the scene;
[0036] Step 3, input the reconstructed model into an attribute compression module for efficient compression processing.
[0037] As Figure 2 shown, the three-dimensional Gaussian sputtering model in this embodiment includes:
[0038] An improved adaptive density control strategy is used to gradually adjust the gradient weight according to the increase in the number of iterations to achieve more accurate density control;
[0039] In this embodiment, the attribute compression module includes:
[0040] Delete the Gaussian basis elements with low opacity;
[0041] Use the sensitivity-aware vector clustering technique to encode and compress the color features, covariance matrix, and opacity parameters of the Gaussian basis elements.
[0042] In the three-dimensional Gaussian sputtering model of this embodiment, in the improved adaptive density control strategy, the co-directional position gradient, pixel weight gradient, and gradually enhanced weight gradient are comprehensively considered.
[0043] In the over-reconstructed area of a complex scene containing high-frequency details, the original adaptive density control strategy only uses the average gradient in the view space as the criterion for whether each Gaussian needs to be densified, resulting in a blurred rendered image. However, the pixel-level gradient directions of each three-dimensional Gaussian basis element are not the same, which will cause gradient conflicts, resulting in gradient attenuation, thus ignoring the densification of a certain Gaussian. The co-directional position gradient only retains the magnitude of the gradient, regardless of the direction, and takes the absolute value of each gradient component for accumulation. Specifically, the co-directional position gradient is specifically expressed as:
[0044]
[0045] Among them, L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian point participates in every 100 iterations, represents the co-directional position gradient of each Gaussian basis element.
[0046] In addition, the number of pixels covered by some large Gaussians is too large, which will make the final average gradient too small, and these points are difficult to clone and split, resulting in an unsatisfactory modeling effect in these areas. Therefore, pixels are used as the weight of the gradient. Specifically expressed as:
[0047]
[0048] Among them, L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian point participates in every 100 iterations, represents the number of pixels covered by the three-dimensional Gaussian basis element of the i-th Gaussian basis element under view k.
[0049] Meanwhile, this embodiment believes that after each densification, as the number of iterations increases, the three-dimensional Gaussian basis elements are continuously optimized, and the later gradients should have higher weights. Therefore, it is proposed to add the number of iterations to the cross-view gradient averaging to amplify the gradient contribution of the three-dimensional Gaussian basis elements with higher iteration numbers. Specifically, it is expressed as:
[0050]
[0051] Among them, since densification is performed every 100 times, d i represents the current iteration number during these 100 densification processes, and M represents the total number of views that the Gaussian points participate in every 100 iterations.
[0052] Finally, this embodiment combines the co-directional position gradient, pixel weight gradient, and gradually enhanced weight gradient to form the total gradient of the Gaussian basis elements, which is used to control the densification of the Gaussian basis elements and generate the final three-dimensional scene model. Specifically, it is expressed as:
[0053]
[0054] Among them, L k is the loss under view k, v i,k represents the planar coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian points participate in every 100 iterations, represents the number of pixels covered by the three-dimensional Gaussian basis element of the i-th Gaussian basis element under view k, represents the co-directional position gradient of each Gaussian basis element, and d i represents the current iteration number during every 100 densification processes.
[0055] Subsequently, the reconstructed model is input into the attribute compression module for efficient compression processing. In this embodiment, the attribute compression module includes:
[0056] Delete the three-dimensional Gaussian basis elements with opacity lower than 0.05 every 100 iterations;
[0057] Use the sensitivity-aware vector clustering technology to encode and compress the color features, covariance matrix, and opacity parameters of the Gaussian basis elements.
[0058] In this embodiment, through observation, in the original 3DGS, every 100 iterations, the three-dimensional Gaussian basis elements with an opacity lower than 0.005 are deleted. However, after deleting the three-dimensional Gaussian basis elements with low opacity below a certain threshold, the image quality does not decrease significantly, which indicates that there are a large number of redundant three-dimensional Gaussian basis elements with extremely low opacity in the three-dimensional scene model generated by the original 3DGS. Through analysis and testing, since the human eye cannot distinguish when the PSNR difference is within 0.5, 0.5 is selected as the threshold for deleting opacity.
[0059] To optimize the storage efficiency, the present invention performs clustering encoding on the parameters with low sensitivity to the image among the color features, covariance matrices, and opacities of the Gaussian basis elements, so that these Gaussian basis elements share the parameters in the codebook, effectively reducing the storage requirements. Specifically:
[0060] The sensitivity perception technology is specifically expressed as
[0061]
[0062] Among them, N represents the total number of views in the training set for scene reconstruction, P k is the total number of pixels in view k, and E k is the total energy of view k, that is, the sum of the GRB components over all pixels. To measure the sensitivity of E k to the change in q, the gradient of E k with respect to q is used for representation, and q represents the color feature, covariance matrix, and opacity parameters.
[0063] The vector clustering technology is specifically as follows: For the three-dimensional Gaussian basis elements with higher sensitivity to the color feature, covariance matrix, and opacity than the threshold, the K-means clustering algorithm is used for processing. First, randomly and uniformly initialize the codebook of the opacity parameters of these three-dimensional Gaussian basis elements, and then update the codebook in an iterative optimization manner to achieve a more stable result. The codebook size is set to 2048.
[0064] After successful encoding, overall fine-tuning is also required to make the encoded three-dimensional model perform better.
[0065] Finally, the fine-tuned model is further compressed. Specifically, the DEFLATE compression format is used, which mainly combines the LZ77 algorithm and Huffman coding.
[0066] It is not difficult to find that the present invention adopts an adaptive density control strategy formed by comprehensively combining a same-direction position gradient, a pixel weight gradient, and a gradually enhanced weight gradient to optimize and adjust the spatial distribution of three-dimensional Gaussian basis elements, and effective results are obtained in terms of reducing storage. During attribute compression, first, extremely low opacity is deleted, then the sensitivity-aware vector clustering technology is used to encode the color features, covariance matrix, and opacity parameters of the Gaussian basis elements, and finally, the DEFLATE compression format is used for compression. The present invention effectively solves the problem of high storage requirements in the process of three-dimensional scene reconstruction using three-dimensional Gaussian sputtering.
Claims
1. A storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering, characterized in that: The following steps are involved: Capture 2D scene images from multiple viewpoints; Input 2D scene images from multiple perspectives into the 3D Gaussian sputtering model to achieve 3D reconstruction of the scene; The reconstructed model is input into the attribute compression module for efficient compression processing; wherein the three-dimensional Gaussian sputtering model includes: Improved adaptive density control strategy, which is used to gradually adjust the gradient weight according to the number of iterations to achieve more accurate density control; The attribute compression module comprises: Remove low opacity high-skirts; The sensitivity-aware vector clustering technique is used to further encode and compress the opacity parameters of the Gaussian primitives based on the encoded color features and covariance matrix.
2. The storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering according to claim 1 is characterized in that: The improved adaptive density control strategy includes: We progressively enhance the weight gradient and propose to add the iteration number to the gradient average across views to amplify the gradient contribution of 3D Gaussian primitives with higher iteration number.
3. The storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering according to claim 2 is characterized in that: The improved adaptive density control strategy is specifically: Among them, L k is the loss under view k, v i,k represents the plane coordinates of the i-th Gaussian point under view k, M represents the total number of views that the Gaussian point participates in every 100 iterations, represents the number of pixels covered by the 3D Gaussian primitive under view k for the i-th Gaussian primitive, represents the same-direction position gradient of each Gaussian basis element, d i Indicates the current iteration number in the process of densification every 100 times.
4. The storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering according to claim 1, characterized in that: The attribute compression module comprises: Every 100 iterations, remove the 3D Gaussian primitives with opacity lower than 0.05; The sensitivity-aware vector clustering technique is used to further encode and compress the opacity parameters of the Gaussian primitives based on the encoded color features and covariance matrix.
5. The storage optimization method for three-dimensional scene reconstruction based on three-dimensional Gaussian sputtering according to claim 4 is characterized in that: The opacity parameter of the Gaussian primitive is further encoded, specifically: Where N is the total number of views in the training set used for scene reconstruction, and P k is the total number of pixels in view k, E k is the total energy of view k, that is, the sum of the RGB components of all pixels. In order to measure E k Sensitivity to changes in q, using E k The gradient relative to q is expressed, and q represents the opacity parameter. For three-dimensional Gaussian primitives with opacity sensitivity higher than the threshold, the K-means clustering algorithm is used for processing.
Citation Information
Cited By
Improved AD-GS three-dimensional reconstruction method based on 2DGS
CN120807798A
Three-dimensional model reconstruction method and device based on Gaussian sputtering model
CN121392161A
Operation scene real-time dynamic reconstruction method and system based on three-dimensional Gaussian sputtering
CN121962396A