Efficient 3D Reconstruction Method Based on Data-Driven Compression and Adaptive Splitting
By introducing data-driven compression and adaptive splitting methods into 3DGS technology, the problems of Gaussian distribution adjustment and computing efficiency in complex scenarios are solved, and efficient three-dimensional reconstruction and high-fidelity image rendering are achieved.
Patent Information
- Application Number
- CN202411676372.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing 3D Gaussian Splatting (3DGS) technology has storage and computing efficiency issues when dealing with complex scenarios, and it is difficult to dynamically adjust the Gaussian distribution to preserve image details.
Using data-driven compression and adaptive splitting methods, 3D Gaussians are adaptively split by image gradient information, redundant Gaussian culling is used using generative adversarial network (GAN), and Gaussian contribution is dynamically evaluated through importance weight sampling technology.
It significantly reduces the number of redundant 3D Gaussians, reduces computing resource consumption, while maintaining the high fidelity and visual effects of the scene, and improving the efficiency of three-dimensional reconstruction.
Smart Images

Figure CN119206088B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and specifically relates to an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting. Background Art
[0002] Voice 3D scene reconstruction has important application values in fields such as virtual reality, augmented reality, and film and television special effects. Traditional three-dimensional scene reconstruction methods mainly rely on geometric representations such as voxel grids and triangular grids. Although these methods can achieve good results in specific scenarios, they face problems of storage and computational efficiency when dealing with complex scenes. With the progress of computing technology, the 3D Gaussian Splatting (3DGS) technology is regarded as a new, feasible, and efficient three-dimensional scene representation method. 3DGS constructs a scene by using millions of elliptical 3D Gaussians and can perform image rendering in a high-quality manner. Each 3D Gaussian is defined by parameters such as its position, transparency, covariance matrix, and spherical harmonic coefficients, thereby realizing high-fidelity image generation.
[0003] However, existing 3DGS methods may have problems such as image blurring and detail loss in the high-frequency region. Generally, 3DGS needs to manage and process a large number of 3D Gaussians, which will lead to huge memory consumption and computational resource requirements. Especially in real-time applications, this poses a great challenge to the performance of the system. In addition, most existing methods rely on fixed rules to reduce the number of 3D Gaussians and lack the ability to dynamically adjust the Gaussian distribution according to scene details, resulting in possible insufficient retention of details in key regions of the image.
[0004] Therefore, the present invention proposes an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting, aiming to make full use of key information in the scene, reduce redundant 3D Gaussians, and maintain the high fidelity of the scene at the same time. First, based on image gradient information, the 3D Gaussians are adaptively split, and more Gaussians are retained in regions with rich image details to reduce image blurring. Secondly, a pruning mechanism based on the generative adversarial network (GAN) is used to effectively eliminate redundant 3D Gaussians. Finally, the contribution of 3D Gaussians to the rendering quality is dynamically evaluated through importance weight sampling technology, further optimizing the selection and retention of Gaussians, thereby significantly improving the efficiency of scene reconstruction. Summary of the Invention
[0005] The object of the present invention is to provide an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting. This method effectively reduces the number of redundant Gaussians and adaptively adjusts the Gaussian distribution, achieving high-precision modeling of the scene with a small number of Gaussians, reducing computational resource consumption, and maintaining the visual effect at the same time.
[0006] To achieve the above functions, the present invention designs an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting, and performs the following steps S1 - S7 to complete the three-dimensional reconstruction of the target scene:
[0007] Step S1: Obtain real images of the target scene from different perspectives, generate a sparse point cloud through the SfM method, and initialize the sparse point cloud as multiple 3D Gaussians. The attributes of the 3D Gaussians include the central position, size, shape in space, color, transparency, and learnable features in the three-dimensional space;
[0008] Step S2: Construct a mask generation network and a discriminator network. The mask generation network selects the 3D Gaussians to be eliminated based on the learnable features of the 3D Gaussians, generates a rendered image based on data-driven, and outputs a mask; the discriminator network is used to judge the difference between the rendered image and the real image;
[0009] Step S3: Construct a differentiable raster renderer for optimizing the attributes of the 3D Gaussians;
[0010] Step S4: Randomly select a perspective, input the real image of this perspective into the differentiable raster renderer, and render the 3D Gaussians under this perspective through the differentiable raster renderer to generate a complete 3D Gaussian rendered image;
[0011] Step S5: The mask generation network outputs the importance scores of the 3D Gaussians according to the learnable features of the 3D Gaussians and the perspective information, generates a mask according to a threshold; uses the generated mask to eliminate specific 3D Gaussians; obtains an image after Gaussian elimination, renders the image after Gaussian elimination, and generates a rendered image of the eliminated part of the 3D Gaussians;
[0012] Step S6: The discriminator network compares the complete 3D Gaussian rendered image, the rendered image of the eliminated part of the 3D Gaussians, and the real image; constructs loss functions for the mask generation network, the discriminator network, and scene reconstruction, and through backpropagation, uses the Adam optimizer to realize iterative optimization and update optimization of the parameters of the mask generation network and the discriminator network and the attributes of the 3D Gaussians;
[0013] Step S7: Judge whether the preset number of iterations is reached. If so, enter the Gaussian reduction stage. Based on the mask generated by the mask generation network and the importance weighted sampling method, calculate the total importance score of the 3D Gaussians, evaluate the contribution of the 3D Gaussians to the target scene, screen out the 3D Gaussians with a total importance score greater than the preset value for retention, eliminate other 3D Gaussians, and output the optimized Gaussian representation;
[0014] If the preset number of iterations is not reached, the Gaussian growth stage is entered. Incomplete 3D Gaussians for reconstruction are screened and processed according to their scale sizes. 3D Gaussians with a scale larger than the preset value are adaptively split, and the others are duplicated. Subsequently, return to step S2 to continue the iteration until the preset number of iterations is reached, completing the three-dimensional reconstruction of the target scene.
[0015] Advantageous effects: Compared with the prior art, the advantages of the present invention include:
[0016] The present invention proposes an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting, which fully utilizes the key details in the scene, reduces the number of redundant 3D Gaussians, thereby reducing the storage requirements and computational costs. This method uses a data-driven mask generation adversarial network to dynamically reduce the number of Gaussians, ensuring high fidelity of the scene while reducing the number of Gaussians. By combining the mask generation network and the generative adversarial network, the Gaussian distribution of the scene is effectively optimized. This method has significantly improved both the speed and quality of three-dimensional reconstruction and is applicable to virtual reality, augmented reality, and other scenarios that require efficient rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of an efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting according to an embodiment of the present invention;
[0018] Figure 2 is a schematic structural diagram of a mask generation network according to an embodiment of the present invention;
[0019] Figure 3 is a schematic structural diagram of a discriminant network according to an embodiment of the present invention;
[0020] Figure 4 is a flowchart of the adversarial training of a mask generation network and a discriminant network according to an embodiment of the present invention;
[0021] Figure 5 is a schematic diagram of the 3D Gaussian adaptive density adjustment process according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0023] The efficient three-dimensional reconstruction method based on data-driven compression and adaptive splitting provided by the embodiment of the present invention, referring to Figure 1 , performs the following steps S1 - step S7 to complete the three-dimensional reconstruction of the target scene:
[0024] Step S1: Obtain real images of the target scene from different perspectives, generate a sparse point cloud through the SfM method, and initialize the sparse point cloud into multiple 3D Gaussians. The attributes of the 3D Gaussians include the central position in three-dimensional space, size, shape in space, color, transparency, and learnable features.
[0025] The 3D Gaussian described in Step S1 is defined as follows:
[0026] ;
[0027] In the formula, G i ( x ) is a 3D Gaussian, defined as an ellipsoid with specific attributes. x represents each point within the region of the ellipsoid. μ i , Σ i are the attributes of the 3D Gaussian. Among them, μ i is the central position of the 3D Gaussian in three-dimensional space; Σ i represents the covariance matrix of size , which defines the size of the 3D Gaussian and its shape in space;
[0028] The attributes of the 3D Gaussian also include color, transparency, and learnable features. Among them:
[0029] Color: Represented by the coefficients of a 4th-order spherical harmonic function, which describes the color change of the 3D Gaussian under different lighting conditions;
[0030] Opacity: Represented by a scalar α i , which describes the visibility of the 3D Gaussian in the final rendered image. This attribute is used to control the visual contribution of the 3D Gaussian and affects the superposition and fusion effects between different 3D Gaussians;
[0031] Learnable Feature: Represented by a nine-dimensional first-order tensor , which is passed as input to the mask generation network to describe the importance of the 3D Gaussian; it is optimized during the training process to help generate the mask and determine which 3D Gaussians need to be retained or eliminated.
[0032] Step S2: Construct a mask generation network and a discriminator network. The mask generation network selects the 3D Gaussians to be eliminated based on the learnable features of the 3D Gaussians, generates a rendered image with some 3D Gaussians eliminated based on data-driven, and outputs a mask. The discriminator network is used to judge the difference between the rendered image with some 3D Gaussians eliminated and the real image.
[0033] The structures and functions of the mask generation network and the discriminator network described in Step S2 are as follows:
[0034] The mask generation network is implemented using tiny-cuda-nn, which is a neural network library designed for efficient parallel computing. It uses ReLU as the activation function to enhance the non-linear processing ability. The input of the mask generation network includes the learnable features of the 3D Gaussians and view information. The core function of the mask generation network is to generate an accurate mask based on these inputs, thereby optimizing the representation and rendering quality of the scene. Among them, the mask generation network decides whether each 3D Gaussian should be retained or eliminated according to the learnable features; the view information helps the mask generation network process the 3D Gaussians from the best angle and optimize the rendering; the mask generation network processes the 3D Gaussians according to the view information, generates a rendered image with some 3D Gaussians eliminated, and outputs a mask.
[0035] The structure of the mask generation network is as Figure 2 shown. Inside the mask generation network, the input learnable features and view information pass through a series of hidden layers and finally generate a mask, which is used to decide which 3D Gaussians should be retained or eliminated. This mechanism can improve the rendering efficiency and quality by optimizing the Gaussian distribution.
[0036] The discriminator network adopts a convolutional neural network (CNN) structure. The structure of the discriminator network is as Figure 3 shown. It extracts features and evaluates the authenticity of the generated rendered image with some 3D Gaussians eliminated through multiple convolutional layers and pooling layers, compares the generated rendered image with some 3D Gaussians eliminated with the real image, and outputs a discriminant result. The mask generation network gradually optimizes the generated mask by minimizing the feedback of the discriminator network. This adversarial learning ensures that the mask generation network can maintain high accuracy and visual effects of the scene while reducing the number of 3D Gaussians. This adversarial learning mechanism is shown in Figure 4 where the interaction process between the mask generation network and the discriminator network ensures efficient mask generation and image quality improvement.
[0037] Step S3: Construct a differentiable raster renderer for optimizing the attributes of 3D Gaussians.
[0038] Step S4: Randomly select a view, input the real image of this view into the differentiable raster renderer, and render the 3D Gaussians under this view through the differentiable raster renderer to generate a complete 3D Gaussian rendered image.
[0039] Step S5: The mask generation network outputs the importance score of the 3D Gaussian based on the learnable features of the 3D Gaussian and the view information, generates a mask according to a threshold; uses the generated mask to eliminate specific 3D Gaussians; obtains the image after Gaussian elimination, and renders the image after Gaussian elimination to generate a rendered image with some 3D Gaussians eliminated;
[0040] The specific process of step S5 is as follows:
[0041] Step S5.1: The mask generation network adopts a two-layer multi-layer perceptron (MLP), denoted as , and the mask generation network inputs the learnable features and the view information to generate the corresponding importance score of the 3D Gaussian. The importance score generation process is as follows:
[0042] ;
[0043] wherein, represents concatenating the learnable features F i and the view information θ i as the input of the mask generation network, I i represents the importance score of the 3D Gaussian; the importance score reflects the importance of each 3D Gaussian for scene rendering under the current view;
[0044] Step S5.2: The mask generation process not only depends on the application of the threshold, but also combines a more complex strategy; the importance score I i of each 3D Gaussian is used to generate the mask M i , and the generation formula is as follows:
[0045] ;
[0046] wherein, is the indicator function, is the stop gradient operation, and threshold represents the threshold;
[0047] The threshold threshold is calculated by the following formula:
[0048] ;
[0049] wherein, , , N is the current number of 3D Gaussians, is the proportionality coefficient, is the maximum number of 3D Gaussians;
[0050] The above formula performs a logarithmic transformation on the current number of 3D Gaussians N so that the threshold can be smoothly adjusted according to the change in the number of 3D Gaussians, thereby effectively screening out 3D Gaussians with less contribution.
[0051] Step S5.3: The generated mask M i is used to select and process the corresponding 3D Gaussians. The 3D Gaussians are screened using the threshold threshold. The screened 3D Gaussians will continue to be rendered through the 3D Gaussian point algorithm to generate a rendered image with some 3D Gaussians removed.
[0052] Step S6: The discriminative network compares the complete 3D Gaussian rendered image, the rendered image with some 3D Gaussians removed, and the real image; constructs the loss functions of the mask generation network, the discriminative network, and the scene reconstruction. Through backpropagation, the Adam optimizer is used to achieve iterative optimization and update optimization of the parameters of the mask generation network and the discriminative network, as well as the attributes of the 3D Gaussians;
[0053] The loss function of the mask generation network in Step S6 Loss G is as follows:
[0054] ;
[0055] where represents the L2 norm of the learnable features, σ represents the Sigmod operation; by minimizing the loss function Loss G , the mask generation network can dynamically optimize the generated mask to ensure that the importance of the 3D Gaussians is effectively evaluated. Select as many 3D Gaussians as possible to retain more scene information, but at the same time ensure that the number of these 3D Gaussians does not have a negative impact on the final reconstruction quality. The loss function Loss G effectively reduces the number of unnecessary 3D Gaussians by controlling the selection of 3D Gaussians, improving the computational efficiency.
[0056] The loss function of the discriminative network Loss D is as follows:
[0057] ;
[0058] where Loss R represents the loss of the real image, Loss F represents the loss of the rendered image with some 3D Gaussians removed;
[0059] The discriminative network receives two types of images as input: real images Img R and the rendered image with the 3D Gaussian removed Img F ; the loss of the real image Loss R and the loss of the rendered image with the 3D Gaussian removed Loss F The calculation formulas are as follows:
[0060] ;
[0061] ;
[0062] where D( Img R ) and D( Img F ) respectively represent the outputs of the discriminative network for the real image and the rendered image with the 3D Gaussian removed, BCE represents the binary cross-entropy loss, and MSE represents the mean squared error loss; by minimizing the loss of the discriminative network, the discriminative network can improve its ability to distinguish between real images and rendered images with the 3D Gaussian removed, thus providing more powerful feedback to the mask generation network and prompting it to generate higher-quality masks.
[0063] The loss function for scene reconstruction Loss Restruction is as follows:
[0064] ;
[0065] where Img represents the complete 3D Gaussian rendering, Img R is the real image, λ is a hyperparameter; L1 represents the L1 loss, and SSIM represents the structural similarity loss; the loss function Loss Restruction combines the L1 loss and the structural similarity loss to ensure that, in the case of 3D Gaussian removal, the difference between the complete Gaussian rendering image and the real image is minimized. The L1 loss measures the pixel difference between images, while SSIM ensures the structural consistency of the images.
[0066] The total loss function Loss total is as follows:
[0067] ;
[0068] where λ GRepresent the hyperparameters of the mask generation network, λ D represent the hyperparameters of the discriminative network. The total loss function combines the loss function for scene reconstruction Loss Restruction , the loss function of the mask generation network Loss G and the loss function of the discriminative network Loss D . Loss Restruction Ensure that the difference between the generated rendering map of the eliminated part of the 3D Gaussian and the real image is minimized, Loss G optimize the selection and elimination of 3D Gaussians, while Loss D help the mask generation network generate more realistic images through the feedback of the discriminative network. The hyperparameters λ G and λ D control the Loss G and Loss D weights in Loss total , thus ensuring that the mask generation network can maintain the image quality while reducing the number of 3D Gaussians.
[0069] Step S7: Determine whether the preset number of iterations is reached. If so, enter the Gaussian reduction phase. Based on the mask generated by the mask generation network and the importance weighted sampling method, calculate the total importance score of the 3D Gaussians, evaluate the contribution of the 3D Gaussians to the target scene, filter out the 3D Gaussians with a total importance score greater than the preset value for retention, eliminate other 3D Gaussians, and output the optimized Gaussian representation;
[0070] If the preset number of iterations is not reached, enter the Gaussian growth phase. Filter out the 3D Gaussians with incomplete reconstruction, process them according to the scale size, adaptively split the 3D Gaussians with a scale greater than the preset value, and copy the other 3D Gaussians; then return to step S2 to continue the iteration until the preset number of iterations is reached, and complete the three-dimensional reconstruction of the target scene.
[0071] In one embodiment, the preset number of iterations is half of the total number of iterations; that is, when the number of iterations is greater than or equal to half of the total number of iterations, it is regarded as the Gaussian reduction phase, and when the number of iterations is less than half of the total number of iterations, it is regarded as the Gaussian growth phase.
[0072] The Gaussian growth phase in step S7 specifically includes the following process:
[0073] Step S7.1.1: Screen out 3D Gaussians with scales larger than the preset value, define them as large Gaussians, and use the Scharr edge detection operator to calculate the complete 3D Gaussian rendering Img and the real image Img R gradients. The formula is as follows:
[0074] ;
[0075] ;
[0076] In the formula, G img represents the gradient of the complete 3D Gaussian rendering Img ; G ImgR represents the gradient of the real image Img R ;
[0077] Calculate the gradient difference G d :
[0078] ;
[0079] Step S7.1.2: Project the center position of each large Gaussian from three-dimensional space onto the image plane, and calculate the image gradient at the projection point; if the gradient difference value at the projection point of the 3D Gaussian is lower than 25% of the maximum gradient difference, then this 3D Gaussian is marked as incompletely reconstructed. The formula is as follows:
[0080] ;
[0081] In the formula, Threshold represents the threshold;
[0082] Step S7.1.3: Use a voxel grid to discretize the target scene into small spatial regions (voxels). Each voxel represents a specific spatial region. Calculate the density of the space by counting the number of 3D Gaussians within each voxel. The formula is as follows:
[0083] ;
[0084] Among them, d v is the density of voxel v ; is the number of Gaussians in voxel v ; vol ( v ) is the volume of voxel v ; In this way, a 3D Gaussian distribution density map of the entire scene can be obtained, providing a basis for subsequent splitting operations;
[0085] Step S7.1.4: For the large Gaussians marked as incompletely reconstructed, dynamically adjust the number of Gaussian splits based on the density of the space where they are located. The formula for the number of Gaussian splits is as follows:
[0086] ;
[0087] where, N i represents the number of splits of the i th large Gaussian, d i is the density of the space where the i th large Gaussian is located, and min(D) and max(D) are the minimum and maximum space densities respectively;
[0088] Step S7.1.5: Screen out the 3D Gaussians whose scale is less than or equal to the preset value, and define them as small Gaussians. Calculate the gradient of the center position of the small Gaussians through backpropagation, and determine whether these gradients are less than the preset threshold, so as to determine whether these small Gaussians fail to fully reconstruct details; for the small Gaussians with gradient values lower than the preset threshold, perform a copy operation to enhance their expression ability in reconstruction, thereby improving detail recovery.
[0089] Refer to Figure 5 , Figure 5 for the schematic diagram of the 3D Gaussian adaptive density adjustment process, which
[0090] shows two processing cases of 3D Gaussian splitting and copying.
[0091] Step S7.2.1: At the beginning of the Gaussian reduction stage, every 25 iterations, the mask generation network generates an importance score G i for each 3D Gaussian I i ;
[0092] Step S7.2.2: For the importance score I i output by the mask generation network, perform binarization processing to generate a mask M i , and the mask is used to identify which 3D Gaussians should be retained and which 3D Gaussians should be eliminated; the mask is generated every 25 iterations;
[0093] Step S7.2.3: In order to maintain the image quality while reducing the number of 3D Gaussians, perform importance-weighted sampling every 1000 iterations, and randomly select 95% of the 3D Gaussians to be retained each time; calculate the mixing weights and hit counts of each 3D Gaussian on each pixel in the view j ;
[0094] 3D Gaussian G i The mixing weight of ω i is calculated as follows:
[0095] ;
[0096] where, is an indicator function used to determine whether the 3D Gaussian G i intersects with the ray r u,v starting from the camera center and passing through the pixel (u, v); σ i is the opacity of the 3D Gaussian G i , H and W are the height and width of the image respectively;
[0097] The hit count of the 3D Gaussian G i c i represents the number of times the 3D Gaussian covers the pixel and is calculated as follows:
[0098] ;
[0099] Calculate the total importance score G i of the 3D Gaussian S i as follows:
[0100] ;
[0101] In the formula, N v represents the total number of viewpoints;
[0102] Normalize the total importance score for each 3D Gaussian, and the formula is as follows:
[0103] ;
[0104] In the formula, is the total importance score of the normalized 3D Gaussian G i . Through the total importance score of the normalized 3D Gaussian G i , during each weighted sampling, the probability of the 3D Gaussian being selected is proportional to its importance, thus ensuring that the 3D Gaussian with high importance has a higher chance of being retained and retaining the key details of the image.
[0105] To verify the effectiveness of the method of the present invention, comparative experiments and ablation experiments were conducted. First, the datasets used and training details are introduced, and then the comparative experiment results of different algorithms on the datasets are shown, and a series of ablation experiments are carried out to evaluate the effectiveness of the adversarial network composed of the mask generation network and the discriminator network, importance weighted sampling, and the adaptive splitting module.
[0106] The initial learning rate of the model was set to 0.0001 and gradually adjusted to 0.00005 during the training process. The batch size used was 1. To optimize the reduction of the number of Gaussians, no Gaussian reduction operation was performed in the first half of the training, and the focus was entirely on the training of the network. After reaching half of the total training iterations, the Gaussian reduction began. During the reduction phase, only the reconstruction loss was retained, and the adversarial network was no longer trained. During the reduction process, the Gaussian reduction frequencies of the adversarial network and importance sampling were different: the adversarial network performed Gaussian reduction every 25 iterations, while importance sampling performed reduction every 1000 iterations. This different reduction frequency ensured that in the later stage of model training, unimportant Gaussians could be more accurately reduced, thereby improving the rendering efficiency of the model while maintaining a high reconstruction accuracy.
[0107] Comparative experiments were conducted on the method proposed in the present invention and several current mainstream 3D reconstruction techniques. The experimental datasets included MipNeRF360 and Deep Blending. In the comparative experiments, the better results in the experimental data and pre-trained models were selected, covering models such as Compact3DGS, LightGaussian, and Mini-Splatting. The experimental results are shown in Table 1:
[0108]
[0109] To further verify the effectiveness of each module in the model, ablation experiments were conducted. In the experiments, the adversarial network, importance weighted sampling, and adaptive splitting module were removed respectively to evaluate the impact of these modules on the overall effect, and the performance of the model after removing different modules was compared with that of the complete model. The ablation experiment results are shown in Table 2. The tick in Table 2 indicates the retention of the corresponding module:
[0110]
[0111] Among the evaluation metrics, PSNR and SSIM represent the peak signal-to-noise ratio and structural similarity respectively, FPS represents the number of frames per second, and LPIPS measures the perceptual similarity. PSNR mainly reflects the clarity of image reconstruction. The higher the value, the closer the reconstructed image is to the original image. SSIM takes into account the similarity of the image in terms of brightness, contrast, and structure. LPIPS is used to measure the perceptual difference between images. The lower the value, the closer the reconstructed image is to the real image perceptually.
[0112] As can be seen from Table 1, in the MipNeRF360 dataset, the method proposed in the present invention has performance of 28.20, 0.845, and 0.226 in terms of PSNR, SSIM, and LPIPS respectively, and is particularly outstanding in the optimization of the number of Gaussians, with only 0.49M Gaussians. These results have obvious advantages compared with other comparative methods, especially while maintaining high-quality reconstruction, significantly reducing the number of Gaussians. In the Deep Blending dataset, the PSNR, SSIM, and LPIPS of the present invention are 35.56, 0.939, and 0.203 respectively, and also maintain excellent reconstruction quality when the number of Gaussians is reduced to 0.51M.
[0113] As can be seen from Table 2, after adding the adversarial network, the method proposed in the present invention has increased by 0.74% (from 28.10 to 28.20), 0.25% (from 0.845 to 0.847), and 0.52% (from 0.226 to 0.225) in terms of PSNR, SSIM, and LPIPS respectively. Thus, it can be seen that the adversarial network significantly improves the reconstruction quality of the model, especially in terms of detail retention. After introducing the adaptive splitting module, PSNR and SSIM have increased by 1.12% (from 28.20 to 28.52) and 0.52% (from 0.847 to 0.851) respectively, indicating that the module has a more obvious effect in Gaussian distribution optimization and visual detail enhancement. Importance weighted sampling further improves the performance of the model, enabling PSNR and SSIM to maintain high precision while reducing the number of Gaussians. Combining the results of Table 2, each module in the method proposed in the present invention has brought stable improvement to the model in the ablation experiment, verifying the effectiveness of each module in improving the reconstruction quality and reducing the number of Gaussians.
[0114] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. An efficient 3D reconstruction method based on data-driven compression and adaptive splitting, characterized in that: Execute the following steps S1 to S7 to complete the three-dimensional reconstruction of the target scene: Step S1: Acquire real images of the target scene from different perspectives, generate sparse point clouds through the SfM method, and initialize the sparse point clouds into multiple 3D Gaussians. The attributes of the 3D Gaussians include the center position, size, shape, color, transparency, and learnable features in the three-dimensional space. Step S2: construct a mask generation network and a discriminant network. The mask generation network selects the 3D Gaussian to be eliminated based on the learnable features of the 3D Gaussian, generates a rendering image with some 3D Gaussian eliminated based on data-driven, and outputs a mask; the discriminant network is used to judge the difference between the rendering image with some 3D Gaussian eliminated and the real image; Step S3: construct a differentiable raster renderer to optimize the properties of the 3D Gaussian; Step S4: randomly select a viewing angle, input the real image of the viewing angle into the differentiable grating renderer, render the 3D Gaussian at the viewing angle through the differentiable grating renderer, and generate a complete 3D Gaussian rendering image; Step S5: the mask generation network outputs the importance score of the 3D Gaussian according to the learnable features and perspective information of the 3D Gaussian, and generates a mask according to the threshold; uses the generated mask to eliminate the specific 3D Gaussian; obtains the image after the Gaussian elimination, renders the image after the Gaussian elimination, and generates a rendering image with part of the 3D Gaussian eliminated; Step S6: the discriminant network compares the complete 3D Gaussian rendering, the rendering with some 3D Gaussians eliminated, and the real image; Construct the loss function of the mask generation network, the discriminant network and the scene reconstruction, and use the Adam optimizer to iteratively optimize and update the mask generation network and the discriminant network parameters and the properties of the 3D Gaussian through back propagation; Step S7: Determine whether the preset number of iterations has been reached. If so, proceed to the Gaussian reduction stage. Based on the mask generated by the mask generation network and the importance weighted sampling method, calculate the total importance score of the 3D Gaussian, evaluate the contribution of the 3D Gaussian to the target scene, select the 3D Gaussian with a total importance score greater than the preset value for retention, eliminate other 3D Gaussians, and output the optimized Gaussian representation; If the preset number of iterations is not reached, the Gaussian growth stage is performed to screen out incompletely reconstructed 3D Gaussians, process them according to their scale, adaptively split the 3D Gaussians whose scales are greater than the preset value, and copy the other 3D Gaussians; then return to step S2 to continue iterating until the preset number of iterations is reached, and the three-dimensional reconstruction of the target scene is completed.
2. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The 3D Gaussian in step S1 is defined as follows: ; In the formula, G i ( x ) is a 3D Gaussian, defined as an ellipsoid with specific properties, x Represents each point in the area where the ellipsoid lies, μ i , Σ i are the properties of a 3D Gaussian, where μ i is the center position of the 3D Gaussian in three-dimensional space; Σ i Indicates size The covariance matrix defines the size of the 3D Gaussian and its shape in space; The attributes of 3D Gaussian also include color, transparency, and learnable features, among which the learnable feature F i Using a nine-dimensional first-order tensor express.
3. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The structures and functions of the mask generation network and the discriminant network described in step S2 are as follows: The mask generation network is implemented using tiny-cuda-nn, using ReLU as the activation function. The input of the mask generation network includes the learnable features of 3D Gaussian. , viewing angle information, the mask generation network determines whether each 3D Gaussian should be retained or eliminated according to the learnable features; the mask generation network processes the 3D Gaussian according to the viewing angle information and generates a rendering image that eliminates some 3D Gaussians, and outputs a mask; The discriminant network adopts a convolutional neural network structure, compares the generated rendering image with part of the 3D Gaussian eliminated with the real image, and outputs the discrimination result. The mask generation network gradually optimizes the generated mask by minimizing the feedback of the discriminant network.
4. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The specific process of step S5 is as follows: Step S5.1: The mask generation network uses a two-layer multilayer perceptron, denoted as , the mask generates the network input learnable features and viewing angle information , generating the corresponding 3D Gaussian importance score , the importance score generation process is as follows: ; in, Indicates that the features can be learned F i and viewing angle information θ i Concatenate as input to the mask generation network, I i represents the importance score of 3D Gaussian; Step S5.2: Importance score for each 3D Gaussian I i Used to generate the mask M i , the generation formula is as follows: ; in, is the indicator function, It stops the gradient operation, and threshold represents the threshold value; The threshold is calculated by the following formula: ; in, , , N is the current number of 3D Gaussians, is the proportionality coefficient, is the maximum number of 3D Gaussians; Step S5.3: Generated mask M i It is used to select and process the corresponding 3D Gaussians, and use the threshold threshold to filter the 3D Gaussians. The filtered 3D Gaussians will continue to be rendered through the 3D Gaussian point algorithm to generate a rendering image with some 3D Gaussians eliminated.
5. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The loss function of the mask generation network in step S6 Loss G As follows: ; in, represents the L2 norm of the learnable feature, and σ represents the Sigmod operation; Loss function of the discriminative network Loss D As follows: ; in, Loss R represents the loss of the real image, Loss F Indicates the loss of rendering by eliminating some 3D Gaussians; The discriminative network receives two types of images as input: real images Img R And the rendering with some 3D Gaussians removed Img F ; Loss of real images Loss R And eliminate some of the loss of 3D Gaussian renderings Loss F The calculation formula is as follows: ; ; in, D ( Img R )and D ( Img F ) represent the output of the discriminant network for the real image and the rendering image with some 3D Gaussians eliminated, BCE represents the binary cross entropy loss, and MSE represents the mean square error loss; Loss function for scene reconstruction Loss Restruction As follows: ; in, Img Represents a complete 3D Gaussian rendering. Img R is a real image, λ is a hyperparameter; L1 represents L1 loss, SSIM represents structural similarity loss; Total loss function Loss total As follows: ; in, λ G represents the hyperparameters of the mask generation network, λ D Represents the hyperparameters of the discriminative network.
6. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The Gaussian growth phase in step S7 specifically includes the following processes: Step S7.1.1: Filter out 3D Gaussians whose scale is larger than a preset value, define them as large Gaussians, and use the Scharr edge detection operator to calculate the complete 3D Gaussian rendering Img and real images Img R The gradient of is as follows: ; ; In the formula, G img Represents a complete 3D Gaussian rendering Img The gradient of G ImgR Represents a real image Img R The gradient of Compute gradient differences G d : ; Step S7.1.2: Project the center position of each large Gaussian from the three-dimensional space to the image plane, and calculate the image gradient at the projection point; if the gradient difference value at the projection point of the 3D Gaussian is less than 25% of the maximum gradient difference, the 3D Gaussian is marked as incompletely reconstructed, and the formula is as follows: ; Where, Threshold represents the threshold; Step S7.1.3: Discretize the target scene into voxels using a voxel grid, and calculate the density of the space by calculating the number of 3D Gaussians within each voxel, as follows: ; in, d v It is a voxel v The density of It is a voxel v The number of Gaussians in vol ( v ) is a voxel v Volume; Step S7.1.4: For the large Gaussian that is marked as incompletely reconstructed, the number of Gaussian splits is dynamically adjusted based on the density of the space where it is located. The formula for the number of Gaussian splits is as follows: ; in, N i Indicates i The number of splits of a large Gaussian, d i It is i The density of the space where the large Gaussian is located, min(D) and max(D) are the minimum and maximum spatial densities respectively; Step S7.1.5: Filter out 3D Gaussians whose scale is less than or equal to the preset value, define them as small Gaussians, calculate the gradient of the center position of the small Gaussians through back propagation, determine whether these gradients are less than the preset threshold, and perform a copy operation on the small Gaussians whose gradient values are lower than the preset threshold.
7. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The Gaussian reduction stage in step S7 specifically includes the following processes: Step S7.2.1: At the beginning of the Gaussian reduction phase, every 25 iterations, the mask generation network generates a 3D Gaussian G i Generate its importance score I i ; Step S7.2.2: Importance score for mask generation network output I i , perform binarization and generate a mask M i , the mask is generated every 25 iterations; Step S7.2.3: Perform importance weighted sampling every 1000 iterations, randomly select and retain 95% of the 3D Gaussians each time; calculate the value of each 3D Gaussian at the viewing angle j The blending weights and hit counts on each pixel in; 3D Gaussian G i The mixed weight ω i The calculation of is as follows: ; in, Is the indicator function, used to judge the 3D Gaussian G i Is it consistent with the ray r that passes through the pixel (u,v) from the center of the camera? u,v intersect; σ i For 3D Gaussian G i The opacity, H and W are the height and width of the image respectively; 3D Gaussian G i Hit count c i The calculation of is as follows: ; Calculate 3D Gaussian G i Total importance score S i As follows: ; In the formula, N v Indicates the total number of viewing angles; The total importance score of each 3D Gaussian is normalized as follows: ; In the formula, is the normalized 3D Gaussian G i The total importance score of .
8. The efficient 3D reconstruction method based on data-driven compression and adaptive splitting according to claim 1, characterized in that: The preset number of iterations in step S7 is half of the total number of iterations.
Citation Information
Patent Citations
3D modeling reconstruction system, method and device based on point cloud information and Gaussian cloud cluster
CN118196306A
4D content generation method and device, equipment, medium and computer program product
CN118413715A