Three-dimensional Gaussian splashing new visual angle synthesis method based on human eye perception
By introducing perceptual sensitivity modeling and adaptive density enhancement strategies in the three-dimensional Gaussian splashing technology, the problem of difficult to balance reconstruction quality and efficiency is solved, and efficient and accurate image synthesis of new perspectives is achieved.
Patent Information
- Application Number
- CN202510592346.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing three-dimensional Gaussian splattering technology is difficult to balance between reconstruction quality and efficiency, especially in areas with rich details or complex boundaries, and the existence of parameter redundancy leads to inefficient computing.
By constructing perceptual sensitivity modeling, the distribution of Gaussian primitives is optimized by using human eye perception characteristics, the density enhancement strategy guided by sensitivity and scene adaptive depth reinitialization are adopted to control the number and distribution of Gaussian primitives, reduce redundancy and improve reconstruction efficiency.
It realizes the matching of Gaussian primitive distribution with human eye perception, improves the balance of reconstruction quality and efficiency, reduces model storage and computing overhead, and is suitable for scenarios with limited resources or high real-time requirements.
Smart Images

Figure CN120107447A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of new-viewing angle image generation, and in particular relates to a three-dimensional Gaussian splashing new-viewing angle synthesis method based on human eye perception. Background Art
[0002] New perspective image synthesis technology uses known multi-perspective two-dimensional images to model three-dimensional spatial scenes to generate images from any new perspective. It has long been a research focus in the field of computer vision and has been further promoted with the growing demand for applications such as virtual reality, augmented reality, and digital twins.
[0003] Compared with traditional algorithms that rely on expensive equipment, new perspective image synthesis methods based on deep learning have gradually become the focus of research. Among them, 3D Gaussian Splatting (3DGS) has attracted more and more scholars due to its realistic rendering effect and real-time rendering speed. The algorithm explicitly models the three-dimensional scene into a large number of ellipsoidal three-dimensional Gaussian primitives, optimizes the position, shape, opacity and color of these Gaussian primitives at different perspectives through training, and uses differentiable tile rasterization to render two-dimensional images at specific perspectives.
[0004] Specifically, 3DGS stores a set of parameterized 3D Gaussian primitives, projects them onto the 2D image plane through affine transformation, sorts them according to their distance from the camera, and finally generates the final image using transparency blending. It is expressed as: , (1) in, It is Gaussian primitives, through their center coordinates , covariance matrix , Opacity Spherical harmonic coefficients Characterize its geometric shape and color. Given a Gaussian basis of and , its geometric shape It can be defined as: , (2) in, is a coordinate in 3D space.
[0005] Single Gaussian In perspective The color below Its spherical harmonic coefficients It is calculated that after depth sorting of all Gaussian primitives, the viewing angle Lower Pixel Rendering color Determined by the render function: , (3) , (4) in, and The three-dimensional Gaussian basis element was calculated In perspective Under the oval transparent weights and geometry.
[0006] Unlike traditional neural networks with fixed parameters, 3DGS initializes the position of Gaussian primitives through the point cloud generated by Structure-from-Motion (SfM), and uses adaptive density control to optimize the distribution of Gaussian primitives in local areas. On the one hand, this strategy increases the number of parameters in areas with poor reconstruction effects through two densification operations, cloning and splitting, and on the other hand, removes low-opacity Gaussian primitives through cropping operations to reduce redundancy. The combination of the two operations can control the position distribution of Gaussian primitives and effectively improve the accuracy of reconstruction.
[0007] Existing technology and its shortcomings: Although 3D Gaussian splatting and its subsequent work have made breakthroughs in the task of new-view image synthesis, these methods cannot efficiently distribute Gaussian primitives and still have limitations in reconstruction quality, efficiency, and their balance. First, traditional 3D Gaussian splatting is prone to blurring and artifacts in areas with rich details or complex boundaries, making it difficult to achieve good reconstruction quality in some areas. At the same time, the 3D Gaussian splatting method is prone to parameter redundancy, which increases the model storage overhead and reduces computational efficiency.
[0008] In order to improve the reconstruction quality and efficiency of 3D Gaussian splatting, some improved methods proposed optimization methods of adaptive density control strategies to make the spatial distribution of Gaussian primitives more reasonable.
[0009] For example, in terms of improving reconstruction quality, Pixel-GS uses the number of pixels covered by the Gaussian primitives as the weight for gradient calculation, avoiding ignoring primitives that need to be densified due to the average of gradients from multiple perspectives; Mini-Splatting directly performs additional densification on large-sized primitives, allowing more primitives to represent details. Although these methods can significantly improve reconstruction quality, they usually require the introduction of more Gaussian primitives, which greatly reduces reconstruction efficiency.
[0010] On the contrary, some methods improve reconstruction efficiency by avoiding the densification of useless primitives. For example, color-cued3DGS considers the color gradient when calculating the gradient, which greatly reduces the number of Gaussian primitives. However, although such methods are close to the original methods in terms of PSNR indicators, due to the significant reduction in the number of Gaussian primitives, the quality perceived by the human eye is seriously lost.
[0011] In summary, existing methods have obvious limitations in balancing reconstruction quality and reconstruction efficiency, which limits their application potential in scenarios with limited resources or high real-time requirements. Summary of the invention
[0012] In view of the shortcomings of the existing three-dimensional Gaussian splashing technology, in order to achieve a balance between reconstruction quality and efficiency, the present invention summarizes the following problems and proposes optimization: (1) The problem of mismatch between Gaussian primitive distribution and human eye perception: Most methods based on 3D Gaussian splashing only consider the geometry and color reconstruction effects of the scene during the training process, and do not use the relevant characteristics of the human visual system to optimize the training of the model, resulting in a mismatch between the distribution of Gaussian primitives in the reconstructed scene and human eye perception. The present invention aims to construct a perceptual sensitivity model of the scene, consider the perceptual sensitivity of the human eye to different spatial regions during the training process, construct a more reasonable spatial distribution of Gaussian primitives, and improve the quality and efficiency of reconstruction.
[0013] (2) Densification strategies lack scene adaptability: Existing densification strategies often use the same standard to select Gaussian primitives that need to be densified for different types of scenes. Although this method can improve the reconstruction quality in smaller scenes, when the scale is expanded, this fixed strategy often fails to achieve the expected effect. The present invention aims to use the perceptual sensitivity of each Gaussian primitive obtained through learning to dynamically determine whether each Gaussian primitive needs to be densified, and at the same time use the sensitivity to calculate scene-related characteristics, so as to adaptively select specific densification operations and adjust the distribution of Gaussian primitives, and obtain better robustness in different types of scenes.
[0014] (3) Improving reconstruction quality causes Gaussian primitive redundancy: Since the number of Gaussian primitives is closely related to the quality of scene reconstruction, while promoting densification to improve visual effects, the number of model parameters will also increase. On the one hand, the present invention uses the human eye perception characteristics to limit the densification of Gaussian primitives in non-sensitive areas, and on the other hand, it promotes the cutting of redundant Gaussian primitives through opacity attenuation, thereby improving the reconstruction quality while controlling the model complexity and achieving higher computational efficiency.
[0015] The present invention proposes a new perspective synthesis method of three-dimensional Gaussian splashing based on human eye perception, which integrates multi-perspective human eye perception sensitivity into the training process to optimize the distribution of Gaussian primitives and improve the quality and efficiency of scene reconstruction. Specifically, the multi-perspective sensitivity map is pre-calculated by perceptual sensitivity extraction, and the color and sensitivity loss are used to constrain the training process through dual-branch rendering of color and sensitivity, so that each Gaussian primitive learns the sensitivity of its represented area, and realizes the human eye perception modeling that can learn the scene. On this basis, a perception-based Gaussian primitive distribution strategy is proposed, including sensitivity-guided densification, so that a sufficient number of Gaussian primitives are used to represent perceptually critical and poorly effective areas; and scene-adaptive deep reinitialization, which further improves the performance of the model on the sparse scene of the initial point cloud. In addition, the opacity attenuation mechanism of the cloned primitive proposed in the present invention greatly reduces the storage and computational overhead of the model by prompting the clipping of redundant Gaussian primitives while maintaining the rendering quality.
[0016] Technical Solution A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception includes the following steps: Step 1: Training data processing Input a set of multi-view 2D images of a 3D scene, estimate the camera position of each image through the SfM algorithm, and use perceptual sensitivity extraction to obtain the sensitivity map corresponding to each image, and finally output the training data; Step 2 Initialize 3DGS Use the initial point cloud obtained by the SfM algorithm to initialize the 3DGS Gaussian basis ellipsoid and all parameters of the Gaussian basis ellipsoid; Step 3: Dual-branch rendering Based on the 3DGS model obtained in step 2, the 3D Gaussian primitive is projected to 2D to complete the rendering process. The rendering branch includes two branches: color branch and sensitivity branch. Among them, the color branch renders the scene RGB map (using spherical harmonics to calculate the RGB color of the 3D Gaussian primitive projection), and the sensitivity branch renders the sensitivity map (using the sensitivity parameter to calculate the sensitivity of the 3D Gaussian primitive projection); Calculate the overall loss function of the model to learn the color, geometry, and perceptual modeling of the scene; Step 4 Density Control Density control of the model using adaptive density control and sensitivity-guided densification; Step 5: Scene-adaptive depth reinitialization Whether the initial point cloud of the scene is sparse is determined by learning the sensitivity of the Gaussian primitives, so as to adaptively optimize the position distribution of the Gaussian primitives during the training process to avoid the Gaussian primitives with incorrect initial positions misleading subsequent training.
[0017] Compared with the prior art, the present invention has the following beneficial effects: (1) Gaussian basis element distribution is consistent with the human eye perception characteristics: The learnable human eye perception modeling proposed in the present invention can efficiently capture the sensitivity of the human eye to different areas of the scene, and has the potential to be applied to other tasks related to new perspective image synthesis based on three-dimensional Gaussian splashing, such as model lightweighting.
[0018] (2) Balance between scene reconstruction quality and efficiency: The cloned primitive opacity attenuation mechanism proposed in the present invention can significantly reduce the number of Gaussian primitives required for reconstruction while maintaining the quality of model reconstruction, thus achieving a better balance between quality and efficiency.
[0019] (3) Adaptability and robustness of the densification strategy: This paper proposes a novel perception-based Gaussian primitive distribution strategy, which can robustly improve the reconstruction quality and efficiency even in large-scale urban scenes, and can adaptively adjust the densification strategy according to the scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is an overall flow chart of the present invention; FIG2 is a diagram of dual-branch rendering according to an embodiment of the present invention; FIG3 is a schematic diagram of the process of step 4 of an embodiment of the present invention; FIG4 is a schematic diagram of the process of step 5 of an embodiment of the present invention; FIG5 is a comparison of visualization results of the embodiment of the present invention and the comparative method on the BungeeNeRF dataset; FIG6 is a comparison of visualization results of the embodiment of the present invention and the comparative method on the Mip-NeRF 360 dataset; FIG. 7 is a comparison of the robustness of the embodiment of the present invention and the comparative method in a large-scale urban scene. DETAILED DESCRIPTION
[0021] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0022] A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception, firstly performs training data processing through SfM and perceptual sensitivity extraction before training begins (step 1). During the training process, the invention first defines and initializes a 3DGS model containing additional sensitivity parameters (step 2), and then simultaneously learns the color, geometry and perceptual modeling of the scene through dual-branch rendering (step 3), and constrains it through multi-view RGB images and sensitivity maps. Then, in the qualified training rounds, the invention will perform density control (step 4) and scene adaptive depth reinitialization (step 5), respectively adjusting the density distribution and position distribution of Gaussian primitives, while reducing the cloning of redundant primitives, to achieve a new perspective synthesis with a balance between quality and efficiency.
[0023] A new perspective synthesis method of three-dimensional Gaussian splashing based on human eye perception includes the following steps: (e.g. Figure 1 ) Step 1: Training data processing Input a set of multi-view 2D images of a 3D scene, estimate the camera position of each image through the SfM algorithm, and use perceptual sensitivity extraction to obtain the sensitivity map corresponding to each image, and finally output the training data. The details are as follows: Step 1.1 Camera position estimation Use existing SfM algorithms to handle perspective Two-dimensional image of , get the camera position of this view .
[0024] Step 1.2 Perceptual sensitivity extraction First, the traditional image edge extraction algorithm is used to process the two-dimensional image , obtain the gradient amplitude map, and then enhance and smooth it according to the relevant characteristics of the human visual system, and finally obtain each two-dimensional image Corresponding, easy-to-learn scene sensitivity map , for supervised sensitivity learning of Gaussian primitives.
[0025] Step 1.2.1 Extract image edges For each 2D image , using the Sobel operator to extract local structure. Sobel operator horizontal and vertical gradient convolution kernel and Defined as: , (5) The final edge response map (i.e. gradient magnitude map) Obtained by calculating the color image and the convolution kernel in two directions: , (6) in, Represents a convolution operation.
[0026] Step 1.2.2 Binarization ( Threshold Considering the characteristics of human visual system perception, in order to avoid ignoring the details of the structure whose absolute value of the pixel response value is relatively small but can still be perceived by the human eye, the extracted gradient amplitude map is enhanced so that it only retains the binary information of whether the human eye is sensitive to each pixel: , (7) in, The enhanced sensitivity map is in pixels The pixel value on is the indicator function, , is the enhancement threshold for each pixel. In this embodiment, .
[0027] Step 1.2.3 Average Pooling use Average pooling smoothes the enhanced sensitivity map to simulate the human visual characteristics demonstrated by eye movement experiments.
[0028] Step 1.2.4 Binarization ( Threshold Sensitivity to smoothing Figure 2 Binarization, to obtain the final easy-to-learn binary perceptual sensitivity map, binary threshold , this embodiment sets .
[0029] For each viewing angle, the sensitivity map generated after processing in this step and the original two-dimensional image are provided to step 2 as training data.
[0030] Step 2 Initialize 3DGS The initial point cloud obtained by the SfM algorithm initializes the 3DGS Gaussian basis ellipsoid and all the parameters of the Gaussian basis ellipsoid.
[0031] The present invention is that each Gaussian unit Additional learnable perception parameters , used to render the perceptual sensitivity map, the 3DGS model of the present invention is defined as: , (8) Step 3 Dual-branch rendering (such as Figure 2 ) During the 3DGS rendering process, the 3D Gaussian primitives need to be projected into 2D to achieve efficient rendering.
[0032] Based on the 3DGS model in step 2, the rendering branch of the model of the present invention includes two branches: a color branch and a sensitivity branch.
[0033] Among them, the color branch renders the scene RGB map (using spherical harmonics to calculate the RGB color of the 3D Gaussian primitive projection), and the sensitivity branch renders the sensitivity map (using the sensitivity parameters to calculate the sensitivity of the 3D Gaussian primitive projection).
[0034] Compute the overall loss function of the model to learn the color, geometry, and perceptual modeling of the scene.
[0035] The details are as follows: Step 3.1 Render RGB image Rendering process of color branch: Rendering RGB images using 3DGS model and differentiable tile rasterization.
[0036] Step 3.2 Rendering sensitivity map The sensitivity branch maps the multi-view sensitivity information from the two-dimensional image to the three-dimensional Gaussian primitives through learning, and constrains the sensitivity of each primitive to be uniform over the spatial range covered. The sensitivity map of the scene can be rendered through the sensitivity parameters of each primitive and differentiable tile rasterization: , (9) in, The sensitivity map for rendering in pixels The value of is the sigmoid function.
[0037] Step 3.3 Calculate the loss During the optimization process, the present invention utilizes the losses of the two rendering branches, color and sensitivity, to jointly supervise model training.
[0038] For the color branch, the same as the original three-dimensional Gaussian splash is adopted The loss function is obtained by weighted summation of L1 loss and structural consistency loss.
[0039] For the sensitivity branch, binary cross entropy loss is used Supervision, so that the sensitivity map of the rendering And the corresponding true sensitivity map Consistent. Sensitivity loss The specific definition is: , (10) Finally, the overall loss function of the model is The weighted sum of color loss and sensitivity loss is obtained: , (11) in, is the weight of sensitivity loss.
[0040] Step 3.4 Update 3DGS parameters After calculating the loss in each round of training, the model parameters are updated by calculating the gradient through back propagation.
[0041] Step 4 Density control (the overall process is as follows Figure 3 shown) The density control process was performed every 100 epochs from 500 to 15,000 training iterations.
[0042] The present invention improves the densification operation of the standard 3DGS and improves the quality and efficiency of scene reconstruction by optimizing the density distribution of Gaussian primitives. The details are as follows: Step 4.1 Adaptive density control The processing of 3DGS adaptive density control includes the splitting, cloning and cutting of primitives. The present invention optimizes the cloning operation and prompts more redundant primitives to be cut to improve the reconstruction efficiency, as follows: Step 4.1.1 Screening large position gradient and large size primitives The same method as standard 3DGS is used to screen out primitives that need to be split using gradient and size thresholds.
[0043] Step 4.1.2 Split Perform splitting operations on the primitives selected in step 4.1.1 using the same method as standard 3DGS.
[0044] Step 4.1.3 Screening large position gradient small size primitives The same method as standard 3DGS was used to screen out the motifs that need to be cloned using gradient and size thresholds.
[0045] Step 4.1.4 Opacity attenuation of the primitive to be cloned Before the Gaussian primitives are cloned, the present invention introduces an opacity attenuation mechanism, the purpose of which is to promote the deletion of redundant Gaussian primitives during the clipping process by reducing the opacity of the primitives to be cloned, so as to avoid negative impact on the training process.
[0046] Specifically, the opacity of the primitive to be cloned is determined by the opacity attenuation function Performs a non-linear transformation to achieve the effect of opacity attenuation.
[0047] The principle is explained as follows: According to the transparency blending, when the two opacities are When the Gaussian primitives overlap, the opacity of the corresponding spatial area It can be expressed as: , (12) Assume that before the cloning operation, the opacity of the spatial region represented by a Gaussian primitive is The goal of the opacity attenuation mechanism is to reduce the opacity of the area after the cloning operation is completed. Specifically, the opacity of the cloned primitive will be reduced by the opacity attenuation function Performs a non-linear transformation to achieve the effect of opacity attenuation.
[0048] Finally, the opacity of the two cloned Gauss primitives It can be obtained by solving the following equation: , (13) It can be solved , that is, adjust the opacity of the primitive to be cloned to .
[0049] Opacity Falloff Function The choice of must satisfy the following requirements: a larger attenuation is performed on the smaller opacity to encourage the Gaussian primitive to be clipped; and a small reduction is performed on the higher opacity to avoid removing important Gaussian primitives, which would cause a cliff drop in model performance.
[0050] Specific, requirements exist The scope meets the following conditions: (1) , indicating that the transformed value is not greater than the original value; (2) , indicating that the transformed value is still within the scope; (3) ,in For function The stationary point of , meaning that a smaller opacity will get a larger falloff.
[0051] As an example, a power function is used As .
[0052] To determine The best value of various indexes was tested. Finally, this embodiment preferably , which can balance reconstruction quality and efficiency.
[0053] Step 4.1.5 Cloning Perform cloning operations on the primitives filtered out in step 4.1.3 using the same method as standard 3DGS.
[0054] Step 4.1.6 Gaussian basis element clipping The same approach as for standard 3DGS is used, removing primitives whose opacity is below a threshold of 0.005.
[0055] Step 4.2 Sensitivity-guided densification On the basis of step 4.1, the present invention utilizes the scene perception sensitivity learned by each primitive to guide the spatial distribution of Gaussian primitives in different regions, so as to promote more Gaussian primitives to be distributed in regions that require more accurate reconstruction.
[0056] Step 4.2.1 Screening high-sensitivity Gaussian primitives Since the real sensitivity map to be learned is binary, the sensitivity of the ideal Gaussian primitives should be close to 0 or 1, where primitives with a sensitivity close to 1 represent the sensitive areas of the human eye in the scene, and primitives close to 0 represent the insensitive areas. Specifically, a high-sensitivity Gaussian primitive Can pass the threshold Filter, you can choose a larger value, this embodiment sets : , (14) This step is repeated every 500-15000 training iterations. The round is executed once, and when it is not executed is an empty set. In this embodiment, , to balance reconstruction quality and efficiency.
[0057] Step 4.2.2 Screening of medium-sensitivity high basis units For high primitives with large differences in sensitivity under different viewing angles, the training process easily makes the middle values of the sensitivity areas of these primitives average the differences between different viewing angles. Therefore, it is impossible to represent such complex areas using only a single primitive, and these primitives need to be split into smaller units.
[0058] Using Thresholds and Screening medium sensitivity Gaussian : , (15) This step is repeated every 500-15000 training iterations. The round is executed once, and when it is not executed is an empty set. In this embodiment, , to balance reconstruction quality and efficiency.
[0059] Step 4.2.3 Screening high-weight primitives In order to prevent too many Gaussian primitives from appearing inside the object, the present invention imposes weight restrictions on the Gaussian primitives selected according to sensitivity. Defined as: , (16) , (17) , (18) , (19) in, Calculated primitives In perspective Pixels The weight on and Respectively represent different weight thresholds for high sensitivity and medium sensitivity Gaussian primitives, It represents the primitives that need to be densified, which are selected from the high-sensitivity Gaussian primitives. It indicates the primitives that need to be densified, which are selected from the medium-sensitivity high primitives; The largest element in the set is selected. Representation perspective All pixels in For all perspectives in the scene Middle-Gaussian The maximum weight of .
[0060] Step 4.2.4 Calculate scene sensitivity The present invention determines the specific densification operation of the selected primitives by scene sensitivity. Specifically, the scene sensitivity Can be defined as all viewing angles Average pixel sensitivity: , (20) , (twenty one) in Pixel The sensitivity of Representation perspective The average sensitivity.
[0061] Step 4.2.5 Densification of high-sensitivity Gaussian primitives The original three-dimensional Gaussian splashing performs splitting and cloning operations according to the size of the Gaussian primitives during the densification process. In order to better capture the details of the scene, the present invention only uses the scene sensitivity Below threshold When, yes The cloning operation described in step 4.1.4 and step 4.1.5 and the splitting operation described in step 4.1.2 are simultaneously applied. Instead, only the selected Gauss primitives are split, regardless of their size, to use more primitives of smaller size to represent complex details.
[0062] Step 4.2.6 Densification of medium-sensitivity Gaussian primitives This step is Densify using the same splitting operation as in step 4.1.2.
[0063] Step 5: Scene adaptive depth reinitialization (the overall process is as follows Figure 4 shown) This step determines whether the initial point cloud of the scene is sparse by learning the sensitivity of the Gaussian primitives, so as to adaptively optimize the position distribution of the Gaussian primitives during the training process to avoid the Gaussian primitives with incorrect initial positions misleading subsequent training.
[0064] Step 5.1 Calculate the proportion of large primitives with medium sensitivity For scenes where the initial point cloud is too sparse, a large-sized Gaussian is usually used to represent multiple regions with different spatial sensitivities, resulting in poor sensitivity learning of the primitive. Using this feature, the present invention uses a ratio of medium-sensitivity Gaussian primitives in large-sized Gaussian primitives after 600 rounds of warm-up training. , as an indicator to judge whether the initial point cloud is sparse: , (twenty two) in, , (twenty three) in for Scaling of the longest axis, represents the set of all Gaussian primitives with their longest axis scaled, Represents the third quartile, used to identify the top 25% of the largest Gaussian primitives .
[0065] like Exceeding a predefined threshold , then the deep reinitialization operation is performed on this scene in subsequent training, otherwise it is not performed.
[0066] Specifically, when the number of training rounds is 5000, 10000, and 15000, the present invention adopts a scene-adaptive depth reinitialization strategy, which is not executed in other rounds.
[0067] Step 5.2 Deep reinitialization This step is consistent with the depth reinitialization step in Mini-Splatting. This step finds the Gaussian primitive with the largest weight for each pixel, and takes the depth of the midpoint of the two intersections of the ray from the camera to the pixel and the Gaussian primitive ellipsoid with the largest weight as the depth of the pixel, thus obtaining the depth-color pair of each pixel. Then, 35,000 non-repeating depth-color pairs are randomly sampled as new 3D points to reinitialize the Gaussian primitive.
[0068] Example Dataset The method proposed in the present invention is tested on 21 scenes of Mip-NeRF 360, Tanks&Temples, Deep Blending and BungeeNeRF datasets, covering various types of scenes such as indoor, outdoor and large-scale cities.
[0069] Among them, the Mip-NeRF 360 dataset is designed for borderless 360-degree scenes. It contains 9 complex indoor and outdoor scenes (such as streets, buildings, and natural landscapes) and supports dynamic perspective changes.
[0070] Tanks&Temples is a large-scale real-scene dataset that contains high-resolution indoor and outdoor videos and laser scanning data. It is used to evaluate the reconstruction accuracy and generalization ability of NeRF-type models. It includes 14 complex scenes (such as churches, castles, and forests) covering different lighting conditions and dynamic elements (such as pedestrians and vehicles). The Deep Blending dataset contains 19 complex scenes (such as indoors, streets, and natural landscapes), each of which provides 12 to hundreds of images.
[0071] BungeeNeRF is a model and supporting dataset designed specifically for extreme multi-scale urban scenes. It supports seamless rendering from satellite to ground perspectives, includes 12 global city landmarks (such as New York and Tokyo), and thousands of multi-resolution images for each scene. Method application All experiments in this example are performed on a single NVIDIA RTX 4090 GPU. In order to achieve a better balance between quality and efficiency, different weight thresholds are used for high-sensitivity and medium-sensitivity Gaussian primitives, marked as and The specific values of the hyperparameters in this embodiment are shown in Table 1.
[0072] Table 1 Value settings of various hyperparameters in this embodiment Comparison of implementation effects The comparison methods of this embodiment include: Pixel-GS, Mini-Splatting-D, Taming-3DGS and 3DGS. Among them, Pixel-GS is a density control optimization method for 3D Gaussian splatting, which solves the blurring and needle-like artifact problems in sparse point cloud areas through pixel-aware gradients and gradient field scaling; Mini-Splatting-D achieves ultra-fast optimization through radical Gaussian densification and visibility culling; Taming-3DGS proposes a budget-controlled Gaussian optimization framework to solve efficiency problems in resource-constrained scenarios. The first three methods all optimize the densification strategy, while the last method is the original three-dimensional Gaussian splatting method.
[0073] The indicators used are: SSIM (Structural Similarity Index), LPIPS (Learned Perceptual Image Patch Similarity), number of Gaussian primitives (#G), and performance comparison results on quality-efficiency balance (QEB). Among them, QEB is an indicator that comprehensively considers reconstruction quality and efficiency, and is defined as: , (twenty four) Table 2 Comparison experimental results of reconstruction quality between the present invention and other models Table 2 shows the quantitative comparison results of the reconstruction quality of the present invention and the comparative methods on different data sets. In the four data sets, the present invention performs well in SSIM and perception-related LPIPS indicators. Unlike Pixel-GS and Taming-3DGS, which introduce too many Gaussian primitives in large-scale scenes and cause CUDA memory overflow (OOM) problems, the present invention adaptively allocates primitives according to the perceptual sensitivity of different regions, achieving a better quality-efficiency trade-off.
[0074] Table 3 Comparison experimental results of reconstruction efficiency of the present invention and other models Table 3 provides a quantitative analysis of the model complexity and the trade-off between quality and efficiency of the present invention and the comparative methods on different datasets. Compared with other quality-focused methods, the present invention shows significant improvements in efficiency and quality-efficiency balance, rendering realistic new perspective images at a faster speed and with fewer Gaussian primitives.
[0075] The effect of the processing is shown in the figure Figure 5 , Figure 6 and Figure 7As shown in FIG. 1 , the present invention allocates more Gaussian primitives to object details and edges, effectively reducing the blurriness of the scene. Figure 5 As shown in FIG. 1 , the present invention accurately reconstructs the roads, buildings and grass at the scene boundary, avoiding the distortion phenomenon that occurs in other methods. Figure 6 In the images, the proposed method reconstructs the ground texture more realistically, while other methods tend to produce more artifacts in these areas.
[0076] like Figure 7 The robustness of the present invention and the comparative methods (Pixel-GS and Mini-Splatting-D) on large-scale urban scenes is compared. The numbers in brackets in the figure represent the LPIPS index and the number of Gaussian primitives (in millions, M). For example, the index (0.116, 5.03) of the present invention indicates that the LPIPS index of the present invention on this image is 0.116, and the number of Gaussian primitives is 5.03M. The method of the present invention shows extremely high robustness in large-scale scenes, avoiding the introduction of too many Gaussian primitives like Pixel-GS and the reconstruction failure in Mini-Splatting-D. This is due to the dual-branch rendering and sensitivity-guided densification of the present invention, which limit the number of densified Gaussian primitives while adaptively identifying areas that require more primitives.
[0077] Key points and advantages of the present invention Learnable human eye perception modeling: Before the training begins, the present invention operates on the multi-view color image of the scene through perceptual sensitivity extraction, and pre-calculates the binary sensitivity map of the three-dimensional scene for supervising the training of the sensitivity parameters of the Gaussian model. During the training process, the sensitivity parameters of each Gaussian primitive are rendered and optimized in a similar manner to its color and opacity parameters, and finally the human eye perception modeling of the scene can be learned. Since the sensitivity map used for training is binary, compared with directly learning a continuous gradient amplitude map, the Gaussian primitive is easier to capture the spatial perception sensitivity and avoid the performance degradation caused by poor learning. In addition, compared with the statistical three-dimensional Gaussian primitive in the two-dimensional projection of multiple perspectives to cover the sensitivity pixel value and to judge the sensitivity of the Gaussian primitive, the learnable human eye perception modeling proposed by the present invention can promote the spatial perception sensitivity represented by the same Gaussian primitive to be similar, which is conducive to avoiding the densification of Gaussian primitives in non-sensitive areas, and more accurately responding to sensitive areas or poorly learned Gaussian primitives for subsequent additional densification operations.
[0078] Perception-based Gaussian primitive distribution strategy: Compared with the relatively fixed methods of existing work, the densification strategy proposed in the present invention shows good robustness in different types of scenes, and can adaptively adjust the strategy according to the characteristics of the scene to achieve better reconstruction effects. Since the splitting strategy can increase the number of small-sized Gaussian primitives in the scene, for scenes containing a large number of sensitive areas, splitting operations only in sensitive areas can accurately represent scene details, thereby improving reconstruction quality. For scenes with sparse initial point clouds, after 600 rounds of warm-up training, there will be more large-sized Gaussians with medium sensitivity, because the sparse initial point cloud makes it possible for spatial regions of different sensitivities to be represented by only a small number of primitives and cannot be accurately reconstructed. These primitives often contain incorrect spatial distributions. Although blindly densifying can improve the overall reconstruction quality to a certain extent, it will cause subsequent training to continuously distribute new Gaussian primitives in the wrong positions, hindering further improvement of model performance. For this type of scene, selectively introducing a deep reinitialization strategy to timely adjust the positions of erroneous primitives in the position distribution can effectively avoid misleading subsequent training. At the same time, it will not cause the original well-reconstructed areas to be destroyed due to random sampling, and can show good robustness in different scenes.
[0079] Opacity attenuation mechanism of cloned primitives: The cloning operation of the original three-dimensional Gaussian splash will copy a Gaussian with all parameters completely consistent with the cloned primitive, which will cause the opacity of the spatial area represented by it to increase, making the area that was originally poorly learned play a more important role in rendering, hindering model training. The present invention promotes the clipping of poorly learned redundant primitives by nonlinearly reducing the opacity of primitives with different opacities after the cloning operation, avoiding excessive cloning of redundant primitives, resulting in reduced storage and computing efficiency. Specifically, for Gaussian primitives with lower original opacity, their opacity is reduced more significantly after cloning, because transparent poorly learned Gaussian primitives have less impact on the rendering results and are more likely to be redundant primitives. On the contrary, for Gaussian primitives with higher original opacity, their opacity is reduced to a smaller extent after cloning to further determine whether they are redundant in the subsequent training process, avoiding the rapid deletion of important Gaussian primitives causing a cliff-like drop in model quality.
[0080] The above description is only a description of the preferred embodiments of the present application, and is not intended to limit the scope of the present application. Any changes or modifications made by any person skilled in the art based on the above disclosed technical contents shall be deemed as equivalent effective embodiments and shall fall within the scope of protection of the technical solution of the present application.
Claims
1. A new perspective synthesis method of three-dimensional Gaussian splashing based on human eye perception, characterized in that: The following steps are involved: Step 1: Training data processing Input a set of multi-view 2D images of a 3D scene, estimate the camera position of each image through the SfM algorithm, and use perceptual sensitivity extraction to obtain the sensitivity map corresponding to each image, and finally output the training data; Step 2 Initialize 3DGS Use the initial point cloud obtained by the SfM algorithm to initialize the 3DGS Gaussian basis ellipsoid and all parameters of the Gaussian basis ellipsoid; Step 3: Dual-branch rendering Based on the 3DGS model obtained in step 2, the 3D Gaussian primitive is projected to 2D to complete the rendering process. The rendering branch includes two branches: the color branch and the sensitivity branch. The color branch renders the scene RGB map, and the sensitivity branch renders the sensitivity map. Calculate the overall loss function of the model to learn the color, geometry, and perceptual modeling of the scene; Step 4 Density Control Density control of the model using adaptive density control and sensitivity-guided densification; Step 5: Scene-adaptive depth reinitialization Whether the initial point cloud of the scene is sparse is determined by learning the sensitivity of the Gaussian primitives, so as to adaptively optimize the position distribution of the Gaussian primitives during the training process to avoid the Gaussian primitives with incorrect initial positions misleading subsequent training.
2. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 1, characterized in that: Step 1 is as follows: Step 1.1 Camera position estimation Use SFM algorithm to process perspective Two-dimensional image of , get the camera position of this view ; Step 1.2 Perceptual sensitivity extraction First, the image edge extraction algorithm is used to process the two-dimensional image , obtain the gradient amplitude map, enhance and smooth it, and finally obtain each two-dimensional image Corresponding sensitivity map , for supervised sensitivity learning of Gaussian primitives; Step 1.2.1 Extract image edges For each 2D image , using the Sobel operator to extract local structure, Sobel operator horizontal and vertical gradient convolution kernel and Defined as: , (5) Final edge response map Obtained by calculating the color image and the convolution kernel in two directions: , (6) in, Represents the convolution operation; Step 1.2.2 Binarization The extracted gradient magnitude map is enhanced so that it only retains the binary information of whether the human eye is sensitive to each pixel: , (7) in, The enhanced sensitivity map is in pixels The pixel value on is the indicator function, , is the enhancement threshold for each pixel; Step 1.2.3 Average Pooling use Average pooling smoothes the enhanced sensitivity map to simulate human visual characteristics; Step 1.2.4 Binarization Binarize the smoothed sensitivity map to obtain the final easy-to-learn binary sensitivity map and the binary threshold ; For each viewing angle, the sensitivity map generated after processing in this step and the original two-dimensional image are provided to step 2 as training data.
3. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 2, characterized in that: set up , .
4. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 1, characterized in that: In step 2, for each Gaussian basis element Added learnable perception parameters , used to render the sensitivity map, the 3DGS model is defined as: , (8)。 5. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 1, characterized in that: Step 3 is as follows: Step 3.1 Render RGB image Render RGB images using the 3DGS model and differentiable tile rasterization; Step 3.2 Rendering sensitivity map The sensitivity branch maps the multi-view sensitivity information from the two-dimensional image to the three-dimensional Gaussian primitives through learning, and constrains the sensitivity of each primitive to be uniform across the spatial range. The sensitivity map of the scene is rendered through the sensitivity parameters of each primitive and differentiable tile rasterization: , (9) in, The sensitivity map for rendering is in pixels The value of is the sigmoid function, The three-dimensional Gaussian basis In perspective The weight of the lower ellipse is transparent, are learnable perception parameters; Step 3.3 Calculate the loss During the optimization process, the losses of the two rendering branches, color and sensitivity, are used to jointly supervise the model training; For the color branch, the same as the original three-dimensional Gaussian splash is adopted The loss function is obtained by weighted summation of L1 loss and structural consistency loss; For the sensitivity branch, binary cross entropy loss is used Supervision, so that the sensitivity map of the rendering And the corresponding true sensitivity map Consistency; loss of sensitivity The specific definition is: , (10) Finally, the overall loss function of the model is The weighted sum of color loss and sensitivity loss is obtained: , (11) in, is the weight of sensitivity loss; Step 3.4 Update 3DGS parameters After calculating the loss in each round of training, the model parameters are updated by calculating the gradient through back propagation.
6. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 1, characterized in that: Step 4 is as follows: During the training iterations from 500 to 15,000, the density control process was performed every 100 rounds. Step 4.1 Adaptive density control The processing of 3DGS adaptive density control includes the splitting, cloning and clipping of primitives, as follows: Step 4.1.1 Screening large position gradient and large size primitives Use gradient and size thresholds to select primitives that need to be split; Step 4.1.2 Split Perform splitting operation on the primitives filtered out in step 4.1.1; Step 4.1.3 Screening large position gradient small size primitives Use gradient and size thresholds to filter out primitives that need to be cloned; Step 4.1.4 Opacity attenuation of the primitive to be cloned Before the Gaussian primitive is cloned, the opacity of the primitive to be cloned is reduced by the opacity attenuation function. Perform nonlinear transformation to achieve the effect of opacity attenuation; Step 4.1.5 Cloning Perform cloning operation on the primitives selected in step 4.1.3; Step 4.1.6 Gaussian basis element pruning Remove primitives with opacity below a threshold of 0.005; Step 4.2 Sensitivity-guided densification Based on step 4.1, the scene perception sensitivity learned by each primitive is used to guide the spatial distribution of Gaussian primitives in different regions, so that more Gaussian primitives are distributed in the regions that need more accurate reconstruction. Step 4.2.1 Screening high-sensitivity Gaussian primitives In 500-15000 training iterations, The round performs a high sensitivity Gaussian primitive screening operation once; Pass Threshold Screening for high sensitivity Gaussian : , (14) When no filtering is performed is an empty set; Step 4.2.2 Screening of medium-sensitivity high basis units In 500-15000 training iterations, The round performs a screening of high sensitivity base element operations; Using Thresholds and Screening medium sensitivity Gaussian : , (15) When no filtering is performed is an empty set; Step 4.2.3 Screening high-weight primitives Apply weight restrictions to the Gaussian primitives selected based on sensitivity, and additional densification of Gaussian primitives is required Defined as: , (16) , (17) , (18) , (19) in, Calculated primitives In perspective Pixels The weight on and Respectively represent different weight thresholds for high sensitivity and medium sensitivity Gaussian primitives, It represents the primitives that need to be densified, which are selected from the high-sensitivity Gaussian primitives. It indicates the primitives that need to be densified, which are selected from the medium-sensitivity high primitives; The largest element in the set is selected. Representation perspective All pixels in For all perspectives in the scene Middle-Gaussian The maximum weight of Step 4.2.4 Calculate scene sensitivity The specific densification operation of the primitives selected by scene sensitivity judgment; scene sensitivity Defined as all viewpoints Average pixel sensitivity: , (20) , (21) in Pixel The sensitivity of Representation perspective The average sensitivity of Step 4.2.5 Densification of high-sensitivity Gaussian primitives When scene sensitivity Below threshold When, yes Simultaneously use the cloning operation described in step 4.1.4 and step 4.1.5 and the splitting operation described in step 4.1.2; conversely, only split the selected Gaussian primitives; Step 4.2.6 Densification of medium-sensitivity Gaussian primitives right Densify using the same splitting operation as in step 4.1.
2.
7. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 6, characterized in that: Using power function As ,in .
8. A three-dimensional Gaussian splash new perspective synthesis method based on human eye perception as claimed in claim 6, characterized in that: Step 5 is as follows: Step 5.1 Calculate the proportion of medium-sensitivity large primitives The ratio of medium-sensitivity Gaussian primitives in large-size Gaussian primitives after 600 rounds of warm-up training , as an indicator to judge whether the initial point cloud is sparse: , (22) in, , (23) in for Scaling of the longest axis, represents the set of all Gaussian primitives with their longest axis scaled, Represents the third quartile, used to identify the top 25% of the largest Gaussian primitives ; like Exceeding a predefined threshold , then the deep re-initialization operation is performed on the scene in subsequent training, otherwise it is not performed; Step 5.2 Deep reinitialization This step is consistent with the depth reinitialization step in Mini-Splatting; this step finds the Gaussian primitive with the largest weight for each pixel, and takes the depth of the midpoint of the two intersections of the ray from the camera to the pixel and the Gaussian primitive ellipsoid with the largest weight as the depth of the pixel, thus obtaining the depth-color pair of each pixel; then randomly sample 35,000 non-repeating depth-color pairs as new three-dimensional points to reinitialize the Gaussian primitive.
Citation Information
Patent Citations
Three-dimensional human body generation method based on text prompt
CN118229860A
Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion
CN118351252A
Three-dimensional object generation method and device, equipment and storage medium
CN118429531A
Face high-fidelity and drivable reconstruction method based on three-dimensional Gaussian splashing
CN118736108A
Three-dimensional temperature field reconstruction method based on 3D Gaussian splashing
CN119379925A
Cited By
Gradient control-based Gaussian rendering image degradation processing method
CN120580335A