A New Perspective Synthesis Method of 3D Gaussian Splashing Based on Human Eye Perception

Through human eye perception sensitivity modeling and adaptive density control, the Gaussian primitive distribution is optimized, and the problem of unbalanced reconstruction quality and efficiency in the three-dimensional Gaussian splashing method is solved, and efficient new perspective image synthesis is achieved.

CN120107447BActive Publication Date: 2025-07-18TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510592346.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-18
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing three-dimensional Gaussian splashing method has limitations in the balance of reconstruction quality and efficiency in new perspective image synthesis. The distribution of Gaussian primitives does not match the perception of the human eye, the density enhancement strategy lacks scene adaptability, and the redundancy of the number of Gaussian primitives leads to inefficient computing efficiency.

Method used

By introducing human eye perception sensitivity modeling, multi-view sensitivity maps are used for training data processing, dual-branch rendering and adaptive density control are adopted, combined with cloned primitive opacity attenuation mechanism, Gaussky distribution and position adjustment are optimized, and a scene-adaptive density enhancement strategy is realized.

Benefits of technology

It improves the balance between reconstruction quality and efficiency, reduces the number of Gaussian primitives, improves the robustness and computing efficiency of the model, and is suitable for new perspective image synthesis in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107447B_ABST
    Figure CN120107447B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of novel view image generation, and proposes a three-dimensional Gaussian splash novel view synthesis method based on human eye perception. The method includes the following steps: Step 1, training data processing; Step 2, initializing 3DGS; Step 3, dual-branch rendering; Step 4, density control; Step 5, scene-adaptive depth re-initialization. The Gaussian basis element distribution obtained by the method of the present invention conforms to the characteristics of human eye perception; the proposed opacity attenuation mechanism of the cloned basis element can significantly reduce the number of Gaussian basis elements required for reconstruction on the basis of maintaining the model reconstruction quality, achieving a better balance between quality and efficiency; the perception-based Gaussian basis element distribution strategy of the present invention can robustly improve the reconstruction quality and efficiency even in large-scale urban scenes, and can adaptively adjust the densification strategy according to the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of novel view image generation, and particularly relates to a novel view synthesis method of three-dimensional Gaussian splatting based on human eye perception. Background Art

[0002] The novel view image synthesis technology uses known multi-view two-dimensional images to model a three-dimensional space scene to generate images from any novel view. It has long been a research focus in the field of computer vision and has been further promoted with the growth of application requirements such as virtual reality, augmented reality, and digital twins.

[0003] Compared with traditional algorithms that rely on expensive equipment, the novel view image synthesis method based on deep learning has gradually become the focus of research. Among them, three-dimensional Gaussian splatting (3DGS) has attracted more and more scholars' research due to its realistic rendering effect and real-time rendering speed. This algorithm explicitly models the three-dimensional scene as a large number of ellipsoidal three-dimensional Gaussian basis elements, and optimizes the positions, shapes, opacities, and colors under different views of these basis elements through training, and uses differentiable tile rasterization to render two-dimensional images from a specific view.

[0004] Specifically, 3DGS stores a series of parameterized three-dimensional Gaussian basis elements, projects them onto the two-dimensional image plane through affine transformation, sorts them according to their distances to the camera, and finally uses transparency blending to generate the final image. In a scene, the set of Gaussian basis elements is expressed as:

[0005] , (1)

[0006] where is the th Gaussian basis element, and its geometric shape and color are characterized by its center coordinates , covariance matrix , opacity and spherical harmonic coefficients . Given the and of the Gaussian basis element , its geometric shape can be defined as:

[0007] , (2)

[0008] where is a coordinate in 3D space.

[0009] The color of a single Gaussian basis element under the view ​ can be calculated through its spherical harmonic coefficients After performing a depth sort on all Gaussian basis elements, the perspective of the pixel The rendering color is determined by the rendering function:

[0010] , (3)

[0011] , (4)

[0012] where and respectively calculate the weights and geometries of the three-dimensional Gaussian basis elements at the perspective with elliptical transparency

[0013] Different from traditional neural networks with fixed numbers of parameters, 3DGS initializes the positions of Gaussian basis elements through the point cloud generated by Structure-from-Motion (SfM) and adopts adaptive density control to optimize the distribution of Gaussian basis elements in local regions. On the one hand, this strategy increases the number of parameters in regions with poor reconstruction effects through two densification operations, cloning and splitting. On the other hand, it removes Gaussian basis elements with low opacity through cropping operations to reduce redundancy. The combination of the two operations can control the position distribution of Gaussian basis elements and effectively improve the accuracy of reconstruction.

[0014] Existing technologies and their disadvantages:

[0015] Although three-dimensional Gaussian splatting and its subsequent works have made breakthroughs in the new view image synthesis task, these methods cannot distribute Gaussian basis elements efficiently and still have limitations in reconstruction quality, efficiency, and their balance. First, traditional three-dimensional Gaussian splatting is prone to blurring and artifacts in regions with rich details or complex boundaries, and it is difficult to achieve good reconstruction quality in some regions. At the same time, the three-dimensional Gaussian splatting method is prone to parameter redundancy, which reduces the computational efficiency while increasing the model storage overhead.

[0016] To improve the reconstruction quality and efficiency of three-dimensional Gaussian splatting, some improved methods have proposed optimization methods for the adaptive density control strategy to make the spatial distribution of Gaussian basis elements more reasonable.

[0017] For example, from the perspective of improving the reconstruction quality, Pixel-GS uses the number of pixels covered by Gaussian basis elements as the weight for gradient calculation, avoiding the neglect of basis elements that need to be densified due to the averaging of gradients from multiple perspectives; Mini-Splatting directly densifies large-sized basis elements additionally, promoting more basis elements to represent details. Although these methods can significantly improve the reconstruction quality, they usually require introducing more Gaussian basis elements, resulting in a substantial reduction in the reconstruction efficiency.

[0018] On the contrary, some methods achieve an improvement in the reconstruction efficiency by avoiding densifying useless basis elements. For example, color-cued3DGS considers color gradients when calculating gradients, significantly reducing the number of Gaussian basis elements. However, although such methods are close to the original methods in terms of the PSNR metric, due to the substantial reduction in the number of Gaussian basis elements, the loss of human eye perception quality is serious.

[0019] In summary, there are obvious limitations in the balance between the reconstruction quality and the reconstruction efficiency of existing methods, restricting their application potential in scenarios with limited resources or high real-time requirements. Summary of the Invention

[0020] Aiming at the deficiencies of the existing three-dimensional Gaussian splatting technology and in order to achieve a balance between the reconstruction quality and the efficiency, the present invention summarizes the following problems and proposes optimizations:

[0021] (1) The problem of mismatch between the distribution of Gaussian basis elements and human eye perception: Most methods based on three-dimensional Gaussian splatting only consider the geometric and color reconstruction effects of the scene during the training process, without using the relevant characteristics of the human visual system to optimize the training of the model, resulting in a mismatch between the distribution of Gaussian basis elements in the reconstructed scene and human eye perception. The present invention aims to construct a perceptual sensitivity model of the scene, consider the perceptual sensitivity of the human eye to different spatial regions during the training process, construct a more reasonable spatial distribution of Gaussian basis elements, and improve the reconstruction quality and efficiency.

[0022] (2) The problem of lack of scene adaptability in the densification strategy: Existing densification strategies often use the same criteria to select Gaussian basis elements that need to be densified for different types of scenes. Although this method can improve the reconstruction quality in scenes with a small scale, when the scale expands, this fixed strategy often fails to achieve the expected effect. The present invention aims to dynamically determine whether each Gaussian basis element needs to be densified by using the perceptual sensitivity learned by each Gaussian basis element, and at the same time use this sensitivity to calculate scene-related characteristics, so as to adaptively select specific densification operations and adjust the distribution of Gaussian basis elements, and obtain better robustness in different types of scenes.

[0023] (3) High redundancy problem of Gaussian basis elements caused by improving reconstruction quality: Since the number of Gaussian basis elements is closely related to the quality of scene reconstruction, while promoting densification to improve the visual effect, the number of model parameters will also increase accordingly. On the one hand, the present invention uses the human eye perception characteristics to limit the densification of Gaussian basis elements in insensitive regions, and on the other hand, promotes the cropping of redundant Gaussian basis elements through opacity attenuation, controls the model complexity while improving the reconstruction quality, and achieves higher computational efficiency.

[0024] The present invention proposes a new perspective synthesis method for three-dimensional Gaussian splash based on human eye perception, integrating multi-view human eye perception sensitivity into the training process to optimize the distribution of Gaussian basis elements and improve the quality and efficiency of scene reconstruction. Specifically, first, pre-compute multi-view sensitivity maps through perception sensitivity extraction, and use color and sensitivity losses for constraint during the training process through dual-branch rendering of color and sensitivity, enabling each Gaussian basis element to learn the sensitivity of its representation area and realizing learnable human eye perception modeling of the scene. On this basis, a perception-based Gaussian basis element distribution strategy is proposed, including sensitivity-guided densification, enabling a sufficient number of Gaussian basis elements to represent perception-critical and poorly performing regions; and scene-adaptive depth re-initialization, further improving the performance of the model in sparse initial point cloud scenes. In addition, the opacity attenuation mechanism of the cloned basis elements proposed in the present invention significantly reduces the storage and computational overhead of the model by promoting the cropping of redundant Gaussian basis elements while maintaining the rendering quality.

[0025] Technical solution

[0026] A new perspective synthesis method for three-dimensional Gaussian splash based on human eye perception, comprising the following steps:

[0027] Step 1 Training data processing

[0028] Input a set of multi-view two-dimensional images of a three-dimensional scene, estimate the camera positions of each image through the SfM algorithm, and obtain the corresponding sensitivity maps for each image by using perception sensitivity extraction, and finally output the training data;

[0029] Step 2 Initialize 3DGS

[0030] Initialize the 3DGS Gaussian basis element ellipsoids using the initial point cloud obtained through the SfM algorithm, and initialize all parameters of the Gaussian basis element ellipsoids;

[0031] Step 3 Dual-branch rendering

[0032] Based on the 3DGS model obtained in Step 2, project the 3D Gaussian basis elements onto 2D to complete the rendering process. The rendering branch includes two branches: a color branch and a sensitivity branch;

[0033] Among them, the color branch renders the RGB image of the scene (calculating the RGB color of the 3D Gaussian basis element projection using spherical harmonics), and the sensitivity branch renders the sensitivity map (calculating the sensitivity of the 3D Gaussian basis element projection using sensitivity parameters);

[0034] Calculate the overall loss function of the model and learn the color, geometry, and perceptual modeling of the scene;

[0035] Step 4 Density control

[0036] Perform density control on the model using adaptive density control and sensitivity-guided densification;

[0037] Step 5 Scene adaptive depth re-initialization

[0038] Judge whether the initial point cloud of the scene is sparse based on the learning situation of the Gaussian basis element sensitivity, and adaptively optimize the position distribution of the Gaussian basis elements during the training process to avoid the Gaussian basis elements with incorrect initial positions from misleading subsequent training.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] (1) The Gaussian basis element distribution conforms to the human eye perception characteristics: The learnable human eye perception modeling proposed by the present invention can efficiently capture the sensitivity of the human eye to different regions of the scene, and has the potential to be applied to other new perspective image synthesis related tasks such as model lightweighting based on three-dimensional Gaussian splashing.

[0041] (2) Balance between scene reconstruction quality and efficiency: The opacity attenuation mechanism of the cloning basis element proposed by the present invention can significantly reduce the number of Gaussian basis elements required for reconstruction while maintaining the reconstruction quality of the model, achieving a better balance between quality and efficiency.

[0042] (3) Adaptability and robustness of the densification strategy: The present invention proposes a novel perception-based Gaussian basis element distribution strategy, which can robustly improve the reconstruction quality and efficiency even in large-scale urban scenes, and can adaptively adjust the densification strategy according to the scene. Brief description of the drawings

[0043] Figure 1 is the overall flowchart of the present invention;

[0044] Figure 2 is a diagram of double-branch rendering in an embodiment of the present invention;

[0045] Figure 3 is a schematic flowchart of Step 4 in an embodiment of the present invention;

[0046] Figure 4 is a schematic flowchart of Step 5 in an embodiment of the present invention;

[0047] Figure 5 shows the comparison of the visualization results between the embodiment of the present invention and the comparative method on the BungeeNeRF dataset;

[0048] Figure 6 shows the comparison of the visualization results between the embodiment of the present invention and the comparative method on the Mip-NeRF 360 dataset;

[0049] Figure 7 shows the comparison effect of the robustness between the embodiment of the present invention and the comparative method on large-scale urban scenes. Detailed implementation manners

[0050] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.

[0051] A novel view synthesis method based on human eye perception for three-dimensional Gaussian splashing first performs training data processing (step 1) before the start of training through SfM and perceptual sensitivity extraction. During the training process, the present invention first defines and initializes a 3DGS model containing additional sensitivity parameters (step 2), and then simultaneously learns the color, geometry, and perceptual modeling of the scene through dual-branch rendering (step 3), and is constrained by multi-view RGB images and sensitivity maps. Then, in the eligible training rounds, the present invention will perform density control (step 4) and scene adaptive depth re-initialization (step 5), respectively adjusting the density distribution and position distribution of the Gaussian basis elements, while reducing the cloning of redundant basis elements, to achieve novel view synthesis with a balance between quality and efficiency.

[0052] A novel view synthesis method based on human eye perception for three-dimensional Gaussian splashing includes the following steps: (as Figure 1 )

[0053] Step 1 Training data processing

[0054] Input a set of multi-view two-dimensional images of a three-dimensional scene, estimate the camera position of each image through the SfM algorithm, and obtain the sensitivity map corresponding to each image by using perceptual sensitivity extraction, and finally output the training data. Specifically as follows:

[0055] Step 1.1 Camera position estimation

[0056] Use the existing SfM algorithm to process the two-dimensional images of the view , and obtain the camera position of this view .

[0057] Step 1.2 Perceptual sensitivity extraction

[0058] First, use the traditional image edge extraction algorithm to process the two-dimensional images to obtain a gradient magnitude map, and then enhance and smooth it according to the relevant characteristics of the human visual system, and finally obtain each two-dimensional image Correspondingly, an easily learnable scene sensitivity map for supervising the sensitivity learning of Gaussian basis elements.

[0059] Step 1.2.1 Extract image edges

[0060] For each two-dimensional image , use the Sobel operator to extract local structures. The horizontal and vertical gradient convolution kernels of the Sobel operator and are defined as:

[0061] , (5)

[0062] The final edge response map (i.e., the gradient magnitude map) is obtained by calculating through the color image and the convolution kernels in two directions:

[0063] , (6)

[0064] where represents the convolution operation.

[0065] Step 1.2.2 Binarization ( threshold)

[0066] Considering the characteristics perceived by the human visual system, in order to avoid ignoring the detailed structures with relatively small absolute numerical values of pixel response values but still perceptible to the human eye, enhance the extracted gradient magnitude map so that it only retains the binarized information on whether the human eye is sensitive to each pixel:

[0067] , (7)

[0068] where is the pixel value of the enhanced sensitivity map at pixel , is the indicator function, , is the enhancement threshold for each pixel, and in this embodiment, is set.

[0069] Step 1.2.3 Average pooling

[0070] Use average pooling to smooth the enhanced sensitivity map to simulate the human visual characteristics proven by eye movement experiments.

[0071] Step 1.2.4 Binarization ( threshold)

[0072] The smoothed sensitivity Figure two is quantized to obtain the final easily learnable binary perception sensitivity map, and the binary threshold , which is set in this embodiment .

[0073] For each view, the sensitivity map and the original 2D image generated after this step are used as training data and provided to Step 2.

[0074] Step 2 Initializes 3DGS

[0075] The initial point cloud obtained by the SfM algorithm is used to initialize the Gaussian basis element ellipsoids of 3DGS, and all parameters of the Gaussian basis element ellipsoids are initialized.

[0076] In the present invention, for each Gaussian basis element an additional learnable perception parameter is added for rendering the perception sensitivity map. Then, the 3DGS model of the present invention is defined as:

[0077] , (8)

[0078] Step 3 Dual-branch rendering (such as Figure 2 )

[0079] During the 3DGS rendering process, it is necessary to project the 3D Gaussian basis elements onto 2D to achieve efficient rendering.

[0080] Based on the 3DGS model in Step 2, the rendering branch of the model of the present invention includes two branches: the color branch and the sensitivity branch.

[0081] Among them, the color branch renders the scene RGB map (calculating the RGB color of the 3D Gaussian basis element projection using spherical harmonics), and the sensitivity branch renders the sensitivity map (calculating the sensitivity of the 3D Gaussian basis element projection using the sensitivity parameter).

[0082] Calculate the overall loss function of the model to learn the color, geometry, and perception modeling of the scene.

[0083] Specifically as follows:

[0084] Step 3.1 Render the RGB map

[0085] Rendering process of the color branch: Use the 3DGS model and differentiable tile rasterization to render the RGB image.

[0086] Step 3.2 Render the sensitivity map

[0087] The sensitivity branch maps multi-perspective sensitivity information from a two-dimensional image to three-dimensional Gaussian basis elements in a learned manner, and constrains the sensitivities of the spatial ranges covered by each basis element to be unified. The sensitivity map of the scene can be rendered through the sensitivity parameters of each basis element and differentiable tile rasterization:

[0088] , (9)

[0089] where, is the value of the rendered sensitivity map at pixel , is the sigmoid function.

[0090] Step 3.3 Calculate the loss

[0091] During the optimization process, the present invention jointly supervises the model training using the losses of the two rendering branches of color and sensitivity.

[0092] For the color branch, the same loss function as the original three-dimensional Gaussian splash is adopted, and is obtained by weighted summation of the L1 loss and the structural consistency loss.

[0093] For the sensitivity branch, binary cross-entropy loss is used for supervision, so that the rendered sensitivity map is consistent with the corresponding ground-truth sensitivity map . The sensitivity loss is specifically defined as:

[0094] , (10)

[0095] Finally, the overall loss function of the model is obtained by weighted summation of the color loss and the sensitivity loss:

[0096] , (11)

[0097] where, is the weight of the sensitivity loss.

[0098] Step 3.4 Update the 3DGS parameters

[0099] After calculating the loss in each round of training, the model parameters are updated by backpropagation to calculate the gradients.

[0100] Step 4 Density control (the overall process is as Figure 3 shown)

[0101] During the training iterations from 500 to 15000 rounds, the density control process is performed every 100 rounds.

[0102] The present invention improves the densification operation of the standard 3DGS. By optimizing the density distribution of the Gaussian basis elements, the quality and efficiency of scene reconstruction are enhanced. Specifically as follows:

[0103] Step 4.1 Adaptive density control

[0104] The processing of 3DGS adaptive density control includes the splitting, cloning, and cropping of basis elements. The present invention optimizes the cloning operation to promote more cropping of redundant basis elements to improve the reconstruction efficiency. Specifically as follows:

[0105] Step 4.1.1 Screening large position gradient and large size basis elements

[0106] Using the same method as the standard 3DGS, basis elements that need to be split are screened out using gradient and size thresholds.

[0107] Step 4.1.2 Splitting

[0108] Using the same method as the standard 3DGS, the splitting operation is performed on the basis elements screened out in step 4.1.1.

[0109] Step 4.1.3 Screening large position gradient and small size basis elements

[0110] Using the same method as the standard 3DGS, basis elements that need to be cloned are screened out using gradient and size thresholds.

[0111] Step 4.1.4 Opacity attenuation of basis elements to be cloned

[0112] Before the cloning operation of the Gaussian basis elements, the present invention introduces an opacity attenuation mechanism. The purpose is to promote the deletion of redundant Gaussian basis elements during the cropping process by reducing the opacity of the basis elements to be cloned, and to avoid having a negative impact on the training process.

[0113] Specifically, the opacity of the basis elements to be cloned is non-linearly transformed through the opacity attenuation function to achieve the effect of opacity attenuation.

[0114] The principle is explained as follows:

[0115] According to transparency blending, when two Gaussian basis elements with an opacity of overlap, the opacity of the corresponding spatial region can be expressed as:

[0116] , (12)

[0117] Assume that before the cloning operation, the opacity of the spatial region represented by a Gaussian basis element, that is, the opacity of this basis element, is , the goal of the opacity attenuation mechanism is to reduce the opacity of the area after the cloning operation. Specifically, the opacity of the cloned primitive will be non-linearly transformed through the opacity attenuation function to achieve the effect of opacity attenuation.

[0118] Finally, the opacities of the two cloned Gaussian primitives can be obtained by solving the following equation:

[0119] , (13)

[0120] It can be solved to get , that is, adjust the opacity of the primitive to be cloned to .

[0121] The selection of the opacity attenuation function needs to meet the following requirements: a larger attenuation is applied to the primitive with a smaller opacity to encourage the Gaussian primitive to be cropped; while a smaller reduction is made to the higher opacity to avoid removing important Gaussian primitives and causing a cliff-like drop in the model performance.

[0122] Specifically, it is required that satisfies the following conditions within the range of :

[0123] (1) , indicating that the value after transformation is not greater than the original value;

[0124] (2) , indicating that the value after transformation is still within the range of ;

[0125] (3) , where is the stationary point of the function and satisfies its first derivative , indicating that a smaller opacity will obtain a larger attenuation.

[0126] As an example, the power function is used as .

[0127] To determine the optimal value of , various exponents are tested. Finally, in this example, is preferred, which can balance the reconstruction quality and efficiency.

[0128] Step 4.1.5 Cloning

[0129] Using the same method as the standard 3DGS, perform the cloning operation on the primitives screened in step 4.1.3.

[0130] Step 4.1.6 Gaussian basis element clipping

[0131] Adopt the same method as the standard 3DGS to delete the basis elements with opacity lower than the threshold of 0.005.

[0132] Step 4.2 Sensitivity-guided densification

[0133] Based on Step 4.1, the present invention uses the scene perception sensitivity learned by each basis element to guide the spatial distribution of Gaussian basis elements in different regions, prompting more Gaussian basis elements to be distributed in the regions that require more accurate reconstruction.

[0134] Step 4.2.1 Screening of high-sensitivity Gaussian basis elements

[0135] Since the learned true sensitivity map is binary, the sensitivity of Gaussian basis elements with ideal learning conditions should be close to 0 or 1. Among them, the basis elements with sensitivity close to 1 represent the human eye sensitive regions in the scene, and vice versa, the basis elements close to 0 represent non-sensitive regions. Specifically, high-sensitivity Gaussian basis elements can be screened by the threshold . A larger value can be selected. In this embodiment, it is set :

[0136] , (14)

[0137] This step is executed once every round during 500 - 15000 rounds of training iterations. When not executed, is an empty set. In this embodiment, it is set to balance the reconstruction quality and efficiency.

[0138] Step 4.2.2 Screening of medium-sensitivity Gaussian basis elements

[0139] For the Gaussian basis elements with large sensitivity differences under different perspectives, the training process is likely to make the intermediate values of the sensitivity regions of these basis elements to average the differences of different perspectives. Therefore, using only a single basis element cannot represent such complex regions, and these basis elements need to be split into smaller units.

[0140] Use the thresholds and to screen medium-sensitivity Gaussian basis elements :

[0141] , (15)

[0142] This step is executed once every round during 500 - 15000 rounds of training iterations. When not executed, is an empty set. In this embodiment, it is set , to balance the reconstruction quality and efficiency.

[0143] Step 4.2.3 Screen high-weight primitives

[0144] To prevent too many Gaussian primitives from appearing inside the object, the present invention imposes a weight limit on the Gaussian primitives screened according to sensitivity. Finally, the Gaussian primitives that need to be additionally densified Are defined as:

[0145] , (16)

[0146] , (17)

[0147] , (18)

[0148] , (19)

[0149] Where Calculates the weight of the primitive At the perspective Of the pixels On, And Represent different weight thresholds for highly sensitive and medium-sensitive Gaussian primitives respectively, Represents the primitives that need to be densified screened from the highly sensitive Gaussian primitives, Represents the primitives that need to be densified screened from the medium-sensitive Gaussian primitives; Selects the largest element in the set, Represents the perspective Of all pixels in, For all perspectives in the scene Of the Gaussian primitive The maximum weight.

[0150] Step 4.2.4 Calculate the scene sensitivity

[0151] The present invention determines the specific densification operation of the screened primitives through the scene sensitivity. Specifically, the scene sensitivity Can be defined as the average pixel sensitivity of all perspectives :

[0152] , (20)

[0153] , (21)

[0154] Where Is the pixel The sensitivity at represents the viewing angle of the average sensitivity.

[0155] Step 4.2.5 Densifying High-Sensitivity Gaussian Primitives

[0156] During the densification process, the original three-dimensional Gaussian splash performs splitting and cloning operations according to the size of the Gaussian primitive. To better capture scene details, the present invention only performs both the cloning operation described in Step 4.1.4 and the splitting operation described in Step 4.1.2 when the scene sensitivity is lower than the threshold Otherwise, only the selected Gaussian primitives are split, regardless of their size, to represent complex details using more small-sized primitives.

[0157] Step 4.2.6 Densifying Medium-Sensitivity Gaussian Primitives

[0158] This step densifies using the same splitting operation as in Step 4.1.2.

[0159] Step 5 Scene Adaptive Depth Re-Initialization (the overall process is as Figure 4 shown)

[0160] This step determines whether the initial point cloud of the scene is sparse based on the learning of the sensitivity of the Gaussian primitives, and adaptively optimizes the position distribution of the Gaussian primitives during the training process to avoid misleading subsequent training by Gaussian primitives with incorrect initial positions.

[0161] Step 5.1 Calculating the Proportion of Medium-Sensitivity Large Primitives

[0162] For scenes where the initial point cloud is too sparse, a large-sized Gaussian is usually used to represent multiple regions with different spatial sensitivities, resulting in poor learning of the sensitivity of this primitive. Taking advantage of this feature, after 600 rounds of warm-up training, the present invention uses the proportion of medium-sensitivity Gaussian primitives among the large-sized Gaussian primitives

[0163] as an indicator for determining whether the initial point cloud is sparse:

[0164] where,

[0165] (23)

[0166] where is the scaling of the longest axis, Denote the set of the longest axis scaling of all Gaussian basis elements, which represents the third quartile and is used to identify the 25% largest Gaussian basis elements .

[0167] If exceeds the predefined threshold , then perform the depth re - initialization operation on this scenario in subsequent training, otherwise do not perform it.

[0168] Specifically, when the number of training epochs is 5000, 10000, and 15000, the present invention adopts the scenario - adaptive depth re - initialization strategy and does not perform it in other epochs.

[0169] Step 5.2 Depth Re - initialization

[0170] This step is the same as the depth re - initialization step in Mini - Splatting. This step finds the Gaussian basis element with the largest weight for each pixel and takes the mid - point depth of the two intersection points of the ray from the camera to this pixel and the ellipsoid of the Gaussian basis element with the largest weight as the depth of this pixel, thus obtaining the depth - color pair for each pixel. Then randomly sample 35000 non - repeating depth - color pairs as new 3D points to re - initialize the Gaussian basis elements.

[0171] Embodiment

[0172] Dataset

[0173] The method proposed by the present invention is tested on 21 scenarios of Mip - NeRF 360, Tanks&Temples, Deep Blending, and BungeeNeRF datasets, covering various different types of scenarios such as indoor, outdoor, and large - scale cities.

[0174] Among them, the Mip - NeRF 360 dataset is designed for unbounded 360 - degree scenarios and contains 9 complex indoor and outdoor scenarios (such as streets, buildings, natural landscapes) and supports dynamic view changes.

[0175] Tanks&Temples is a large - scale real - world scenario dataset that contains indoor and outdoor high - resolution videos and laser scan data and is used to evaluate the reconstruction accuracy and generalization ability of NeRF - like models. It includes 14 complex scenarios (such as churches, castles, forests) and covers different lighting conditions and dynamic elements (such as pedestrians, vehicles).

[0176] The Deep Blending dataset contains 19 complex scenarios (such as indoor, streets, natural landscapes), and each scenario provides 12 to hundreds of images.

[0177] BungeeNeRF is a model and a supporting dataset designed specifically for extremely multi-scale urban scenes, enabling seamless rendering from satellite view to ground view. It includes 12 global city landmarks (such as New York and Tokyo), with thousands of multi-resolution images for each scene.

[0178] Method Application

[0179] All experiments in this embodiment were completed on a single NVIDIA RTX 4090 GPU. To achieve a better balance between quality and efficiency, different weight thresholds were used for high-sensitivity and medium-sensitivity Gaussian basis elements, labeled respectively as and . The specific values of each hyperparameter in this embodiment are shown in Table 1.

[0180] Table 1 Value Settings of Each Hyperparameter in this Embodiment

[0181]

[0182] Comparison of Implementation Effects

[0183] The comparison methods in this embodiment include: Pixel-GS, Mini-Splatting-D, Taming-3DGS, and 3DGS. Among them, Pixel-GS is an optimization method for density control of 3D Gaussian splashing, which solves the blurring and needle-like artifact problems in sparse point cloud regions through pixel-aware gradients and gradient field scaling; Mini-Splatting-D achieves ultra-fast optimization through aggressive Gaussian densification and visibility culling; Taming-3DGS proposes a budget-controlled Gaussian optimization framework to solve the efficiency problem in resource-constrained scenarios. The first three methods all optimize the densification strategy, while the last method is the original three-dimensional Gaussian splashing method.

[0184] The metrics used are: performance comparison results on SSIM (Structural Similarity Index), LPIPS (Learned Perceptual Image Patch Similarity), the number of Gaussian basis elements (#G), and Quality-Efficiency Balance (QEB). Among them, QEB is a metric that combines reconstruction quality and efficiency, defined as:

[0185] , (24)

[0186] Table 2 Experimental Results of Reconstruction Quality Comparison between the Invention and Other Models

[0187]

[0188] Table 2 shows the quantitative comparison results of the reconstruction quality between the present invention and the comparative methods on different datasets. Among the four datasets, the present invention performs excellently in terms of SSIM and the perception-related LPIPS metric. Different from Pixel-GS and Taming-3DGS, which suffer from CUDA out-of-memory (OOM) problems due to introducing too many Gaussian basis elements in large-scale scenes, the present invention adaptively allocates basis elements according to the perceptual sensitivity of different regions, achieving a better quality-efficiency trade-off.

[0189] Table 3 Comparative experimental results of the reconstruction efficiency between the present invention and other models

[0190]

[0191] Table 3 provides a quantitative analysis of the model complexity and the trade-off between quality and efficiency between the present invention and the comparative methods on different datasets. Compared with other quality-oriented methods, the present invention shows significant improvements in terms of efficiency and quality-efficiency balance, rendering realistic novel-view images at a faster speed and with fewer Gaussian basis elements.

[0192] The processed effect diagrams are as Figure 5 、 Figure 6 and Figure 7 shown. The present invention allocates more Gaussian basis elements to object details and edges, effectively reducing the blurriness of the scene. As shown in Figure 5 , the present invention accurately reconstructs roads, buildings, and grasslands at the scene boundary, avoiding the distortion phenomena that occur in other methods. Similarly, in Figure 6 , the method of the present invention reconstructs the ground texture more realistically, while other methods tend to produce more artifacts in these regions.

[0193] As Figure 7 shows the effect comparison of the robustness between the present invention and the comparative methods (Pixel-GS and Mini-Splatting-D) on large-scale urban scenes. The numbers in parentheses in the figure represent the LPIPS metric and the number of Gaussian basis elements (in millions, M). For example, the metric (0.116, 5.03) of the present invention indicates that the method of the present invention has an LPIPS metric of 0.116 and a Gaussian basis element number of 5.03M for this image. The method of the present invention shows extremely high robustness in large-scale scenes, avoiding introducing too many Gaussian basis elements like Pixel-GS and reconstruction failures in Mini-Splatting-D. This benefits from the dual-branch rendering and sensitivity-guided densification of the present invention, which limit the number of densified Gaussian basis elements while adaptively identifying regions that require more basis elements.

[0194] Key points and advantages of the present invention

[0195] Learnable Human Eye Perception Modeling: Before the start of training, the present invention operates on the multi-view color images of the scene by extracting the perceptual sensitivity, and pre-computes the binary sensitivity map of the three-dimensional scene for training the sensitivity parameters of the Gaussian model. During the training process, the sensitivity parameters of each Gaussian basis element are rendered and optimized in a similar manner to its color and opacity parameters, and finally, the human eye perception modeling of the scene can be learned. Since the sensitivity map used for training is binary, compared with directly learning the continuous gradient magnitude map, the Gaussian basis element is easier to capture the spatial perception sensitivity, avoiding performance degradation caused by poor learning. In addition, compared with statistically judging the sensitivity of the Gaussian basis element by covering the sensitivity pixel values of the two-dimensional projection of the three-dimensional Gaussian basis element in multiple views, the learnable human eye perception modeling proposed by the present invention can make the spatial perception sensitivities represented by the same Gaussian basis element similar, which is beneficial to avoiding the densification of Gaussian basis elements in insensitive regions, and at the same time more accurately reflecting the sensitive regions or Gaussian basis elements with poor learning for subsequent additional densification operations.

[0196] Perception-based Gaussian Basis Element Distribution Strategy: Compared with the relatively fixed method of the existing work, the densification strategy proposed by the present invention shows good robustness in different types of scenes, and can adaptively adjust the strategy according to the scene characteristics to achieve a better reconstruction effect. Since the splitting strategy can increase more small-sized Gaussian basis elements in the scene, for a scene containing a large number of sensitive regions, only splitting operations in the sensitive regions can accurately represent the scene details, thereby improving the reconstruction quality. For a scene with sparse initial point clouds, after 600 rounds of warm-up training, more large-sized Gaussians will have medium sensitivity, because the sparse initial point clouds make the spatial regions with different sensitivities can only be represented by a small number of basis elements, and accurate reconstruction cannot be achieved. These basis elements often contain incorrect spatial distributions. Although densification can improve the overall reconstruction quality to a certain extent, it will cause subsequent training to continuously distribute new Gaussian basis elements in the wrong positions, hindering the further improvement of the model performance. Selectively introducing the deep re-initialization strategy for this type of scene to timely adjust the positions of the basis elements with incorrect position distributions can effectively avoid misleading subsequent training, and at the same time will not damage the well-reconstructed regions due to random sampling, and can show good robustness in different scenes.

[0197] Opacity attenuation mechanism of cloned primitives: The cloning operation of the original three-dimensional Gaussian splash will copy a Gaussian with all parameters completely consistent with the cloned primitive, which will cause the opacity of the spatial area represented by it to increase, making the area that was originally poorly learned play a more important role in rendering, hindering model training. The present invention promotes the clipping of poorly learned redundant primitives by nonlinearly reducing the opacity of primitives with different opacities after the cloning operation, avoiding excessive cloning of redundant primitives, resulting in reduced storage and computing efficiency. Specifically, for Gaussian primitives with lower original opacity, their opacity is reduced more significantly after cloning, because transparent poorly learned Gaussian primitives have less impact on the rendering results and are more likely to be redundant primitives. On the contrary, for Gaussian primitives with higher original opacity, their opacity is reduced to a smaller extent after cloning to further determine whether they are redundant in the subsequent training process, avoiding the rapid deletion of important Gaussian primitives causing a cliff-like drop in model quality.

[0198] The above description is only a description of the preferred embodiments of the present application, and is not intended to limit the scope of the present application. Any changes or modifications made by any person skilled in the art based on the above disclosed technical contents shall be deemed as equivalent effective embodiments and shall fall within the scope of protection of the technical solution of the present application.

Claims

1. A new perspective synthesis method for three-dimensional Gaussian splash based on human eye perception, characterized in that It includes the following steps: Step 1: Training data processing Input a set of multi-view two-dimensional images of a three-dimensional scene, estimate the camera position of each image through the SfM algorithm, and use perceptual sensitivity extraction to obtain the sensitivity map corresponding to each image, and finally output the training data; Step 2: Initialize 3DGS Use the initial point cloud obtained by the SfM algorithm to initialize the 3DGS Gaussian basis element ellipsoid, and initialize all parameters of the Gaussian basis element ellipsoid; Step 3: Dual-branch rendering Based on the 3DGS model obtained in Step 2, project the 3D Gaussian basis elements onto 2D to complete the rendering process. The rendering branch includes two branches: the color branch and the sensitivity branch; among them, the color branch renders the RGB map of the scene, and the sensitivity branch renders the sensitivity map; Calculate the overall loss function of the model and learn the color, geometry, and perceptual modeling of the scene; Step 4: Density control Use adaptive density control and sensitivity-guided densification to perform density control on the model; Step 5: Scene adaptive depth re-initialization Judge whether the initial point cloud of the scene is sparse through the learning situation of the Gaussian basis element sensitivity, and adaptively optimize the position distribution of the Gaussian basis elements during the training process to avoid the Gaussian basis elements with incorrect initial positions from misleading subsequent training; The specific content of Step 3 is as follows: Step 3.1: Render the RGB map Use the 3DGS model and differentiable tile rasterization to render the RGB image; Step 3.2: Render the sensitivity map The sensitivity branch maps the multi-view sensitivity information from the two-dimensional image to the three-dimensional Gaussian basis elements through learning, and constrains the sensitivity of each basis element to cover the same spatial range; render the sensitivity map of the scene through the sensitivity parameters of each basis element and differentiable tile rasterization: Among them, is the value of the rendered sensitivity map at pixel u, σ is the sigmoid function, is a three-dimensional Gaussian basis element is the weight of the elliptical transparency under the viewing angle v, ∈ i is a learnable perception parameter; Step 3.3: Calculate the loss During the optimization process, use the losses of the color and sensitivity rendering branches to jointly supervise the model training; For the color branch, adopt the same loss function, which is obtained by weighted summation of L1 loss and structural consistency loss; For the sensitivity branch, binary cross-entropy loss is adopted for supervision, so that the rendered sensitivity map is consistent with the corresponding ground-truth sensitivity map ; the sensitivity loss is specifically defined as: Finally, the overall loss function of the model is obtained by the weighted sum of the color loss and the sensitivity loss: where λ S is the weight of the sensitivity loss; Step 3.4: Update 3DGS parameters After calculating the loss in each round of training, calculate the gradient through backpropagation to update the model parameters.

2. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 1, wherein The specific content of Step 1 is: Step 1.1: Camera position estimation Process the 2D image of view i using the SFM algorithm Obtain the camera position of this view Step 1.2: Perceptual sensitivity extraction First, use an image edge extraction algorithm to process the two-dimensional image Obtain the gradient magnitude map, then enhance and smooth it, and finally obtain each two-dimensional image The corresponding sensitivity map For supervising the sensitivity learning of Gaussian basis elements; Step 1.2.1: Extract image edges For each two-dimensional image Extract local structures using the Sobel operator. The horizontal and vertical gradient convolution kernels G of the Sobel operator x and G y are defined as: The final edge response map G is obtained by calculating the color image and convolution kernels in two directions: Among them, represents a convolution operation; Step 1.2.2: Binarization Enhance the extracted gradient magnitude map so that it only retains the binarized information of whether the human eye is sensitive to each pixel: Among them, G E (u) is the pixel value of the enhanced sensitivity map at pixel u, is an indicator function, τ e ∈[0,1], which is the enhancement threshold for each pixel; Step 1.2.3: Average pooling Use 2×2 average pooling to smooth the enhanced sensitivity map to simulate human visual characteristics; Step 1.2.4: Binarization Binarize the smoothed sensitivity map to obtain the final binarized sensitivity map that is easy to learn, with the binarization threshold τ s ∈[0, 1]; For each view, the sensitivity map generated after this step of processing and the original two-dimensional image are used as training data together and provided to Step 2.

3. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 2, wherein Set τ e = 0.05, τ s = 0.

5.

4. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 1, characterized in that, In step 2, for each Gaussian basis element a learnable perception parameter ∈ i is added to render the sensitivity map. Then, the 3DGS model is defined as:

5. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 1, characterized in that, The specific content of Step 4 is: During the training iteration from 500 rounds to 15000 rounds, perform density control processing every 100 rounds; Step 4.1: Adaptive density control The processing of 3DGS adaptive density control includes the splitting, cloning, and cropping of basis elements, which are specifically as follows: Step 4.1.1: Screen large-position-gradient large-size basis elements Use gradient and size thresholds to screen out the basis elements that need to be split; Step 4.1.2: Split Perform split operations on the basis elements screened out in Step 4.1.1; Step 4.1.3 Filter large position gradient and small-size primitives Filter out the primitives to be cloned using gradient and size thresholds; Step 4.1.4 Opacity attenuation of primitives to be cloned Before performing the cloning operation on the Gaussian primitives, the opacity of the primitives to be cloned is non-linearly transformed through the opacity attenuation function OD(·) to achieve the effect of opacity attenuation; Step 4.1.5 Cloning Perform the cloning operation on the primitives filtered in Step 4.1.3; Step 4.1.6 Gaussian primitive cropping Delete the primitives with opacity lower than the threshold of 0.005; Step 4.2 Sensitivity-guided densification Based on Step 4.1, use the scene perception sensitivity learned by each primitive to guide the spatial distribution of Gaussian primitives in different regions, promoting more Gaussian primitives to be distributed in the regions that require more accurate reconstruction; Step 4.2.1 Filter high-sensitivity Gaussian primitives Perform the operation of screening high-sensitivity Gaussian basis elements once every Iter during 500 - 15000 training iterations h round; Through the threshold τ h Select high-sensitivity Gaussian basis elements When screening is not performed it is an empty set; Step 4.2.2 Filter medium-sensitivity Gaussian primitives Perform a high-sensitivity Gaussian basis element operation in the screening once every Iter during 500 - 15000 training iterations m round; Using threshold τ h and τ l to screen highly sensitive Gaussian basis elements When screening is not performed it is an empty set; Step 4.2.3 Filter high-weight primitives Apply weight restrictions to the Gaussian basis elements selected according to sensitivity, and the Gaussian basis elements that need to be additionally densified It is defined as: Among them, the weight of primitive i on pixel u of view v is calculated, and respectively represent different weight thresholds for high-sensitivity and medium-sensitivity Gaussian primitives, represents the primitive to be densified selected from high-sensitivity Gaussian primitives, represents the primitive to be densified selected from medium-sensitivity Gaussian primitives; MAX(·) selects the largest element in the set, pix v represents all pixels in view v, is the maximum weight of Gaussian primitives in all views V in the scene; Step 4.2.4 Calculate scene sensitivity Judge the specific densification operation of the primitives filtered through scene sensitivity; the scene sensitivity β is defined as the average pixel sensitivity of all viewpoints V: where v(u) is the sensitivity at pixel u, and avg v represents the average sensitivity of the viewing angle v; Step 4.2.5 Densify high-sensitivity Gaussian primitives When the scene sensitivity β is lower than the threshold τ β then, for both the cloning operation described in Step 4.1.4 and the splitting operation described in Step 4.1.2 are adopted simultaneously; conversely, only the screened Gaussian basis elements are split; Step 4.2.6 Densify medium-sensitivity Gaussian primitives Pair Densify by performing the same splitting operation as in step 4.1.

2.

6. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 5, wherein Using the power function x k as OD(·), where k = 1.

2.

7. The three-dimensional Gaussian splash new perspective synthesis method based on human eye perception according to claim 5, wherein Specifically in Step 5 as follows: Step 5.1 Calculate the proportion of medium-sensitivity large primitives After 600 rounds of warm-up training, use the proportion γ of large-size medium-sensitivity Gaussian primitives as an indicator to judge whether the initial point cloud is sparse: Wherein, wherein is the scaling of the longest axis, S max represents the set of the longest axis scalings of all Gaussian basis elements, and Q3 represents the third quartile, which is used to identify the 25% largest Gaussian basis elements If γ exceeds the predefined threshold τ γ , then perform a deep re-initialization operation on this scenario during subsequent training, otherwise do not perform it; Step 5.2 Depth re-initialization This step is the same as the depth re-initialization step in Mini-Splatting; this step finds the Gaussian primitive with the largest pixel weight for each pixel, and takes the midpoint depth of the two intersection points of the ray from the camera to the pixel and the ellipsoid of the Gaussian primitive with the largest weight as the depth of the pixel, so as to obtain the depth-color pair for each pixel; then randomly sample 35000 non-repeating depth-color pairs as new 3D points to re-initialize the Gaussian primitives.

Citation Information

Patent Citations

  • Three-dimensional human body generation method based on text prompt

    CN118229860A

  • Three-dimensional object generation method and device, equipment and storage medium

    CN118429531A