Double-branch three-dimensional Gaussian rendering optimization method
Through the dual-branch three-dimensional Gaussian rendering optimization method, the dual-path rendering refinement strategy is used to independently optimize the insertion target and background environment, solving the problem of edge blur in three-dimensional scene combination editing, significantly improving the rendering quality and visual realism.
Patent Information
- Application Number
- CN202510103462.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
AI Technical Summary
In 3D scene combination editing, there are difficulties in edge optimization and detail enhancement of insertion targets and backgrounds, resulting in blurring or distortion of edges, reducing rendering quality and scene realism.
The dual-branch three-dimensional Gaussian rendering optimization method is adopted. By introducing a dual-path rendering refinement strategy, the insertion target and background environment are independently optimized. The three-dimensional Gaussian representation form combined with the optimization scheme of independent rendering paths is used to achieve the separation of foreground and background.
It significantly improves the accuracy of inserting the Gaussian ellipsoid into the target boundary, eliminates floating noise in the boundary area, improves the overall quality and visual reality of the rendering results, and has higher geometric accuracy and scene restoration capabilities.
Smart Images

Figure CN120088380A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, computer graphics, and three-dimensional representation optimization, and specifically relates to a dual-branch three-dimensional Gaussian rendering optimization method. Background Art
[0002] In recent years, with the continuous development of artificial intelligence and graphics rendering technologies, three-dimensional scene modeling and editing have gradually become one of the important technologies driving digital innovation. Whether in the fields of game design, film and television production, intelligent driving simulation, and virtual reality and augmented reality, high-quality three-dimensional scene modeling is indispensable. Especially with the rise of the metaverse concept, how to quickly generate three-dimensional scenes with high realism and visual consistency has become a key challenge for industrial development. However, in the combination editing of three-dimensional scenes, the edge optimization and detail enhancement of the inserted object and the background have always been technical problems to be solved urgently.
[0003] In actual scene editing tasks, since the Gaussian of the inserted object needs to be jointly optimized and rendered with the Gaussian of the background, the edge area of the inserted object is prone to blurring or distortion. This not only reduces the visual quality of the final rendering but also has a negative impact on the scene realism and coherence.
[0004] The current solutions for three-dimensional scene combination editing mainly include the following:
[0005] 1. Optimization method based on mask and segmentation
[0006] The optimization method based on mask and segmentation completes the separation and subsequent optimization of the foreground and background by providing an interest area mask for the source object or segmenting the Gaussian subset in the three-dimensional space. However, the optimization method based on mask and segmentation is prone to introducing floating noise during the separation process. Especially for the boundary areas in complex scenes, the noise will damage the accurate reconstruction of the scene boundary, resulting in edge blurring and visual incoherence in the final rendering result.
[0007] 2. Joint optimization method based on global supervision
[0008] The joint optimization method based on global supervision imposes unified supervision constraints on the entire scene (including the foreground and background) during the optimization process to ensure the overall consistency of the scene. However, since it ignores the specific interaction relationship between the foreground object and the background Gaussian, it is easy to cause local detail loss or texture distortion of the inserted object, and it is difficult to eliminate the blurring phenomenon in the boundary area.
[0009] 3. Local optimization method based on independent paths
[0010] The local optimization method based on independent paths optimizes and refines the foreground and background independently to reduce mutual interference and improve the rendering quality. Although it improves the detailed performance of the scene to a certain extent, the independent optimization often lacks the coordinated processing between the foreground and the background, which easily leads to a sense of visual fragmentation in the final rendering result.
[0011] 4. Generation method based on hybrid strategy
[0012] The generation method based on hybrid strategy combines local supervision and global consistency optimization strategy, and balances the scene editing efficiency and visual quality through the combination of data-driven and rule-driven methods. However, the implementation complexity of the hybrid strategy is relatively high, and there are still problems with insufficiently fine optimization when dealing with the boundary between the inserted object and the background, and it is impossible to completely avoid the phenomena of edge blurring or structural distortion. Summary of the invention
[0013] In view of this, the present invention provides a dual-branch three-dimensional Gaussian rendering optimization method. By introducing a dual-path rendering refinement strategy, this method effectively solves the optimization problem of the Gaussian ellipsoids at the edges of the inserted object in the multi-three-dimensional Gaussian representation combination task.
[0014] In order to solve the above technical problems, the present invention is implemented as follows.
[0015] A dual-branch three-dimensional Gaussian rendering optimization method includes the following steps:
[0016] Step 1: Obtain the multi-view image data of the inserted object and the multi-view image data of the environmental background, perform preprocessing, and obtain the initial point cloud of the inserted object and the initial point cloud of the environmental background;
[0017] Step 2: Based on the three-dimensional Gaussian representation form, use the multi-view image data and the initial point cloud of the inserted object and the environmental background obtained in Step 1 as inputs, and respectively construct three-dimensional Gaussian models of the inserted object and the environmental background and corresponding three-dimensional Gaussian model renderers;
[0018] Step 3: Through the three-dimensional Gaussian models of the inserted object and the environmental background obtained in Step 2, perform spatial alignment and integration on the inserted object and the environmental background to generate a combined scene three-dimensional Gaussian model;
[0019] Step 4: Through the three-dimensional Gaussian model renderers of the inserted object and the environmental background obtained in Step 2, respectively construct independent foreground rendering paths and background rendering paths;
[0020] Step 5: Through the foreground rendering path and the background rendering path, optimize and fine-tune the combined scene three-dimensional Gaussian model respectively according to the inserted target mask image and the environmental background mask image extracted from the corresponding multi-view image data; the two fine-tuning trainings share the same total loss function, and this total loss function is constructed based on the rendering outputs and the ground truths of the models of the two fine-tuning trainings.
[0021] Step 6: After the training is completed, obtain the combined scene three-dimensional Gaussian model, and use the combined scene three-dimensional Gaussian model renderer to perform combined scene rendering to complete the three-dimensional Gaussian combined scene characterization optimization task.
[0022] Preferably, in the step 1, the multi-view image data is collected by a visible light imaging device; the visible light imaging device uses a digital camera, a mobile phone camera or a drone, and the collection range covers the complete geometric structures of the inserted target and the environmental background.
[0023] Preferably, in the step 1, the multi-view image data is preprocessed by using the Structure from Motion and Multi-View Stereo method SfM-MVS.
[0024] Preferably, in the step 3, the three-dimensional Gaussian models of the inserted target and the environmental background are spatially aligned and integrated by using three-dimensional coordinates to obtain the combined scene three-dimensional Gaussian model; in the combination process, the Gaussian ellipsoid parameters of the inserted target and the background are directly spliced.
[0025] Preferably, in the step 5, the mask images of the inserted target and the background environment are obtained by an automatic or semi-automatic annotation tool, or in a manual annotation form.
[0026] Preferably, in the step 6, the total loss function of the optimization and fine-tuning training is:
[0027]
[0028]
[0029] In the formula, represents the loss function of the background environment, represents the loss function of the inserted target; C bg represents the rendering output of the current round when the combined scene three-dimensional Gaussian model is optimized and fine-tuned by using the background rendering path, and C fg represents the rendering output of the current round when the combined scene three-dimensional Gaussian model is optimized and fine-tuned by using the foreground rendering path; represents the ground truth of the background environment, represents the ground truth of the inserted target, and λ lpips represents the weight coefficient, and LPIPS represents the learned perceptual image patch similarity loss function.
[0030] |||| 1 Represents the L1 loss function.
[0031] Beneficial effects:
[0032] (1) The proposed dual-path rendering refinement strategy in the present invention independently optimizes the insertion target and the background environment, significantly improving the accuracy of the Gaussian ellipsoid at the boundary of the insertion target in the 3D scene combination task, effectively eliminating the floating noise in the boundary area, and enhancing the overall quality and visual realism of the rendering result.
[0033] (2) The dual-path rendering refinement strategy provided by the present invention adopts an optimization scheme combining 3D Gaussian representation forms and independent rendering paths, realizing the separation processing of the foreground and the background. Compared with traditional scene editing methods, the present invention has higher geometric accuracy and scene restoration ability in the modeling and rendering of multi-view scenes, while reducing the computational complexity of the Gaussian model.
[0034] (3) The proposed optimization and fine-tuning training mechanism in the present invention, through the combined use of perceptual loss (LPIPS) and L1 loss, ensures the texture consistency and geometric detail expressiveness of the foreground and background regions. Combining with the optimization strategy of independent paths, it solves the problems of boundary blurring and noise interference in the combined scene of multiple 3D Gaussian representations.
[0035] (4) The dual-path rendering refinement strategy provided by the present invention is applicable to various 3D scene editing and rendering tasks, including virtual reality, augmented reality, game design, film and television production, and architectural visualization. The optimized combined scene 3D Gaussian model has the characteristics of clear boundaries, rich details, and coherent structures, providing an efficient and accurate solution for 3D content generation in multiple fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Is a flowchart of the architecture provided by the present invention;
[0037] Figure 2 Is a schematic diagram of the dual-branch 3D Gaussian rendering optimization of the present invention;
[0038] Figure 3 Is an example diagram of the comparison result of the proposed method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.
[0040] The present invention will be described in detail below with reference to the accompanying drawings.
[0041] As Figure 1As shown in the figure, the present invention provides a dual-branch three-dimensional Gaussian rendering optimization method, which specifically includes the following steps:
[0042] Step 1: Obtain multi-view image data of the insertion target and the environmental background, perform preprocessing, and generate initial point cloud data of the insertion target and the environmental background that contains their basic geometric distribution characteristics.
[0043] Among them, the multi-view image data can be collected by any visible light imaging device, such as digital cameras, mobile phone cameras, and drones, etc., and ensure that the data covers a sufficient viewing range to completely describe the geometric information of the insertion target and the environmental background.
[0044] Preferably, in a preferred embodiment, the Structure-from-Motion—Multi-View Stereo (SfM-MVS) method is used to preprocess the multi-view image data.
[0045] Step 2: Based on the three-dimensional Gaussian representation form, using the multi-view image data and the initial point cloud of the insertion target and the environmental background obtained in Step 1 as inputs, respectively construct three-dimensional Gaussian models of the insertion target and the environmental background and the corresponding three-dimensional Gaussian model renderers. As Figure 2 shown.
[0046] In this step, based on the generated initial point cloud data, the three-dimensional Gaussian representation form is used to generate a set of Gaussian ellipsoids to describe the three-dimensional spatial structure of the insertion target and the environmental background, and obtain the three-dimensional Gaussian model of the insertion target and the three-dimensional Gaussian model of the environmental background, as Figure 2 shown. The parameters of each Gaussian ellipsoid include basic attributes such as the center point position, normal vector, major axis and minor axis lengths, etc.; respectively generate corresponding three-dimensional Gaussian model renderers for the insertion target and the environmental background, which are used to convert the geometric and color information of the Gaussian model into a three-dimensional visualization result, providing support for subsequent rendering and optimization.
[0047] Step 3: Through the three-dimensional Gaussian models of the insertion target and the environmental background obtained in Step 2, perform spatial alignment and integration on the insertion target and the environmental background to generate a combined scene three-dimensional Gaussian model.
[0048] In this step, the three-dimensional Gaussian models of the combined insertion target and the environmental background perform spatial alignment and integration on the Gaussian models of the insertion target and the environmental background according to the three-dimensional space coordinates to generate a combined scene three-dimensional Gaussian model; the combination process directly stitches the Gaussian ellipsoid parameters of the insertion target and the background, while ensuring the spatial consistency and structural coherence of the model.
[0049] Step 4: As Figure 2As shown, for the 3D Gaussian model renderer of the insertion target and the environmental background obtained through Step 2, independent foreground rendering paths and background rendering paths are respectively constructed.
[0050] Use the Gaussian model renderers of the insertion target and the environmental background to respectively construct independent foreground (insertion target) rendering paths and background (environmental background) rendering paths. In subsequent steps, the foreground rendering path is used to render the insertion target part through the Gaussian model renderer R of the insertion target fg and the background rendering path is used to render the background part through the Gaussian model renderer R of the environmental background bg to ensure an independent rendering process where the foreground and background are separated.
[0051] Step 5: Through the foreground rendering path and background rendering path constructed in Step 4, according to the insertion target mask image and the background environmental mask image extracted from the corresponding multi-view image data, the combined scene is respectively optimized and fine-tuned for training.
[0052] In this step, an object mask is used to extract the insertion target mask image from the multi-view image data of the insertion target to distinguish the foreground object in the combined scene; through the foreground rendering path, according to the insertion target mask image, the foreground object part in the 3D Gaussian model of the combined scene is optimized and fine-tuned for training. During fine-tuning training, the insertion target mask image serves as the training ground truth, and at the same time, the insertion target mask area serves as the loss calculation area.
[0053] At the same time, taking the negation of the object mask as the background mask, the background mask is used to extract the environmental background mask image from the multi-view image data of the environmental background to distinguish the background environment in the combined scene; through the background rendering path, according to the environmental background mask image, the environmental background part in the 3D Gaussian model of the combined scene is optimized and fine-tuned for training. Similarly, during fine-tuning training, the environmental background mask image serves as the training ground truth, and at the same time, the environmental background mask area serves as the loss calculation area.
[0054] Based on the corresponding mask images extracted from the multi-view image data, the present invention respectively supervises the foreground object and the background environment, and can remove the noise caused by the separation of the insertion target in the original scene and the combination in the new scene in the combined scene.
[0055] Among them, the mask images of the foreground object and the background environment can be obtained through automatic / semi-automatic annotation tools or manual annotation forms to distinguish the foreground object and the background environment in the multi-view rendered images of the combined scene.
[0056] The two fine-tuning trainings share the same total loss function, which is constructed based on the rendering outputs of the models of the two fine-tuning trainings and the ground truth (mask image). The total loss function is:
[0057]
[0058] Among them, The loss function representing the background environment The loss function representing the inserted target. They have the same representation form, which is:
[0059]
[0060] In the formula, C i ={C bg , C fg}, which respectively represent the rendering outputs of the background environment (C bg ) and the foreground target (C bg ), respectively represent the background environment and the foreground target of the real images, λ lpips represents the weight coefficient, LPIPS represents the Learned Perceptual Image Patch Similarity (LPIPS) loss function, |||| 1 represents the L1 loss function.
[0061] Step 6: After the training is completed, obtain the combined scene three-dimensional Gaussian model, and use the combined scene three-dimensional Gaussian model renderer to perform combined scene rendering to complete the three-dimensional Gaussian combined scene characterization optimization task.
[0062] According to the preset optimized fine-tuning training rounds, terminate the training process in Step 5 to obtain a combined scene three-dimensional Gaussian model with clear boundaries and reduced floating noise; use the combined scene three-dimensional Gaussian model renderer to render the fine-tuned combined scene to generate high-quality three-dimensional scene rendering images.
[0063] Through the above embodiments, the present invention can significantly improve the accuracy of the boundary region and the overall quality of the rendering result in the multi-three-dimensional Gaussian combined scene editing task, and is applicable to the efficient generation and optimization of various three-dimensional scenes.
[0064] Example
[0065] In this example, an indoor room scene is selected as the background scene, and an outdoor vase is selected as the inserted target for experiments, as Figure 3As shown in Column 1. The background scene is an indoor room, containing basic furniture such as a table, chairs, and a table lamp. The light source is uniform ceiling pendant lighting, which is collected by surrounding with a high-definition RGB camera. The shooting angle covers 120° in the front view, with a total of 99 images, and the image resolution is 1061×1886. The insertion target is an outdoor vase with a volume of approximately 20cm×20cm×40cm, the main color is yellow, and the bottleneck area is decorated with brown wood grain. It is collected by a high-definition RGB camera, the shooting angle covers 360°, with a total of 185 images, and the image resolution is 1297×840. Based on the above image data, use the structure from motion and multi-view stereo methods to generate an initial point cloud, and construct a three-dimensional Gaussian model of the environmental background in the form of three-dimensional Gaussian representation. Based on these image data, generate an initial point cloud of the insertion target, and construct a three-dimensional Gaussian model of the insertion target in the form of three-dimensional Gaussian representation.
[0066] According to Steps 2 to 3, construct a three-dimensional Gaussian model and the corresponding renderer for the environmental background and the insertion target respectively, and combine the three-dimensional Gaussian model of the insertion target with the three-dimensional Gaussian model of the environmental background through spatial alignment to generate a three-dimensional Gaussian model of the combined scene. The placement position of the inserted target vase is the three-dimensional position coordinates at the center of the table in the indoor scene
[0067] According to Steps 4 to 6, construct independent rendering paths for the foreground insertion target and the background environment respectively, and perform a total of 5000 steps of optimization and fine-tuning training on the combined scene. During the optimization and training process, calculate the loss functions of the foreground insertion target and the background environment respectively. After the optimization and fine-tuning training is completed, use the renderer of the combined scene three-dimensional Gaussian model to generate the final three-dimensional scene rendering image.
[0068] The experimental results compare the effects of three different combination methods, as Figure 3 shown.
[0069] The traditional combination method directly splices the insertion target with the target domain scene and then renders it, without optimizing the edge Gaussian ellipsoid of the insertion target, as Figure 3 shown in Column 2. The results show that there are obvious floating noise points in the edge area of the inserted target vase, especially in the transition area where it contacts the desktop. The noise causes an unnatural sense of break and texture distortion. The gloss on the surface of the vase does not match the background lighting conditions, the boundary is blurred and details are missing, and the overall fusion effect of the inserted target and the scene is poor.
[0070] When only optimizing the background, the optimization training focuses on the refinement of the background Gaussian model, without specifically optimizing the edge Gaussian ellipsoid of the insertion target, as Figure 3As shown in Column 3. The experimental results show that the lighting and texture details in the background part have been improved, but there are still blur problems in the edge areas of the inserted object. Specifically, the shadow transition at the contact between the vase and the table is unnatural, the boundary area lacks detail clarity, and the floating noise has not been effectively removed, affecting the visual consistency of the inserted object.
[0071] Through the dual-branch three-dimensional Gaussian rendering optimization strategy, the edge areas of the inserted object and the target domain scene have been significantly improved, as Figure 3 shown in Column 4. The results show that under the optimization of the dual-path rendering refinement strategy, the floating noise of the Gaussian ellipsoid at the edge of the inserted object vase has been effectively eliminated, and the edge transition is natural and clear. Especially in the area where the vase contacts the table, the shadow transition is soft and highly consistent with the table texture, showing good geometric alignment and light and shadow adaptation effects. The observation results from multiple perspectives further verify the optimization effect, and there are no artifacts or boundary blurs.
[0072] In summary, it is difficult for traditional combination methods and methods that only optimize the background to solve the edge noise and blur problems of the inserted object. The dual-branch three-dimensional Gaussian rendering optimization strategy effectively improves the clarity and texture consistency of the Gaussian ellipsoid at the edge of the inserted object through independent foreground and background optimization paths, achieving seamless integration of the inserted object and the scene. The experimental results verify the technical advantages of the present invention in the multi-three-dimensional Gaussian characterization combination task, providing technical guarantee for high-quality three-dimensional scene rendering.
[0073] It should be understood that the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A dual-branch three-dimensional Gaussian rendering optimization method, characterized in that: The following steps are involved: Step 1: Obtain the multi-view image data of the inserted target and the environmental background, perform preprocessing, and obtain the initialization point cloud of the inserted target and the initialization point cloud of the environmental background; Step 2: Based on the three-dimensional Gaussian representation form, the multi-view image data of the inserted target and the environmental background obtained in step 1 and the initialization point cloud are used as input to construct the three-dimensional Gaussian model of the inserted target and the environmental background and the corresponding three-dimensional Gaussian model renderer respectively; Step 3: Using the three-dimensional Gaussian models of the inserted target and the environmental background obtained in step 2, spatially align and integrate the inserted target and the environmental background to generate a three-dimensional Gaussian model of the combined scene; Step 4: Using the three-dimensional Gaussian model renderer obtained in step 2 to insert the target and the environment background, construct independent foreground rendering paths and background rendering paths respectively; Step 5: Through the foreground rendering path and the background rendering path, according to the inserted target mask image and the environment background mask image extracted from the corresponding multi-view image data, respectively optimize and fine-tune the combined scene 3D Gaussian model; The two fine-tuning trainings share the same total loss function, which is constructed based on the rendering output and the true value of the two fine-tuning training models; Step 6: After the training is completed, a three-dimensional Gaussian model of the combined scene is obtained, and the combined scene is rendered using a three-dimensional Gaussian model renderer to complete the three-dimensional Gaussian combined scene representation optimization task.
2. A dual-branch three-dimensional Gaussian rendering optimization method according to claim 1, characterized in that: In the step 1, the multi-view image data is collected by a visible light imaging device; the visible light imaging device adopts a digital camera, a mobile phone camera or a drone, and the collection range covers the complete geometric structure of the inserted target and the environmental background.
3. A dual-branch three-dimensional Gaussian rendering optimization method according to claim 1, characterized in that: In the step 1, the multi-view image data is preprocessed using a structure-from-motion and multi-view stereo method SfM-MVS.
4. A dual-branch three-dimensional Gaussian rendering optimization method according to claim 1, characterized in that: In step 3, the three-dimensional Gaussian models of the inserted target and the environmental background are spatially aligned and integrated using three-dimensional coordinates to obtain a three-dimensional Gaussian model of the combined scene; the combination process directly splices the parameters of the inserted target and the background Gaussian ellipsoid.
5. A dual-branch three-dimensional Gaussian rendering optimization method according to claim 1, characterized in that: In step 5, the mask image of the inserted target and background environment is obtained by an automatic or semi-automatic annotation tool, or by manual annotation.
6. A dual-branch three-dimensional Gaussian rendering optimization method according to claim 1, characterized in that: In step 6, the total loss function of fine-tuning training is optimized for: In the formula, Represents the loss function of the background environment, represents the loss function of inserting the target; C bg represents the rendering output of the current round when the background rendering path is used to optimize and fine-tune the combined scene 3D Gaussian model. fg Represents the rendering output of the current round when optimizing and fine-tuning the combined scene 3D Gaussian model using the foreground rendering path; represents the true value of the background environment, represents the true value of the inserted target, λ lpips represents the weight coefficient, LPIPS represents the learning perceptual image block similarity loss function, and ||||1 represents the L1 loss function.